Cross-scale urban population movement behavior generation method and model
Through cross-scale urban population movement behavior generation methods and models, the problem of difficulty in modeling and predicting the complexity of large-scale urban population movement is solved in the existing technology, and more efficient urban movement simulation and more accurate data generation are achieved.
Patent Information
- Application Number
- CN202510123171.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The existing technology is difficult to effectively model and predict the complexity of large-scale urban population movement, resulting in poor performance of urban mobile simulation systems and insufficient fidelity and practicality of simulation data.
The cross-scale urban population movement behavior generation method and model are adopted. The behavior generation model includes a traffic-based trajectory generator, a multi-scale discriminator and a collaborative evaluator to learn individual movement patterns and group flow laws, generate simulated individual and group movement trajectories, and optimize model parameters through reward signals and dominant functions.
It can more comprehensively capture the complexity of urban movement, including the heterogeneity of individual movement, interactions between individuals and complex relationships between individuals and environments, generate real human trajectories that conform to individual movement patterns and accurately reflect group flow, and improve the performance of urban movement simulation systems and the quality of simulation data.
Smart Images

Figure CN120046789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of urban population movement modeling and prediction, and particularly to a method and model for generating cross-scale urban population movement behaviors. Background Art
[0002] The movement of humans is crucial for the operation of cities. It connects various socio-economic activities in different regions, while also causing traffic congestion and air pollution, and promoting the spread of infectious diseases. These dynamic processes reflect two aspects of urban human movement. First, the daily individual movement of each person in time and space, i.e., the trajectory, reveals the way of life and service consumption of each individual in the city. Second, the large-scale movement flow of the overall urban population constitutes the daily rhythm of urban activities, which has an important impact on both individual life and environmental quality.
[0003] Existing studies have shown that the distributions of waiting times and jump lengths of individual movements exhibit long-tailed characteristics, which suggests the relevance of the continuous-time random walk (CTRW) model. Subsequent studies have found that there are repeated visit patterns to a few high-frequency visited locations in human trajectories, indicating that in addition to the above scale-free random walk, there are also regular visit patterns. This discovery inspired the design of the well-known exploration and preference return (EPR) model, which describes human movement as two simple competing mechanisms: exploring new locations and preferentially returning to previously visited locations. Compared with individual movement mainly relying on personal preferences, the anisotropic movement between urban regions is mainly affected by the trade-off between relative accessibility and transportation costs, which can be well captured by the law of gravity. In addition, the distance-frequency power-law law of population movement further allows for fine prediction of the flow related to different visit frequencies.
[0004] Although existing physical models and deep learning methods have been accepted in applications, the simplified descriptions they adopt cannot comprehensively capture the complexity of urban movement. First, empirical data shows that individual travel habits have high heterogeneity, such as the long-tailed radius distribution of trajectories. However, the general mechanisms in these models (such as the single exploration or return process in the EPR model) cannot depict the similarity and diversity of human trajectories. Second, when simulating the impact of the urban environment on population movement, in addition to the simple linear relationship between population and distance in the gravity model, more factors and more complex relationships should be considered.
[0005] Therefore, how to effectively model the movement complexity of large-scale cities to improve the performance of urban movement simulation systems and enhance the fidelity and practicality of simulation data is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a method and model for generating cross-scale urban population movement behaviors to solve at least one of the above-mentioned problems.
[0007] To achieve the above object, the present invention adopts the following scheme:
[0008] According to a first aspect of the present invention, there is provided a method for generating cross-scale urban population movement behaviors, which is executed by a behavior generation model. The generation model includes a flow-based trajectory generator, a multi-scale discriminator, and a collaborative evaluator. The method includes: obtaining the true first individual movement trajectory and the corresponding first group movement flow of the urban population; the trajectory generator uses the first individual movement trajectory and the first group movement flow to learn the individual movement pattern and the group flow law, and generates a simulated second individual movement trajectory; aggregating the second individual movement trajectory into a simulated second group movement flow; the multi-scale discriminator compares the first individual movement trajectory and the second individual movement trajectory, and the first group movement flow and the second group movement flow respectively at the individual level and the group level, and gives reward signals at the individual level and the group level based on the comparison results; the collaborative evaluator evaluates the degree of compliance of the generated second individual movement trajectory with the individual movement pattern and the second group movement flow with the group flow law respectively according to the reward signals, and gives the advantage function of the proximal policy optimization algorithm according to the evaluation results; updating the model parameters of the trajectory generator according to the advantage function and the reward signals.
[0009] As an embodiment of the present invention, the behavior generation model in the above method includes a first optimization objective and a second optimization objective. The first optimization objective is used to generate a true movement decision-making strategy that is statistically similar to the extracted empirical strategy; the second optimization objective is used to minimize the error between the generated flow and the ground truth data.
[0010] The function of the first optimization objective is as follows:
[0011]
[0012] The function of the second optimization objective is as follows:
[0013]
[0014] Where: L dist represents the distance between two distributions; L error represents the loss of the prediction result; P data represents the true individual trajectory distribution; F t,data represents the true group flow distribution; (l t |x <t ) represents the given historical movement x <tIn the case of, accessing location l at time t t The probability of.
[0015] As an embodiment of the present invention, in the above method, the trajectory generator uses the first individual movement trajectory and the first group movement flow to learn the individual movement pattern and the group flow law, and generates a simulated second individual movement trajectory, including: inputting the first individual movement trajectory into a gated recurrent unit network to learn the representation of the individual historical movement trajectory; converting the first group movement flow into a regional-level flow distribution; based on the representation of the individual historical movement trajectory and the regional-level flow distribution, generating the second individual movement trajectory in two stages, where the first stage determines the region and the second stage selects a specific location within the determined region.
[0016] As an embodiment of the present invention, in the above method, the multi-scale discriminator compares the first individual movement trajectory and the second individual movement trajectory, and the first group movement flow and the second group movement flow from the individual level and the group level respectively, and gives reward signals at the individual level and the group level based on the comparison results, including: the individual-level discriminator in the multi-scale discriminator inputs the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron respectively, and obtains corresponding individual-level reward signals; the group-level discriminator in the multi-scale discriminator compares the first group movement flow and the second group movement flow, and obtains a corresponding group-level reward signal by comparing the average error of the two flow matrices.
[0017] As an embodiment of the present invention, in the above method, the collaborative evaluator evaluates the degree of compliance of the generated second individual movement trajectory with the individual movement pattern and the second group movement flow with the group flow law respectively according to the reward signals, and gives the advantage function of the proximal policy optimization algorithm according to the evaluation results, including: the individual evaluation network in the collaborative evaluator evaluates the matching degree of the generated second individual movement trajectory with the individual movement pattern in the empirical data by learning to predict the true value function of the state-action, and outputs the individual value based on the Monte Carlo method; the group evaluation network in the collaborative evaluator evaluates the degree of compliance of the generated second group movement flow with the group flow law in the empirical data, and decouples the joint space optimization by decomposing the flow authenticity evaluation value contributed by the entire population into the sum of the group evaluation networks of each individual, and outputs the group value; gives the advantage function of the proximal policy optimization algorithm according to the individual value and the group value.
[0018] According to a second aspect of the present invention, a cross-scale urban population movement behavior generation model is provided. The behavior generation model includes: a data acquisition unit, an aggregation unit, a parameter update unit, a flow-based trajectory generator, a multi-scale discriminator, and a collaborative evaluator; the data acquisition unit is used to acquire the real individual movement trajectories of urban population and the corresponding first group movement flows; the trajectory generator is used to utilize the first individual movement trajectories and the first group movement flows to learn individual movement patterns and group flow laws, and generate simulated second individual movement trajectories; the aggregation unit is used to aggregate the second individual movement trajectories into simulated second group movement flows; the multi-scale discriminator is used to compare the first individual movement trajectories and the second individual movement trajectories, and the first group movement flows and the second group movement flows respectively at the individual level and the group level, and give reward signals at the individual level and the group level based on the comparison results; the collaborative evaluator is used to evaluate the degree of conformity of the generated second individual movement trajectories to the individual movement patterns and the generated second group movement flows to the group flow laws respectively according to the reward signals, and give the advantage function of the proximal policy optimization algorithm according to the evaluation results; the parameter update unit is used to update the model parameters of the trajectory generator according to the advantage function and the reward signals.
[0019] As an embodiment of the present invention, the above-mentioned behavior generation model includes a first optimization objective and a second optimization objective. The first optimization objective is used to generate a real movement decision-making strategy that is statistically similar to the extracted empirical strategy; the second optimization objective is used to minimize the error between the generated flow and the ground truth data.
[0020] The function of the first optimization objective is as follows:
[0021]
[0022] The function of the second optimization objective is as follows:
[0023]
[0024] Where: L dist represents the distance between two distributions; L error represents the loss of the prediction result; P data represents the real individual trajectory distribution; F t,data represents the real group flow distribution; (l t |x <t ) represents the probability of accessing the location l <t at time t given the historical movement x t .
[0025] As an embodiment of the present invention, the above-mentioned trajectory generator includes: a trajectory learning module, configured to input the first individual movement trajectory into a gated recurrent unit network to learn the representation of the individual's historical movement trajectory; a traffic conversion module, configured to convert the first group movement traffic into a traffic distribution at the regional level; a trajectory generation module, configured to generate the second individual movement trajectory in two stages based on the representation of the individual's historical movement trajectory and the traffic distribution at the regional level, where the first stage determines the region and the second stage selects a specific location within the determined region.
[0026] As an embodiment of the present invention, the above-mentioned multi-scale discriminator includes an individual-level discriminator and a group-level discriminator: the individual-level discriminator is configured to input the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron respectively, and obtain corresponding individual-level reward signals respectively; the group-level discriminator is configured to compare the first group movement traffic and the second group movement traffic, and obtain a corresponding group-level reward signal by comparing the average error of the two traffic matrices.
[0027] As an embodiment of the present invention, the above-mentioned collaborative evaluator includes an individual evaluation network, a group evaluation network, and an advantage function acquisition module; the individual evaluation network is configured to evaluate the matching degree between the generated second individual movement trajectory and the individual movement pattern in the empirical data by learning to predict the true value function of the state-action, and output an individual value based on the Monte Carlo method; the group evaluation network is configured to evaluate the compliance degree between the generated second group movement traffic and the group traffic law in the empirical data, and decouple the joint space optimization by decomposing the flow authenticity evaluation value contributed by the entire population into the sum of the group evaluation networks of each individual, and output a group value; the advantage function acquisition module is configured to give an advantage function of the proximal policy optimization algorithm according to the individual value and the group value.
[0028] According to a third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps of the above method are implemented.
[0029] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0030] As can be seen from the above technical solutions, the cross-scale urban population movement behavior generation method and model provided by the present invention overcome the limitations of the simplified descriptions of existing models and can capture the complexity of urban movement more comprehensively, including the heterogeneity of individual movement, the interaction between individuals, and the complex relationship between individuals and the environment. It can not only generate real human trajectories that conform to individual movement patterns (including spatial movement laws and diurnal preferences in time), but also accurately reflect the group flow and precisely display the daily rhythm of urban activities. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0032] Figure 1 is a schematic flowchart of a cross-scale urban population movement behavior generation method provided by an embodiment of the present application;
[0033] Figure 2 is a schematic flowchart of the trajectory generator generating a simulated second individual movement trajectory provided by an embodiment of the present application;
[0034] Figure 3 is a schematic structural diagram of a flow-based trajectory generator provided by an embodiment of the present application;
[0035] Figure 4 is a schematic flowchart of the multi-scale discriminator giving reward signals at the individual level and the group level provided by an embodiment of the present application;
[0036] Figure 5 is a schematic diagram of the function design of the trajectory discriminator and the flow discriminator provided by an embodiment of the present application;
[0037] Figure 6 is a schematic flowchart of the collaborative evaluator giving the advantage function of the proximal policy optimization algorithm provided by an embodiment of the present application;
[0038] Figure 7 is a schematic diagram of the network architecture of the individual evaluation network and the group evaluation network provided by an embodiment of the present application;
[0039] Figure 8 is the overall framework diagram of the cross-scale urban population movement behavior generation method provided by an embodiment of the present application;
[0040] Figure 9 is a schematic structural diagram of a cross-scale urban population movement behavior generation model provided by an embodiment of the present application;
[0041] Figure 10 is a schematic structural diagram of a trajectory generator provided by an embodiment of the present application;
[0042] Figure 11 is a schematic structural diagram of a multi-scale discriminator provided by an embodiment of the present application;
[0043] Figure 12 is a schematic structural diagram of a collaborative evaluator provided by an embodiment of the present application;
[0044] Figure 13 is a schematic block diagram of the system composition of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further describes the embodiments of the present invention in detail with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0046] As Figure 1 shown is a schematic flow diagram of a cross-scale urban population movement behavior generation method provided by an embodiment of the present application. This method is executed by a behavior generation model, and the generation model includes a flow-based trajectory generator, a multi-scale discriminator, and a collaborative evaluator. The method includes the following steps:
[0047] Step S101: Obtain the real first individual movement trajectory of the urban population and the corresponding first group movement flow.
[0048] The first individual movement trajectory and the first group movement flow collected in this step are a small amount of real data. Here, the small amount is relative to the subsequently generated simulated second individual movement trajectory. The individual movement trajectory data can include, for example, mobile phone positioning data, transportation card data, or GPS data, etc. The group movement flow data can include, for example, road network flow data, regional population density data, etc. After obtaining these data, they can be cleaned and preprocessed to ensure quality and reliability.
[0049] Step S102: Use the first individual movement trajectory and the first group movement flow to learn the individual movement pattern and the group flow law, and generate a simulated second individual movement trajectory.
[0050] This step is executed by the trajectory generator, which is a parametric neural network that can learn individual movement patterns (spatial movement laws, diurnal preferences over time) and group flow laws. That is, the trajectory generator of this application not only considers individual preferences (e.g., exploration and preference return in the EPR model), but also takes into account the influence of group flow. Therefore, it can overcome the limitations of existing models that ignore individual heterogeneity and group influence. In addition, the trajectory generator of this application adopts a hierarchical structure, first selecting regions and then specific locations to solve the problem of scale mismatch.
[0051] Step S103: Aggregate the second individual movement trajectory into a simulated second group movement flow.
[0052] This step aggregates the generated individual trajectories into group flow data. For example, it can be aggregated by counting by region or by section. This application does not limit the specific method of aggregation and can use any existing aggregation method to execute this step.
[0053] Step S104: Compare the first individual movement trajectory and the second individual movement trajectory, and the first group movement flow and the second group movement flow at the individual level and the group level respectively, and give reward signals at the individual level and the group level based on the comparison results.
[0054] This step is executed by the multi-scale discriminator. The multi-scale discriminator evaluates the quality of the generated trajectories and gives reward signals at the individual level (trajectory similarity) and the group level (flow similarity) respectively. It makes up for the deficiency that existing deep learning methods are mainly used for simple object modeling, enabling this application to process complex data at the urban scale.
[0055] Step S105: Evaluate the degree of compliance of the generated second individual movement trajectory with the individual movement pattern and the second group movement flow with the group flow law respectively according to the reward signals, and give the advantage function of the proximal policy optimization algorithm according to the evaluation results.
[0056] This step is executed by the collaborative evaluator. The collaborative evaluator combines the reward signals at the individual and group levels to calculate the advantage function of the PPO algorithm. The advantage function considers individual value (evaluating the matching degree of individual trajectories and empirical data based on the Monte Carlo method) and group value (evaluating the compliance degree of group flow and empirical data), and calculates through the advantage function. The core of this step is to establish the connection between individual-level travel decisions and group-level flows, solving the problem of ignoring the interaction between individuals and groups in existing models.
[0057] Step S106: Update the model parameters of the trajectory generator according to the advantage function and the reward signals.
[0058] As can be seen from the above technical solutions, the cross-scale urban population movement behavior generation method provided by the present invention overcomes the limitations of the simplified descriptions of existing models and can capture the complexity of urban movement more comprehensively, including the heterogeneity of individual movement, the interaction between individuals, and the complex relationship between individuals and the environment. It can not only generate real human trajectories that conform to individual movement patterns (including spatial movement laws and diurnal preferences of time), but also accurately reflect the group flow and precisely display the daily rhythm of urban activities.
[0059] Preferably, the above behavior generation model includes a first optimization objective and a second optimization objective, wherein the first optimization objective is used to generate a real movement decision-making strategy that is statistically similar to the extracted empirical strategy; the second optimization objective is used to minimize the error between the generated flow and the ground truth data.
[0060] In this application, the input of the behavior generation model is the movement trajectories of a small number of individuals and the group movement flow in the city, and the output is the full-scale movement trajectories of the city and the group movement flow.
[0061] The function of the first optimization objective of the behavior generation model is as follows in formula (1):
[0062]
[0063] The function of the second optimization objective is as follows in formula (2):
[0064]
[0065] Where: L dist represents the distance between two distributions, for example, it can be the Kullback-Leibler divergence and the Jensen-Shannon divergence; L error represents the loss of the prediction result, for example, it can be the mean square error, the mean absolute error, etc.; P data represents the real individual trajectory distribution; F t,data represents the real group flow distribution; (l t |x <t ) represents the probability of visiting location l <t at time t under the condition of the given historical movement x t .
[0066] Preferably, as Figure 2 shown, in the above step S102, the trajectory generator uses the first individual movement trajectory and the first group movement flow to learn the individual movement pattern and the group flow law, and the generated simulated second individual movement trajectory includes:
[0067] Step S1021: Input the moving trajectory of the first individual into a gated recurrent unit network to learn the representation of the individual's historical moving trajectory.
[0068] Step S1022: Convert the first group's moving flow into a flow distribution at the regional level.
[0069] Step S1023: Generate the second individual's moving trajectory in two stages based on the representation of the individual's historical moving trajectory and the flow distribution at the regional level, where the first stage determines the region and the second stage selects a specific location within the determined region.
[0070] Specifically, for the structural schematic diagram of the flow-based trajectory generator provided by the embodiments of the present invention, reference can be made to Figure 3 . As Figure 3 can be seen, the individual trajectory is represented by a sequence of positions, while the population flow is described with a coarser spatial resolution, that is, how many people move from one region to another region. To solve the problem of scale mismatch, the embodiments of the present invention adopt a generator with a hierarchical structure, where in the first stage, the individual decides the next region to go to; in the second stage, the individual selects a position belonging to that region. Obviously, the region selection in the first stage is consistent with the spatial resolution of the population flow, that is, the regional level. Therefore, the embodiments of the present invention model the interaction from the population to the individual at this stage.
[0071] The following combines Figure 3 to elaborate on the first stage and the second stage in the above steps in detail:
[0072] (1) The first stage:
[0073] The model first receives historical trajectory data for a period of time as input, Figure 3 and the historical trajectory shown in Figure 3 is represented as a dot sequence. Then the historical trajectory data is processed by a gated recurrent unit (GRU) to extract the feature information in the time series and generate a state representation. The output of the GRU is fed into a multi-head attention mechanism module, which contains multiple attention heads (Head 1, Head 2,..., Head k). Each attention head focuses on different aspects of the historical trajectory and captures information at different time scales and patterns. The output of the multi-head attention mechanism contains an uncertainty variable (var) used to represent the degree of uncertainty of the model's prediction of the region selection. "High" and "low" represent two cases of high uncertainty and low uncertainty, which reflects the model's ability to express the confidence level of the prediction result. Finally, the model outputs a probability distribution representing the possibility of the individual moving to each region. This probability distribution is calculated based on the output of the multi-head attention mechanism and the uncertainty variable. Figure 3 shows two example regions, and the high and low uncertainty regions are inFigure 3 It has also been reflected in
[0074] (2) The second stage:
[0075] The regional information selected in the first stage is converted into a vector representation through an Embedding layer. Then the regional embedding information is concatenated with the state representation generated in the first stage and processed again by GRU to fuse regional information and time series information. The output of GRU is transformed through a multi-layer perceptron (MLP) to extract higher-level features. The output of MLP is converted into a probability distribution through a Softmax layer, representing the possibility of an individual moving to various locations within the selected region. Figure 3 It shows multiple possible next locations and the probability of each location. The final output is the next location and the next action (the next location and region selection together determine the next action).
[0076] Preferably, as Figure 4 shown, in the above step S104, the multi-scale discriminator compares the first individual movement trajectory and the second individual movement trajectory, the first group movement flow and the second group movement flow respectively from the individual level and the group level, and gives reward signals at the individual level and the group level based on the comparison results, including:
[0077] Step S1041: The individual-level discriminator in the multi-scale discriminator inputs the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron respectively, and obtains corresponding individual-level reward signals respectively.
[0078] Step S1042: The group-level discriminator in the multi-scale discriminator compares the first group movement flow and the second group movement flow, and obtains a corresponding group-level reward signal by comparing the average error of the two flow matrices.
[0079] As can be seen from the above, the embodiment of the present invention defines two-level reward functions, comprehensively evaluates the generation results at the individual level and the group level respectively, and all intermediate steps obtain zero rewards. The present invention models these two reward functions respectively, where the first individual-level reward function is modeled by a discriminator based on a neural network, while the second group-level reward function is directly given. As Figure 5 shows the design of these two-level functions, by Figure 5It can be seen that there are a trajectory discriminator and the above-mentioned individual-level discriminator, a traffic discriminator and the above-mentioned group-level discriminator. Among them: The trajectory discriminator is used to evaluate the authenticity of the generated individual trajectories. It receives real trajectories (Real) and generated trajectories (Fake) as inputs. The input data first passes through a Sigmoid layer, then through a multi-layer perceptron (MLP), and finally the features of the real trajectories and the generated trajectories are concatenated (Concat), and finally outputs a probability value for judging whether the generated trajectory is real. This part uses a neural network-based discriminator to judge the quality of the generated trajectory by learning the differences between real trajectories and generated trajectories. The traffic discriminator is used to evaluate the authenticity of the generated group traffic. It receives real traffic and generated traffic as inputs and calculates the distance (Distance) between the two. Figure 5 The distribution of real traffic and generated traffic is shown in the form of an image, and the accuracy of the generated traffic is measured by calculating the distance between the two. The reward function of this part is directly given without using a neural network for modeling, but directly calculating a certain distance metric (such as Euclidean distance or other suitable distance metrics) between real traffic and generated traffic. The smaller the distance, the closer the generated traffic is to the real traffic and the higher the reward.
[0080] It can be seen that Figure 5 The multi-scale discriminator shown evaluates the generation results from the individual and group levels respectively through the trajectory discriminator and the traffic discriminator, and gives corresponding reward signals to guide the training of the trajectory generator, and finally improves the authenticity and accuracy of the generated trajectories and group traffic. Compared with the prior art, the present application more comprehensively evaluates the quality of the generation results and improves the performance of the model.
[0081] Preferably, as Figure 6 shown, in the above step S105, the collaborative evaluator respectively evaluates the degree of conformity of the generated second individual movement trajectory to the individual movement pattern and the generated second group movement traffic to the group traffic law according to the reward signal, and the advantage function of the proximal policy optimization algorithm given according to the evaluation result may include:
[0082] Step S1051: The individual evaluation network in the collaborative evaluator evaluates the matching degree between the generated second individual movement trajectory and the individual movement pattern in the empirical data by learning to predict the true value function of the state-action, and outputs the individual value based on the Monte Carlo method.
[0083] Step S1052: The group evaluation network in the collaborative evaluator evaluates the degree of conformity of the generated second group movement traffic to the group traffic law in the empirical data, and decouples the joint space optimization by decomposing the evaluation value of the flow authenticity contributed by the entire population into the sum of the group evaluation networks of each individual, and outputs the group value.
[0084] Step S1053: Give the advantage function of the proximal policy optimization algorithm according to the individual value and the group value.
[0085] As Figure 7 shown, it shows the network architectures of the above individual evaluation network and group evaluation network. The individual evaluation network shares the same state encoder with the trajectory generator, including an embedding layer and a GRU layer, for obtaining state representations. Then, the embodiment of the present invention uses an MLP to predict the individual state value and outputs the individual value. The group evaluation network is modeled by another MLP network, which receives the concatenation of the state embedding (O N ) and the action embedding (a N ) as input. Specifically, the state embedding is obtained through the network of the individual evaluation network, and the action embedding is obtained through the embedding layer.
[0086] The cross-scale urban population movement behavior generation method proposed by the present invention is to establish the connection between the travel decision-making at the individual level and the aggregated flow at the population level. To strengthen the representation and learning of the movement model, the present invention designs the following two individual-population cooperation mechanisms.
[0087] I. Cooperative movement representation learning
[0088] In the city, people's actions are not independent of each other. The preferences of individuals and collective behaviors will both affect the travel decisions of individuals. For example, the trajectory of an individual shows his / her daily routine activities, indicating that future travel behaviors are affected by historical activities. At the same time, an individual will occasionally visit some unfamiliar locations recommended by friends, which indicates the existence of social interactions between individuals. To consider the individual and population-level factors cooperatively, the present invention designs the population movement model π θ (a t |o t ) as the cooperative cooperation of two parameterized decision-making processes:
[0089]
[0090] Among them, and represent two different decision-making strategies considering individual preferences and population impacts respectively. The first part captures the preferences of individuals by learning the movement rules in the historical trajectory X <t , while the second part describes the impact of population flow shaped by the urban environment on social interactions. The cooperative design between these two parts makes a discrete choice between them according to the parameterized Bernoulli distribution Bernoulli(u t ), where Represents the learning probability of an individual's uncertainty in following their own preferences. Specifically, this strategy can be expressed by the following formula:
[0091]
[0092] If an individual has a high level of uncertainty about historical visits, they are more likely to follow the population-level decision Rather than personal preferences The incorporation of this state uncertainty enables each individual to flexibly and effectively utilize collective decisions.
[0093] II. Cooperative Movement Generation Optimization
[0094] In addition to the representation of the population movement model, the present invention further enhances the individual and population-level cooperative learning methods. Similar to GAIL, the optimization of π θ is guided by an evaluation network V φ which evaluates the value behind each state-action pair. To encourage the policy network to collaboratively capture the individual travel patterns and the regularity of population flow, the present invention divides the evaluation network into the above-mentioned individual evaluation network and group evaluation network, providing accurate valuations related to the current trajectory and the aggregated flow respectively and The individual evaluation network learns to predict the correct value function of the state-action, based on whether π θ (a t |o t ) matches the individual movement patterns in the empirical data. The objective is the empirical average return obtained by the Monte Carlo method.
[0095] For evaluating the authenticity of the flow that assesses the contribution of the entire population, directly optimizing the evaluation network is not feasible because of the high dimensions of the joint actions and observations. Therefore, to solve this problem, the present invention proposes to decouple the joint space optimization by decomposing the from the entire population into the sum of the group evaluation networks of each individual. Formally expressed as:
[0096]
[0097] Then, the present invention uses Monte Carlo policy evaluation to optimize the above-mentioned
[0098] Finally, the individual evaluation network and the group evaluation network jointly evaluate the individual strategy, evaluate its ability to capture travel patterns and its further impact on population travel, providing the optimal direction for improving the current movement model, and using the advantage function of the proximal policy optimization algorithm (PPO) for model optimization as follows:
[0099]
[0100] where η is the weight for balancing the cooperation of the two evaluators.
[0101] The overall framework diagram of the above cross-scale urban population movement behavior generation method can also be seen Figure 8 , for Figure 8 which has been mentioned in the above description and will not be elaborated here.
[0102] As can be seen from the above, the cross-scale urban population movement behavior generation method provided by the present invention overcomes the limitations of the simplified description of existing models and can capture the complexity of urban movement more comprehensively, including the heterogeneity of individual movement, the interaction between individuals, and the complex relationship between individuals and the environment. It can not only generate real human trajectories that conform to individual movement patterns (including spatial movement laws and diurnal preferences of time), but also accurately reflect the group flow and precisely display the daily rhythm of urban activities. In addition, by establishing the connection between individual-level travel decisions and population-level aggregated flows, the present application enhances the representation and learning ability of the model, thereby improving the performance of the urban movement simulation system and the fidelity and practicality of the simulation data. Moreover, the present application adopts a deep collaborative learning framework based on generative adversarial imitation learning, and through the collaborative work of three parameterized neural networks (flow-based trajectory generator, multi-scale movement discriminator, and collaborative evaluator), it realizes the high-precision simulation of complex urban movement behaviors. The quality assessment provided by the discriminator and the application of the proximal policy optimization algorithm (PPO) further improve the performance of the trajectory generator. Finally, the present application considers multi-scale factors, can model at both the individual trajectory and group flow levels, and connects the two through a collaborative mechanism, thereby more comprehensively reflecting the complexity of urban movement.
[0103] As Figure 9 shown in the structural schematic diagram of a cross-scale urban population movement behavior generation model provided by an embodiment of the present application, the behavior generation model includes: a data acquisition unit 910, a flow-based trajectory generator 920, an aggregation unit 930, a multi-scale discriminator 940, a collaborative evaluator 950, and a parameter update unit 960. Among them:
[0104] The data acquisition unit 910 is used to acquire the real first individual movement trajectory of the urban population and the corresponding first group movement flow.
[0105] The trajectory generator 920 is used to learn the individual movement pattern and the group flow law by using the first individual movement trajectory and the first group movement flow, and generate a simulated second individual movement trajectory.
[0106] The aggregation unit 930 is used to aggregate the second individual movement trajectory into a simulated second group movement flow.
[0107] The multi-scale discriminator 940 is used to compare the first individual movement trajectory and the second individual movement trajectory, the first group movement flow and the second group movement flow at the individual level and the group level respectively, and give reward signals at the individual level and the group level based on the comparison results.
[0108] The collaborative evaluator 950 is used to evaluate the degree of compliance of the generated second individual movement trajectory with the individual movement pattern and the second group movement flow with the group flow law according to the reward signal respectively, and give the advantage function of the proximal policy optimization algorithm according to the evaluation results.
[0109] The parameter update unit 960 is used to update the model parameters of the trajectory generator according to the advantage function and the reward signal.
[0110] Preferably, the above-mentioned behavior generation model includes a first optimization objective and a second optimization objective. The first optimization objective is used to generate a real movement decision-making strategy that is statistically similar to the extracted empirical strategy; the second optimization objective is used to minimize the error between the generated flow and the ground truth data.
[0111] The function of the first optimization objective is as follows:
[0112]
[0113] The function of the second optimization objective is as follows:
[0114]
[0115] Where: L dist represents the distance between two distributions; L error represents the loss of the prediction result; P data represents the real individual trajectory distribution; F t,data represents the real group flow distribution; (l t |x <t ) represents the probability of accessing the location l <t at time t given the historical movement x t .
[0116] Preferably, as Figure 10 shown, the above-mentioned trajectory generator 920 includes:
[0117] A trajectory learning module 921, which is used to input the first individual movement trajectory into a gated recurrent unit network to learn the representation of the individual historical movement trajectory;
[0118] A flow conversion module 922, which is used to convert the first group movement flow into a flow distribution at the regional level;
[0119] A trajectory generation module 923, configured to generate the second individual movement trajectory in two stages based on the representation of the individual historical movement trajectory and the traffic distribution at the regional level, where the first stage determines the region and the second stage selects a specific location within the determined region.
[0120] Preferably, as Figure 11 shown, the above-mentioned multi-scale discriminator 940 includes an individual-level discriminator 941 and a group-level discriminator 942:
[0121] The individual-level discriminator 941 is configured to input the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron respectively, and obtain corresponding individual-level reward signals respectively.
[0122] The group-level discriminator 942 is configured to compare the first group movement flow and the second group movement flow, and obtain a corresponding group-level reward signal by comparing the average error of the two flow matrices.
[0123] Preferably, as Figure 12 shown, the above-mentioned collaborative evaluator 950 includes an individual evaluation network 951, a group evaluation network 952, and an advantage function acquisition module 953.
[0124] The individual evaluation network 951 is configured to evaluate the matching degree between the generated second individual movement trajectory and the individual movement pattern in the empirical data by learning to predict the true value function of the state-action, and output an individual value based on the Monte Carlo method.
[0125] The group evaluation network 952 is configured to evaluate the degree of compliance of the generated second group movement flow with the group flow law in the empirical data, and decouple the joint space optimization by decomposing the flow authenticity evaluation value contributed by the entire population into the sum of the group evaluation networks of each individual, and output a group value.
[0126] The advantage function acquisition module 953 is configured to give an advantage function of the proximal policy optimization algorithm according to the individual value and the group value.
[0127] For the detailed descriptions of the above-mentioned various units and modules, reference can be made to the corresponding descriptions in the foregoing method embodiments, and details are not described herein again.
[0128] As can be seen from the above technical solutions, the cross-scale urban population movement behavior generation model provided by the present invention overcomes the limitations of the simplified descriptions of existing models and can capture the complexity of urban movement more comprehensively, including the heterogeneity of individual movement, the interaction between individuals, and the complex relationship between individuals and the environment. It can not only generate real human trajectories that conform to individual movement patterns (including spatial movement laws and diurnal preferences of time), but also accurately reflect group flows and precisely display the daily rhythm of urban activities.
[0129] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.
[0130] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program for executing the above method.
[0131] As Figure 13 shown, the electronic device 600 may further include: a communication module 110, an input unit 120, an audio processor 130, a display 160, and a power supply 170. It should be noted that the electronic device 600 does not necessarily have to include Figure 13 all the components shown in Figure 13 ; in addition, the electronic device 600 may further include Figure 13 components not shown in
[0132] As Figure 13 shown, the central processing unit 100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 100 receives inputs and controls the operations of the various components of the electronic device 600.
[0133] Among them, the memory 140, for example, may be one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store the above information related to failures, and can also store programs for executing relevant information. And the central processing unit 100 can execute the program stored in the memory 140 to implement information storage or processing, etc.
[0134] The input unit 120 provides inputs to the central processing unit 100. The input unit 120 is, for example, a key or a touch input device. The power supply 170 is used to supply power to the electronic device 600. The display 160 is used for displaying display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0135] The memory 140 may be a solid-state memory, for example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that stores information even when power is off, can be selectively erased and has more data. An example of this memory is sometimes referred to as an EPROM, etc. The memory 140 may also be some other type of device. The memory 140 includes a buffer memory 141 (sometimes referred to as a buffer). The memory 140 may include an application / function storage unit 142 for storing application programs and function programs or for storing the processes for operating the electronic device 600 by the central processor 100.
[0136] The memory 140 may also include a data storage unit 143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 144 of the memory 140 may include various drivers for the communication functions of the electronic device and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0137] The communication module 110 is a transmitter / receiver that transmits and receives signals via the antenna 111. The communication module 110 (transmitter / receiver) is coupled to the central processor 100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.
[0138] Based on different communication technologies, multiple communication modules 110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 110 (transmitter / receiver) is also coupled to the speaker 131 and the microphone 132 via the audio processor 130 to provide an audio output via the speaker 131 and receive an audio input from the microphone 132, thereby implementing normal telecommunication functions. The audio processor 130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 130 is also coupled to the central processor 100, so that recording can be performed on the local machine through the microphone 132, and the sound stored on the local machine can be played through the speaker 131.
[0139] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0140] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks
[0141] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks
[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks
[0143] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention
Claims
1. A method for generating cross-scale urban population mobility behavior, characterized in that: The method comprises: Obtain the real first individual movement trajectory and the corresponding first group movement flow of the urban population; Using the first individual movement trajectory and the first group movement flow, learning the individual movement pattern and the group flow law, and generating a simulated second individual movement trajectory; aggregating the second individual movement trajectory into a simulated second group movement flow; Comparing the first individual movement trajectory and the second individual movement trajectory, the first group movement flow and the second group movement flow at the individual level and the group level respectively, and providing reward signals at the individual level and the group level based on the comparison results; Evaluate and generate the degree of conformity of the second individual movement trajectory to the individual movement mode and the second group movement flow to the group flow rule according to the reward signal, and provide the advantage function of the proximal strategy optimization algorithm according to the evaluation results; The trajectory generation parameters are updated according to the advantage function and the reward signal.
2. The method for generating cross-scale urban population mobility behavior according to claim 1, characterized in that: The behavior generation model includes a first optimization objective and a second optimization objective, wherein the first optimization objective is used to generate a real movement decision strategy that is statistically similar to the extracted empirical strategy; and the second optimization objective is used to minimize the error between the generated flow and the ground truth data; The function of the first optimization objective is as follows: The function of the second optimization objective is as follows: Where: L dist Represents the distance between two distributions; L error Represents the loss of prediction results; P data represents the real individual trajectory distribution; F t,data Represents the real group flow distribution; (l t |x <t ) indicates that given the historical movement x <t In the case of visiting position l at time t t probability.
3. The method for generating cross-scale urban population mobility behavior according to claim 1, characterized in that: The using the first individual movement trajectory and the first group movement flow to learn the individual movement pattern and the group flow rule, and generating a simulated second individual movement trajectory includes: Inputting the first individual movement trajectory into a gated recurrent unit network to learn the representation of the individual historical movement trajectory; Converting the first group of mobile traffic into regional-level traffic distribution; Based on the representation of the individual historical movement trajectory and the traffic distribution at the area level, the second individual movement trajectory is generated in two stages, wherein the first stage determines the area and the second stage selects a specific location within the determined area.
4. The method for generating cross-scale urban population mobility behavior according to claim 1, characterized in that: The comparing the first individual movement trajectory and the second individual movement trajectory, the first group movement flow and the second group movement flow at the individual level and the group level respectively, and providing individual level and group level reward signals based on the comparison results includes: Inputting the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron, respectively, and obtaining corresponding individual-level reward signals; The first group movement flow rate is compared with the second group movement flow rate, and the corresponding group-level reward signal is obtained by comparing the average errors of the two flow matrices.
5. The method for generating cross-scale urban population mobility behavior according to claim 1, characterized in that: The step of respectively evaluating and generating the degree of conformity of the second individual movement trajectory to the individual movement mode and the second group movement flow to the group flow rule according to the reward signal, and providing the advantage function of the proximal strategy optimization algorithm according to the evaluation result includes: By learning the true value function of the predicted state-action, and based on the Monte Carlo method, evaluating the matching degree between the generated second individual movement trajectory and the individual movement pattern in the empirical data, the individual value is output; Evaluate the conformity of the generated second group mobility flow with the group flow law in the empirical data, decouple the joint spatial optimization by decomposing the mobility authenticity evaluation value contributed by the entire population into the sum of the group evaluation network of each individual, and output the group value; An advantage function of a proximal strategy optimization algorithm is given according to the individual value and the group value.
6. A cross-scale urban population mobility behavior generation model, characterized by: The behavior generation model includes: a data acquisition unit, an aggregation unit, a parameter updating unit, a flow-based trajectory generator, a multi-scale discriminator, and a collaborative evaluator; The data acquisition unit is used to acquire the real first individual movement trajectory and the corresponding first group movement flow of the urban population; The trajectory generator is used to learn the individual movement pattern and the group flow law by using the first individual movement trajectory and the first group movement flow, and generate a simulated second individual movement trajectory; The aggregation unit is used to aggregate the second individual movement trajectory into a simulated second group movement flow; The multi-scale discriminator is used to compare the first individual movement trajectory and the second individual movement trajectory, the first group movement flow and the second group movement flow at the individual level and the group level respectively, and provide reward signals at the individual level and the group level based on the comparison results; The collaborative evaluator is used to evaluate the conformity of the second individual movement trajectory to the individual movement mode and the second group movement flow to the group flow law according to the reward signal, and to provide an advantage function of the proximal strategy optimization algorithm according to the evaluation results; The parameter updating unit is used to update the model parameters of the trajectory generator according to the advantage function and the reward signal.
7. The cross-scale urban population mobility behavior generation model according to claim 6, characterized in that: The behavior generation model includes a first optimization objective and a second optimization objective, wherein the first optimization objective is used to generate a real movement decision strategy that is statistically similar to the extracted empirical strategy; and the second optimization objective is used to minimize the error between the generated flow and the ground truth data; The function of the first optimization objective is as follows: The function of the second optimization objective is as follows: Where: L dist Represents the distance between two distributions; L error Represents the loss of prediction results; P data represents the real individual trajectory distribution; F t,data Represents the real group flow distribution; (l t |x <t ) indicates that given the historical movement x <t In the case of visiting position l at time t t probability.
8. The cross-scale urban population mobility behavior generation model according to claim 6, characterized in that: The trajectory generator comprises: A trajectory learning module, used for inputting the first individual movement trajectory into a gated recurrent unit network to learn the representation of the individual historical movement trajectory; A traffic conversion module, used for converting the first group of mobile traffic into regional level traffic distribution; The trajectory generation module is used to generate the second individual movement trajectory in two stages based on the representation of the individual historical movement trajectory and the traffic distribution at the regional level, wherein the first stage determines the region and the second stage selects a specific location within the determined region.
9. The cross-scale urban population mobility behavior generation model according to claim 6, characterized in that: The multi-scale discriminator includes an individual level discriminator and a group level discriminator: The individual level discriminator is used to input the first individual movement trajectory and the second individual movement trajectory into a network composed of a gated recurrent unit and a multi-layer perceptron, respectively, and obtain corresponding individual level reward signals respectively; The group-level discriminator is used to compare the first group movement flow and the second group movement flow, and obtain a corresponding group-level reward signal by comparing the average errors of the two flow matrices.
10. The cross-scale urban population mobility behavior generation model according to claim 6, characterized in that: The collaborative evaluator includes an individual evaluation network, a group evaluation network and an advantage function acquisition module; The individual evaluation network is used to evaluate the matching degree between the generated second individual movement trajectory and the individual movement pattern in the experience data by learning the real value function of the predicted state-action based on the Monte Carlo method, and output the individual value; The group evaluation network is used to evaluate the conformity of the generated second group mobile flow with the group flow law in the empirical data, decouple the joint spatial optimization by decomposing the flow authenticity evaluation value contributed by the entire population into the sum of the group evaluation network of each individual, and output the group value; The advantage function acquisition module is used to provide an advantage function of a proximal strategy optimization algorithm according to the individual value and the group value.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Moving track generation method and device
CN113255951A
Crowd movement modeling method and device based on access location clustering
CN115994313A
Travel trajectory generation and deep learning network model and training method
CN116245973A
Urban people flow prediction method and device based on activity space and gravity model
CN116362422A
Human body moving track recovery method and device
CN117992916A