User trajectory generation method and device based on generative adversarial imitation learning

CN116030332BActive Publication Date: 2026-09-25TOYOTA JIDOSHA KK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111246324.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2026-09-25
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

[0005]本公开实施例提供了一种基于生成对抗模仿学习的用户轨迹生成方法及装置,能够解决现有技术中用户轨迹生成精度低、需要大量采样用户轨迹数据的问题

Benefits of technology

[0055]本公开的各种实施例提供的基于生成对抗模仿学习的用户轨迹生成方法及装置,基于马尔科夫决策过程构建更加全面的用户移动行为模型,并结合生成对抗模仿学习得到用户轨迹生成模型,可以生成高精度的用户移动轨迹,且可以基于生成的用户移动轨迹进行用户移动轨迹预测、城市规划等,实用性强,不需要采集大量的真实数据样本。本公开实施例中,通过语义感知的状态转换模型与提取的决策特征相结合,与移动轨迹数据进行交互,使得模型的数据层和决策层能够一定程度得到解耦合,使得用户轨迹生成模型能够捕获不同用户的本质决策策略,而不会受到时空差异、用户差异等因素的干扰。此外,本公开实施例中,不仅考虑用户的低维特征还考虑用户的高层特征,从而构建由即时回报函数和长时回报函数组成的多尺度回报函数,可以生成接近用户真实轨迹的移动轨迹,进一步提高用户轨迹生成的精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030332B_ABST
    Figure CN116030332B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a user trajectory generation method and device based on generative adversarial imitation learning, the method comprising: constructing a policy function network and a reward function network based on generative adversarial imitation learning; training the policy function network and the reward function network based on a user's historical moving trajectory to obtain a user trajectory generation model; inputting a user's current moving trajectory into the user trajectory generation model to generate the user's moving trajectory, wherein the state information at least includes a first user attribute state and a second user attribute state. The user trajectory generation method based on generative adversarial imitation learning provided in the present disclosure constructs a more comprehensive user moving behavior model based on Markov decision process, and obtains a user trajectory generation model in combination with generative adversarial imitation learning, which can capture the essential decision strategy of different users without being disturbed by factors such as space-time difference and user difference, and generate high-precision user moving trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of user trajectory generation technology, specifically to a user trajectory generation method and apparatus based on generative adversarial imitation learning. Background Technology

[0002] Artificially synthesized trajectories are helpful for practical applications such as network service optimization and transportation scheduling. For example, in cellular networks, synthetic trajectories can be used to simulate the detailed processes of network user movement and communication, thereby obtaining network reliability performance analysis. Similarly, we can use artificially synthesized trajectories to simulate traffic congestion before and after the implementation of certain policies (e.g., road widening) for urban planning purposes.

[0003] However, synthesizing human trajectories is challenging due to their high scale, complex spatiotemporal correlations, and the randomness of migration trajectories. The rise of deep learning paradigms offers promising solutions for synthesizing high-quality human trajectories, with the most successful and prominent approach being based on Generative Adversarial Networks (GANs). GANs utilize a generator to synthesize new data and a discriminator to distinguish between generated and real data; then, the network is trained through a game between the generator and discriminator networks to generate data with high similarity to real data. Existing techniques leverage the powerful modeling capabilities of GANs, combining them with CNNs and RNNs to synthesize human trajectories. However, human trajectories are generated by complex human decision-making processes, while GANs are designed to learn data distributions directly from demonstrations without modeling the hidden decision-making processes. This results in lower model accuracy, a lack of interpretability, and poor performance in data transfer. Therefore, how to comprehensively improve model accuracy and rationally design the network structure is a pressing issue.

[0004] The existing trajectory synthesis work has the following limitations: (1) Due to the large differences in actual application environments, the model may not be able to achieve transferability and cannot be used for derivative applications; (2) Since human trajectories are more complex and higher dimensional, using a few physical quantities to predict human trajectories will not meet the requirements for the accuracy of generated trajectories, and the feasibility in practical applications is not high; (3) Existing technologies require a large amount of data to model their trajectories, but in reality, big data collection is difficult. Summary of the Invention

[0005] This disclosure provides a user trajectory generation method and apparatus based on generative adversarial imitation learning, which can solve the problems of low accuracy of user trajectory generation and the need for a large amount of user trajectory data sampling in the prior art.

[0006] According to one of the solutions disclosed herein, a user trajectory generation method based on generative adversarial imitation learning is provided, comprising:

[0007] Construct a policy function network and a reward function network based on generative adversarial imitation learning;

[0008] The policy function network and reward function network are trained based on the user's historical movement trajectory to obtain the user trajectory generation model;

[0009] The user's status information is input into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least a first user attribute status and a second user attribute status. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location.

[0010] In some embodiments, before constructing the policy function network and reward function network based on generative adversarial imitation learning, the method further includes constructing a user mobility behavior model based on a Markov decision process, specifically including:

[0011] Determine the user's state space, wherein the state space includes a time state, a first user attribute state, and a second user attribute state;

[0012] Determine the user's action space, which includes dwelling, residence return, preference return, and exploration;

[0013] Based on the state space and the action space, a user's state transition function is constructed, wherein the user's current state s is represented by the user's current state s. i Perform action a i When, the next state s is reached. i+1 The probability distribution.

[0014] In some embodiments, the user's current state includes the user's current location. i At that time, a user's state transition function is constructed based on the state space and the action space, including:

[0015] When the user's current action is staying or returning to their residence, the next state s i+1 The probability of being in the current location is determined;

[0016] When the user's current action is preference regression, the next state s i+1 The location is determined based on the following probability distribution:

[0017] P(l i+1 =l|s i ,a i=Return)=k l / k all

[0018] Where, k l k represents the number of times a user visits location l. all This represents the total number of times a user visits all locations;

[0019] When the user's current action is exploration, sort the locations the user has not visited according to their distance from the current location, and determine the next state s. i+1 The probability distribution of visiting a new location, expressed as:

[0020] P(l i+1 =l|s i ,a i =Explore)∝k(l,l i ) -α

[0021] Where, k(l,l) i ) represents the sorting size, and the parameter α represents the user's sensitivity to distance.

[0022] In some embodiments, the construction of the policy function network includes:

[0023] Based on the user's mobile behavior model, features are extracted from the user's historical mobile trajectory to obtain the user's target state features;

[0024] The target state features are processed to obtain a dense representation vector;

[0025] A nonlinear network is extracted from the dense representation vector using a self-attention mechanism;

[0026] The nonlinear network is normalized using the softmax activation function, and the policy result is output, wherein the policy result includes the probability distribution of each action in the corresponding state.

[0027] In some embodiments, the construction of the reward function network includes:

[0028] Construct an instant reward function R I ;

[0029] Construct a long-term reward function R L ;

[0030] A multi-scale return function is constructed based on the immediate return function and the long-term return function, wherein the multi-scale return function is expressed as:

[0031] R M =R I +λRL

[0032] Where λ is a parameter that balances the impact of immediate returns and long-term returns, and λ>0.

[0033] In some embodiments, constructing the instant reward function includes:

[0034] The user's state features and action features are encoded, and the state features and action features are mapped into feature vectors that characterize the state features and action features;

[0035] The feature vector is input into a linear layer of the ReLU activation function to extract low-dimensional features;

[0036] The low-dimensional features are processed using the Sigmoid activation function to obtain the instant reward function; wherein, the instant reward function R I Represented as:

[0037] R I (s i ,a i ) = logD I (s is ,a i )

[0038] Among them, D I (s is ,a i () indicates the immediate identification result.

[0039] In some embodiments, constructing the long-term reward function includes:

[0040] Obtain the first and second high-level features from the user state features;

[0041] Based on the mutual information between the first high-level feature and the second high-level feature, fit the first probability distribution of the first high-level feature under the second high-level feature;

[0042] Calculate the KL divergence of the first probability distribution and the second probability distribution based on the first probability distribution and the second probability distribution of the first high-level feature;

[0043] The long-term reward function is obtained based on the discrimination results of the KL divergence and the second high-level feature.

[0044] Optionally, the first high-rise feature includes the user's residential address h. i The second high-level feature includes the user's historical movement trajectory.

[0045] The long-term reward function is expressed as follows:

[0046]

[0047] in, For trajectory identification results, For h i and The KL divergence after fitting the mutual information between the two, β>0, is a hyperparameter used to adjust the influence of high-level user characteristics.

[0048] In some embodiments, the policy function network and reward function network are trained based on the user's historical movement trajectory to obtain a user trajectory generation model, including:

[0049] The policy function network and reward function network are pre-trained using an optimization objective based on the negative log-likelihood function;

[0050] The policy function network is trained by maximizing the total reward of the generated trajectory obtained by downsampling according to the policy function network decision, and the reward function network is trained by optimizing the adversarial optimization value of the reward function network.

[0051] According to one of the solutions disclosed herein, a user trajectory generation device based on generative adversarial imitation learning is also provided, comprising:

[0052] The building module is configured to construct a policy function network and a reward function network based on generative adversarial imitation learning;

[0053] The training module is configured to train the policy function network and reward function network based on the user's historical movement trajectory to obtain a user trajectory generation model;

[0054] The generation module is configured to input the user's status information into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least a first user attribute status and a second user attribute status. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location.

[0055] The user trajectory generation method and apparatus based on generative adversarial imitation learning provided in various embodiments of this disclosure construct a more comprehensive user movement behavior model based on Markov decision processes and combine it with generative adversarial imitation learning to obtain a user trajectory generation model. This model can generate high-precision user movement trajectories and can be used for user movement trajectory prediction, urban planning, etc., based on the generated user movement trajectories. It is highly practical and does not require the collection of a large number of real data samples. In the embodiments of this disclosure, a semantically aware state transition model is combined with extracted decision features and interacts with movement trajectory data. This allows the data layer and decision layer of the model to be decoupled to a certain extent, enabling the user trajectory generation model to capture the essential decision-making strategies of different users without being affected by spatiotemporal differences, user differences, or other factors. Furthermore, in the embodiments of this disclosure, not only low-dimensional features of users are considered, but also high-level features, thereby constructing a multi-scale reward function composed of an immediate reward function and a long-term reward function. This can generate movement trajectories that closely approximate the user's real trajectory, further improving the accuracy of user trajectory generation. Attached Figure Description

[0056] Figure 1 A flowchart illustrating a user trajectory generation method based on generative adversarial imitation learning according to an embodiment of this disclosure is shown.

[0057] Figure 2 A schematic diagram showing the various actions included in the action space of an embodiment of this disclosure;

[0058] Figure 3 This diagram illustrates the system architecture of a user trajectory generation model according to an embodiment of the present disclosure.

[0059] Figure 4 A schematic diagram illustrating the construction of the policy function network according to an embodiment of this disclosure is shown;

[0060] Figure 5 A schematic diagram illustrating the construction of the reward function network according to an embodiment of this disclosure is shown;

[0061] Figure 6 This diagram illustrates a comparison of heatmaps between trajectories generated using different user trajectory generation methods and actual trajectories.

[0062] Figure 7 A schematic diagram of the structure of a user trajectory generation device based on generative adversarial imitation learning according to an embodiment of the present disclosure is shown. Detailed Implementation

[0063] Various embodiments and features of this disclosure are described herein with reference to the accompanying drawings.

[0064] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this disclosure will be apparent to those skilled in the art.

[0065] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present disclosure and, together with the general description of the disclosure given above and the detailed description of the embodiments given below, serve to explain the principles of the disclosure.

[0066] These and other features of this disclosure will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.

[0067] It should also be understood that although this disclosure has been described with reference to specific examples, many other equivalent forms of this disclosure can be readily implemented by those skilled in the art.

[0068] The above and other aspects, features and advantages of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.

[0069] Specific embodiments of this disclosure are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this disclosure, which may be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure this disclosure. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely to serve as the basis and representative basis for the claims to teach those skilled in the art to use this disclosure in a variety of substantially any suitable detailed structures.

[0070] Figure 1 A flowchart illustrating a user trajectory generation method based on generative adversarial imitation learning, according to an embodiment of this disclosure, is shown. Figure 1 As shown, this disclosure provides a user trajectory generation method based on generative adversarial imitation learning, including:

[0071] S101: Construct a policy function network and a reward function network based on generative adversarial imitation learning.

[0072] Generative Adversarial Imitation Learning (GAIL) is a deep learning neural network whose core idea originates from Nash equilibrium in game theory. It consists of a generator (G) and a discriminator (D). In this embodiment, the constructed policy function network serves as the generator, and the reward function network serves as the discriminator.

[0073] Prior to step S101, the method further includes constructing a user mobile behavior model based on a Markov decision process, specifically including:

[0074] S201: Determine the user's state space S, wherein the state space includes a time state, a first user attribute state, and a second user attribute state; wherein the first user attribute state is a spatial state, including the user's location; the second user attribute state includes user-related feature states other than the user's location.

[0075] S202: Determine the user's action space A, wherein the action space includes dwelling, residence return, preference return, and exploration;

[0076] S203: Construct a user's state transition function based on the state space and the action space, wherein the user's state transition function represents the user's current state s. i Perform action a i When, the next state s is reached. i+1 The probability distribution.

[0077] In this embodiment, a policy function is constructed based on a Markov Decision Process (MDP) of mobile behavior, and then a reward function is constructed. Specifically, each user can be considered as an "agent," possessing its own unique coarse-grained reward function. A Markov Decision Process (MDP) is a mathematical framework for decision-making, consisting of a quadruple, i.e.<S,A,P,R> Where S represents the MDP state space; A represents the MDP decision space; P: S×A×S→P represents the state transition probability of the MPD, P(s) i+1 |s i ,a i ) indicates the current state s i ∈S and the currently executing action a i Given ∈A, the next state s i+1 The probability distribution; R: S×A→R is the reward function of the MDP, R(s) i ,a i ) indicates the current state s i Next, execute action a i Instant rewards received under certain circumstances.

[0078] In step S201, the user's state space S is mainly characterized by its spatiotemporal location and is described in two dimensions: time and space. In addition to spatiotemporal information, there are other important factors that affect the user's movement trajectory. Therefore, in this embodiment of the disclosure, the state space S is defined as a triple s = (l, t, c), where l represents the user's current location (first attribute state), t represents the current time, and c represents other important decision state features besides spatiotemporal information (second attribute state).

[0079] Specifically, c can represent the effective features mined based on the user's historical movement trajectory, which may include the user's dwell time τ at the current location. i The number of locations the user has visited in history (m) i User's residential address h i Historical trajectory over a certain period of time These features have been proven to have a significant impact on user movement behavior in practical applications. In other words, in the subsequent construction of user trajectory generation models, effective features can be extracted using prior knowledge or by using the user's historical movement trajectory (expert movement trajectory) to extract effective features through the model's self-learning.

[0080] In step S202, the user's action space A is the decision space of the user's trajectory, which is a set of combinations containing all possible user actions and movements. For example... Figure 2 As shown, in this embodiment, the user's action space is defined as a set consisting of four actions: Stay, Home, Return, and Explore. Stay indicates that the user will remain at their current location in the next time unit; Home return indicates that the user will return to their home in the next time unit; Return indicates that the user will move to a previously visited location (excluding their home) in the next time unit; and Explore indicates that the user will visit a location they have not visited before.

[0081] In step S203, constructing the user's state transition function specifically involves solving for the state transition probability P. The state transition function describes the state transition probability P in the current state s. i With the currently executing action a i Under the condition, the next state s i+1 The probability distribution.

[0082] Since the state space defined in step S201 includes location l i Time t i Effective feature c i Three parts, of which t iThe transition is decisive; with a fixed increase of one time unit, the state changes accordingly; effective feature c i It depends on the historical movement trajectory, and it is only affected by the current state s in the condition variable. i Updated, therefore c i The transfer was also decisive.

[0083] For location l i :

[0084] when a i When the status is changed to a stay or a return to a residence, it can be determined that the next status will still be location l. i and the user's residential address h i That is, l i The transition remains decisive; at this point, the next state s i+1 The probability of the location is fixed at 1.

[0085] And when the user's current action is a i When performing preference regression or exploration, l i+1 Since it is random, this embodiment employs the classic Rank-based Exploration and Preferential Return (r-EPR) model to model the probability of this random distribution. Specifically, when a user decides to return to a previously visited location that is not included in their residence, they will select the location with the following probabilities:

[0086] P(l i+1 =l|s i ,a i =Return)=k l / k all

[0087] Where, k l k represents the number of times a user visits location l. all This represents the total number of times a user visits all locations.

[0088] When a user decides to explore a new location, locations the user has not visited before will be compared to the current location. i The distances are sorted, and then the user's next state s is determined based on the sorting. i+1 The probability distribution of visiting a new location, that is, the new location is selected according to the following probabilities:

[0089] P(l i+1 =l|s i ,a i =Explore)∝k(l,l i ) -α

[0090] Where, k(l,l) i The sorting size is α, which represents the user's sensitivity to distance. A smaller α indicates that the user is less sensitive to the exploration distance, and therefore, the user will be more likely to explore more distant locations.

[0091] The state transition function constructed in step S203 is the initial policy function. In the subsequent construction of the policy function network, a policy function that can predict the user's next action is constructed based on this state transition function.

[0092] Understandably, a reward function is also included in the standard MDP framework, specifically the reward function R(s). i ,a i ) measures the state s i Next, execute action a i The reward received is the value earned when transitioning from state S to state S' by taking decision action A. In the MDP framework used in this disclosure, since the policy function is unknown, the reward function is also unknown and will be determined in the subsequent construction of the reward function network. That is, this disclosure integrates the MDP framework into the construction of the user trajectory generation model, resulting in a more accurate user trajectory generation model.

[0093] In some embodiments, such as Figure 4 As shown, in step S101, the construction of the policy function network includes:

[0094] S1011: Based on the user's mobile behavior model, extract features from the user's historical mobile trajectory to obtain the user's target state features;

[0095] S1012: Process the target state features to obtain a dense representation vector;

[0096] S1013: Extract a nonlinear network from the dense representation vector using a self-attention mechanism;

[0097] S1014: Normalize the nonlinear network using the softmax activation function and output the policy result, wherein the policy result includes the probability distribution of each action in the corresponding state.

[0098] The user's historical movement trajectory (expert movement trajectory) is a pre-determined, real movement trajectory. The input to the policy function network is the user's state features. Specifically, based on the state features of the user's state space determined under the aforementioned MDP framework, the target state features, which serve as the input to the policy function network, are determined. The target state features include temporal features, spatial features, and at least one of the aforementioned effective features, in order to obtain hidden features that have a significant impact on the decision-making strategy.

[0099] In this embodiment, the user's target state features (original state features) that serve as input to the policy function network include the user's current location. i Current time t i User's residential address h i The user's dwell time at the current location τ i And the number of locations the user has visited in history (m) i In practice, the aforementioned state features can be extracted from the user's historical movement trajectory using feature extraction algorithms.

[0100] After obtaining the user's target state features through feature extraction in step S1011, the process proceeds to step S1012, where different types of input state feature data are encoded to convert the original state features s into a specific format used as input to the policy function network. This allows the original state features s (as described above) containing various types of attributes to be encoded. i t i Features such as these are converted into a unified input vector format.

[0101] For example, the original input state features s can be encoded using a one-hot code, with the encoding function determined based on learnable parameters. Then, the original input state features s are transformed into a representation s. e The dense representation vector. In this embodiment, the encoding of different original state features s can adopt a unified encoding method to improve computational efficiency.

[0102] like Figure 4 As shown, in this embodiment, data dimensionality reduction and dense representation can be achieved through an embedding layer, vectorizing the feature data of each state, and then feature fusion can be performed through a concatenation layer to obtain a dense representation vector s. e .

[0103] Furthermore, a policy function network based on a self-attention mechanism is constructed through step S1013.

[0104] Specifically, a scalar dot product attention network is used. The input to the attention network consists of a query vector q, a key vector k, and a value vector v, all of which are derived from the dense vector s using the following method. e Extracting independent nonlinear networks:

[0105] q,k,v=ReLU(W q S,W k S,W v S)

[0106] In matrix operations, the three vectors mentioned above are represented by the encoded dense representation vector s.e With three weight matrices W q W k W v S-multiplication creates the nonlinear network, which is then extracted using the ReLU activation function.

[0107] In specific implementation, such as Figure 4 As shown, it can be done through ⊙ (XOR operation), Tensor product Vector operations such as XOR are used to extract nonlinear networks.

[0108] In this embodiment, the ReLU activation function is used to nonlinearize the neural network. The gradient of the ReLU activation function is either 0 or 1, which can effectively avoid the problems of gradient vanishing and gradient exploding.

[0109] Finally, the generated vector y is processed using a linear layer with a softmax activation function in step S1014. i This allows us to obtain the probability distribution π(a|s) for each action, and use this probability distribution as the policy outcome (the output of the policy function network).

[0110] In constructing the policy function network, a self-attention mechanism is used to capture the complex correlations and regularities within the time-varying movement trajectory. The self-attention mechanism is highly effective in sequence modeling and can better model higher-order and long-term patterns of movement behavior, resulting in a more accurate policy function network.

[0111] As shown above, in this embodiment, the user's movement decision strategy is defined based on the different action spaces of the aforementioned semantics, thereby obtaining the user's essential decision strategy. Each agent / user's decision strategy is described by a policy function π: S→A. The policy function maps the user's state to an action, and then uses a state transition function to realize the transition to the next movement data (action). That is, the generator receives a real user state s and outputs the probability distribution π(a|s) of different user actions. After generating the decision strategy π(a|s), the network randomly selects an action (determining the flexible activity at the next location) based on the learned strategy, and generates the user's trajectory based on each action.

[0112] To improve the accuracy (precision) of user trajectory generation, it is further necessary to construct a reward function network, and obtain the final user trajectory generation model through adversarial imitation learning or adversarial inverse reinforcement learning between the policy function network and the reward function network.

[0113] In some embodiments, such as Figure 5 As shown, in step S101, the construction of the reward function network includes:

[0114] S1015: Constructing the instant reward function R I ;

[0115] S1016: Constructing the long-term reward function R L ;

[0116] S1017: Construct a multi-scale return function based on the instantaneous return function and the long-term return function, wherein the multi-scale return function is expressed as:

[0117] R M =R I +λR L

[0118] Where λ is a parameter that balances the impact of immediate returns and long-term returns, and λ>0.

[0119] In the GAIL framework, the reward function scores the generated trajectory by the policy function by matching the distribution of the generated state-action pairs with the distribution of expert movement data (user historical movement data). The reward function network is constructed as a binary classifier, taking both the generated trajectory and the expert movement trajectory as input, and outputting the discrimination result between the generated trajectory and the expert movement trajectory. That is, the discriminator receives a true user state s. i Realistic actions a i And the action a generated above i+1 And determine action a i+1 The probability of a true outcome is represented by the distribution of the generated state-action pairs. The more similar the distribution of the generated state-action pairs is to the expert distribution, the higher the score / reward will be.

[0120] However, due to the complexity and diversity of human movement, simple reward functions cannot capture the intricate information during movement and face problems such as compound error and covariance drift. Therefore, in this embodiment, user movement information is divided into two types: short-term actions and global movement actions. To better extract this information, immediate reward functions and long-term reward functions are constructed respectively, and a multi-scale reward function incorporating both immediate and long-term reward functions is also constructed.

[0121] Among them, the instant reward function R I The reward value is evaluated directly based on the current state-action sequence, while the long-term reward function R... L Used to distinguish the personal information (hidden information) of mobile users from their trajectories, and to assess their similarity to expert mobile data at the trajectory level.

[0122] In some embodiments, such as Figure 5As shown, construct the instant reward function, including:

[0123] A1: Encode the user's state features and action features, and map the state features and action features into feature vectors that characterize the state features and action features;

[0124] A2: Input the feature vector into a linear layer of the ReLU activation function to extract low-dimensional features;

[0125] A3: The low-dimensional features are processed using the Sigmoid activation function to obtain the instantaneous reward function;

[0126] Wherein, the instant reward function R I Represented as:

[0127] R I (s i ,a i ) = logD I (s is ,a i )

[0128] Among them, D I (s is ,a i () indicates the immediate identification result.

[0129] In this example, the discriminator D... I The authenticity of the generated state-action pairs is measured, i.e., whether the movement trajectory is generated by a real mobile user. To achieve this, we first use an encoding module to map states and actions into representation vectors that fully describe their features; then, this representation vector is input into a linear layer with a ReLU activation function to extract low-dimensional features; finally, it is passed through another linear layer using a Sigmoid function as the activation function to obtain the instantaneous discrimination result D. I (s is ,a i ).

[0130] In this embodiment, since the target state features of the input policy function network include the user's current location l i Current time t i User's residential address h i The user's dwell time at the current location τ i And the number of locations the user has visited in history (m) i Therefore, the state features corresponding to short-term actions are determined from these features and used as input to the discriminator to obtain the discrimination result D. I (l i ,t i ,τ i ,mi ,a i The instantaneous reward function constructed based on the above state characteristics can be expressed as follows:

[0131] R I (s i ,a i ) = logD I (l i ,t i ,τ i ,m i ,a i ).

[0132] Instant Discriminator Network D I While distinguishing between real and generated state-action pairs provides signals for the learning of the policy function, it fails to capture global information and high-level user characteristics, which are crucial for generating high-quality mobile data. Therefore, to capture this information, we define a long-term reward function at the trajectory level to measure the authenticity of the entire trajectory, rather than focusing solely on state-action pairs. Simultaneously, to ensure that high-level user characteristics truly play a significant role in guiding decision-making, we compute additional rewards based on mutual information to guarantee the recovery of user characteristics from the generated mobile trajectories.

[0133] In some embodiments, such as Figure 5 As shown, the long-term reward function is constructed, including:

[0134] B1: Obtain the first and second high-level features from the user state features;

[0135] B2: Based on the mutual information between the first high-level feature and the second high-level feature, fit the first probability distribution of the first high-level feature under the second high-level feature;

[0136] B3: Calculate the KL divergence of the first probability distribution and the second probability distribution based on the first probability distribution and the second probability distribution of the first high-level feature;

[0137] B4: The long-term reward function is obtained based on the discrimination results of the KL divergence and the second high-level feature.

[0138] Specifically, the user's current location i Current time t i Real-time features such as user address h can be obtained instantly. i Compared to information such as time and location, this information is more precise and can therefore be used as the first high-level feature. In this embodiment, to verify the authenticity of the generated user movement trajectory, the user's historical movement trajectory, which represents global information about movement, can be used. As the second highest level feature.

[0139] Using neural networks (such as convolutional neural networks, CNNs) to fit conditions on historical movement trajectories The distribution of user characteristics, i.e., fitting In this embodiment, maximization is used. Learning deep representations using mutual information between h and h is equivalent to minimizing the KL divergence. Where p(h) represents the user's residential address h i The probability distribution is obtained as input. Furthermore, since this disclosure aims to determine whether the generated user movement trajectory is genuine, historical movement trajectories are used... As input, the identification result of the generated movement trajectory is obtained.

[0140] That is, the inputs of the long-term discriminator are the historical movement trajectories. and user's residential address h i The output of the long-term discriminator is the trajectory discrimination result. and KL divergence

[0141] Therefore, in this embodiment, the long-term reward function can be expressed as:

[0142]

[0143] Where β>0, is a hyperparameter used to adjust the influence of high-level user characteristics. In the above equation, N d The number of time units is a given number, meaning that the user will only receive a long-term reward after completing a certain length of movement trajectory.

[0144] In other embodiments, the first high-level user characteristic of a user can be specifically determined according to actual needs, and this disclosure does not specifically limit it. For example, the first high-level user characteristic can be other specific access locations.

[0145] After constructing the policy function network and reward function network, a preliminary user trajectory generation model is established. For example... Figure 3 As shown, the constructed user trajectory generation model consists of a three-layer structure: a data layer, a semantic layer, and a decision layer. In the data layer, the user extracts features from the input movement trajectory data to obtain the target state features used for trajectory generation. The semantic layer is used to determine the semantically perceived state transition distribution based on the semantic action space. The decision layer is used for adversarial generative training to obtain the final user trajectory generation model.

[0146] S102: Train the policy function network and reward function network based on the user's historical movement trajectory to obtain the user trajectory generation model.

[0147] The process of training the model is also the process of continuously optimizing the user trajectory generation model composed of the policy function network and the reward function network.

[0148] During training, the generator's goal is to generate actions that are as realistic as possible. i+1 The goal of the discriminator is to deceive the discriminator into making its trajectory as close as possible to that of the expert, so that the discriminator cannot distinguish whether the trajectory was generated by the expert or the generator.

[0149] Overall, the user trajectory generation model is optimized by maximizing the total reward to select the optimal policy function. The policy function and reward function are jointly optimized through adversarial learning.

[0150] Specifically, in order to ensure that the user trajectory generation model can be trained efficiently and robustly, a two-step training method is adopted. First, pre-training is used to accelerate the training process so that the model quickly converges to a decent result. Then, formal reinforcement learning is used to optimize the objective so that the model converges to the optimal performance.

[0151] Step S1021: First, the preliminary user trajectory generation model (the policy function network and reward function network) is pre-trained using an optimization objective based on the negative log-likelihood function. Its optimization function (optimization of the loss function) is expressed as follows:

[0152]

[0153] Where, π E Indicate expert strategy, This represents the expectation derived from the expert's movement trajectory.

[0154] For any policy π, This represents the expectation obtained using the movement trajectory based on policy π sampling. Specifically, for any reward function f: S×A→R, Among them, a i ~π(·|s i ) and s i+1 ~P(·|s i ,a i ),thus, This represents the expectation derived from the expert's movement trajectory.

[0155] The above describes the optimization of the policy function network. Further, the policy (generated movement trajectory) obtained by the above optimization objective based on the negative log-likelihood function is sampled. The sampled generated trajectory is combined with the real trajectory to pre-train the immediate and long-term discriminators, thus completing the model pre-training process.

[0156] Step S1022: Train the policy function network by maximizing the total reward of the generated trajectory obtained by downsampling according to the policy function network decision, and train the reward function network by optimizing the adversarial optimization value of the reward function network.

[0157] Pre-training alone cannot guarantee the optimality of the final policy function, and because it is based on self-supervised training using the negative log-likelihood function as input, the resulting policy function is susceptible to exposure bias. Therefore, further reinforcement learning training is implemented to optimize the model.

[0158] Therefore, the optimization method of the policy function network is changed to maximizing the total reward of the generated trajectory obtained by downsampling according to its decision, while the reward function network (discriminator D) I and D L ) Optimize by optimizing its counter-optimization value.

[0159] Specifically, the optimization function of this part of the policy function network can be expressed as:

[0160]

[0161] in, This represents the expectation of the reward function with respect to the policy. This indicates the expectation of the identification result. H(π) represents the expectation only over the non-zero terms of the long-term reward function, and H(π) represents the causal entropy of policy π, which is defined as follows:

[0162] For the generator, in order to deceive the discriminator as much as possible, it needs to maximize the discrimination probability D of the generated trajectory. I (s,a), D L (s,a), that is, minimizing log(1-D) I (s,a) and log(1-D) L (s,a), then, calculate the expectation under the condition of maximizing the discrimination probability of the generated trajectory, and in the process of calculating the expectation, fully consider the immediate reward and the long-term reward, and train and optimize the model.

[0163] In the aforementioned pre-training and reinforcement learning training processes, both the generator and discriminator are trained alternately; that is, the discriminator is trained first, followed by the generator, and this process is repeated continuously. Specifically, the policy results generated by the policy function network and the user's historical movement trajectory are first input into the reward function network to train it. Then, the policy results are iteratively generated to train both the policy function network and the reward function network until they converge, thereby optimizing the model parameters. After multiple training iterations, the action 'a' generated by the generator... i+1 It is getting closer and closer to real action, that is, generating a given distribution of expert data based on generative adversarial simulation training.

[0164] In this embodiment, through reinforcement learning training in step S1022, flexible and varied strategies can be sought while ensuring the final reward function is maximized. By optimizing based on this objective function (loss function) and combining reinforcement learning update strategies such as PPO and TRPO, it is possible to ultimately learn the user's mobility decision-making strategy from expert mobility data. Using the decision-making strategy for mobility simulation will help us generate higher quality and more realistic mobility data.

[0165] The discriminator provides a reward for the generator's policy learning. This reward distinguishes the expert policy from the learned policy. After receiving the reward from the discriminator, the TRPO algorithm can be used for policy learning. The two processes alternate to complete the final training.

[0166] S103: Input the user's status information into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least the first user attribute status and the second user attribute status mentioned above. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location.

[0167] After continuously training and optimizing the model's parameters to obtain the final user trajectory generation model, the user's state information can be input into the model to generate the user's movement trajectory.

[0168] The state information input to the user trajectory generation model can include s = (l, c), where l can be the user's initial location information, c can be the user's residential address, and other feature states, etc.

[0169] In some embodiments, the state information further includes time state information, that is, the state information of the input user trajectory generation model includes s = (l, t, c), where l can be the user's initial location information, t is the user's initial time, c can be the user's residential address and other feature states, etc.

[0170] In the above embodiments, the initial location information and the initial time can be either historical time or current time, that is, the user's status information can be either historical state or current state, and this disclosure does not specifically limit it.

[0171] For example, l can be the user's current location and t can be the current time, so as to generate the user's movement trajectory based on the user's current state.

[0172] Figure 6 A schematic diagram comparing heatmaps of motion trajectories generated using the generative adversarial imitation learning-based user trajectory generation method provided in this disclosure, as well as the MoveSim and SeqGAN methods, with real motion trajectories is shown. Figure 6 As shown, the movement trajectory generated by this method is closer to the real movement trajectory, and its loss function has a loss value of only 0.027, which can generate high-precision user movement trajectories.

[0173] The method provided by the embodiments of the present invention can simulate different actions under different states and generate generated samples that can be used to predict user movement trajectories, etc. In this way, only a small amount of real data is needed to achieve accurate prediction of user movement trajectories or to carry out urban planning, etc., without the need to collect a large number of real data samples.

[0174] The user trajectory generation method based on generative adversarial imitation learning provided in this disclosure constructs a more comprehensive user movement behavior model based on Markov decision processes and combines it with generative adversarial imitation learning to obtain a user trajectory generation model. This model can generate high-precision user movement trajectories and, based on the generated user movement trajectories and a small amount of real data, can be used for user movement trajectory prediction, urban planning, etc., demonstrating strong practicality. This disclosure combines a semantically aware state transition model with extracted decision features and interacts with movement trajectory data, allowing the data layer and decision layer of the model to be decoupled to a certain extent. This enables the user trajectory generation model to capture the essential decision-making strategies of different users without being affected by spatiotemporal differences, user differences, or other factors. Furthermore, this disclosure considers not only low-dimensional user features but also high-level user features, thereby constructing a multi-scale reward function composed of immediate and long-term reward functions. This generates movement trajectories that closely approximate the user's actual trajectory, further improving the accuracy of user trajectory generation.

[0175] Figure 7A flowchart of a user trajectory generation apparatus based on generative adversarial imitation learning, according to an embodiment of this disclosure, is shown. Figure 7 As shown, this disclosure provides a user trajectory generation apparatus based on generative adversarial imitation learning, comprising:

[0176] Module 701 is configured to build a policy function network and a reward function network based on generative adversarial imitation learning;

[0177] Training module 702 is configured to train the policy function network and reward function network based on the user's historical movement trajectory to obtain a user trajectory generation model;

[0178] The generation module 703 is configured to input the user's status information into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least the first user attribute status and the second user attribute status mentioned above. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location.

[0179] In some embodiments, the user trajectory generation device further includes a user movement behavior model building module, configured to build a user movement behavior model based on a Markov decision process before building a policy function network and a reward function network based on generative adversarial learning, specifically configured as follows:

[0180] Determine the user's state space, wherein the state space includes a time state, a first attribute state, and a second attribute state;

[0181] Determine the user's action space, wherein the action space includes dwelling, residence return, preference return, and exploration;

[0182] Based on the state space and the action space, a user's state transition function is constructed, wherein the user's current state s is represented by the user's current state s. i Perform action a i When, the next state s is reached. i+1 The probability distribution.

[0183] In some embodiments, the user mobility behavior model building module is further configured to: [details about the user's current state and location]. i At that time, a user's state transition function is constructed based on the state space and the action space, specifically including:

[0184] When the user's current action is staying or returning to their residence, the next state s i+1 The probability of being in the current location is determined;

[0185] When the user's current action is preference regression, the next state s i+1 The location is determined based on the following probability distribution:

[0186] P(l i+1 =l|s i ,a i =Return)=k l / k all

[0187] Where, k l k represents the number of times a user visits location l. all This represents the total number of times a user visits all locations;

[0188] When the user's current action is exploration, sort the locations the user has not visited according to their distance from the current location, and determine the next state s. i+1 The probability distribution of visiting a new location, expressed as:

[0189] P(l i+1 =l|s i ,a i =Explore)∝k(l,l i ) -α

[0190] Where, k(l,l) i ) represents the sorting size, and the parameter α represents the user's sensitivity to distance.

[0191] In some embodiments, the construction module 701 is specifically configured to construct a policy function network, specifically including:

[0192] Based on the user's mobile behavior model, features are extracted from the user's historical mobile trajectory to obtain the user's target state features;

[0193] The target state features are processed to obtain a dense representation vector;

[0194] A nonlinear network is extracted from the dense representation vector using a self-attention mechanism;

[0195] The nonlinear network is normalized using the softmax activation function, and the policy result is output, wherein the policy result includes the probability distribution of each action in the corresponding state.

[0196] In some embodiments, the construction module 701 is further configured to construct a reward function network, specifically including:

[0197] Construct an instant reward function R I ;

[0198] Construct a long-term reward function RL ;

[0199] A multi-scale return function is constructed based on the immediate return function and the long-term return function, wherein the multi-scale return function is expressed as:

[0200] R M =R I +λR L

[0201] Where λ is a parameter that balances the impact of immediate returns and long-term returns, and λ>0.

[0202] Furthermore, construct an instant reward function, including:

[0203] The user's state features and action features are encoded, and the state features and action features are mapped into feature vectors that characterize the state features and action features;

[0204] The feature vector is input into a linear layer of the ReLU activation function to extract low-dimensional features;

[0205] The low-dimensional features are processed using the Sigmoid activation function to obtain the instant reward function; wherein, the instant reward function R I Represented as:

[0206] R I (s i ,a i ) = logD I (s is ,a i )

[0207] Among them, D I (s is ,a i () indicates the immediate identification result.

[0208] Furthermore, a long-term reward function is constructed, including:

[0209] Obtain the first and second high-level features from the user state features;

[0210] Based on the mutual information between the first high-level feature and the second high-level feature, fit the first probability distribution of the first high-level feature under the second high-level feature;

[0211] Calculate the KL divergence of the first probability distribution and the second probability distribution based on the first probability distribution and the second probability distribution of the first high-level feature;

[0212] The long-term reward function is obtained based on the discrimination results of the KL divergence and the second high-level feature.

[0213] Optionally, the first high-rise feature includes the user's residential address h. i The second high-level feature includes the user's historical movement trajectory.

[0214] The long-term reward function is expressed as follows:

[0215]

[0216] in, For trajectory identification results, For h i and The KL divergence after fitting the mutual information between the two, β>0, is a hyperparameter used to adjust the influence of high-level user characteristics.

[0217] In some embodiments, the training module 702 is specifically configured as follows:

[0218] The policy function network and reward function network are pre-trained using an optimization objective based on the negative log-likelihood function;

[0219] The policy function network is trained by maximizing the total reward of the generated trajectory obtained by downsampling according to the policy function network decision, and the reward function network is trained by optimizing the adversarial optimization value of the reward function network.

[0220] The user trajectory generation device based on generative adversarial imitation learning provided in this disclosure is based on the same concept as the user trajectory generation method based on generative adversarial imitation learning in the above embodiments. The specific implementation process is detailed in the method embodiments in the above embodiments, and will not be repeated here.

[0221] This disclosure also provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, implement the above-described method.

[0222] In some embodiments, the processor executing computer-executable instructions may be a processing device that includes one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor that runs other instruction sets, or a processor that runs a combination of instruction sets. The processor may also be one or more special-purpose processing devices, such as an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), a System-on-a-Chip (SoC), etc.

[0223] In some embodiments, a computer-readable storage medium may be a memory such as a read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash drives or other forms of flash memory, cache, registers, static memory, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape cassette or other magnetic storage devices, or any other possible non-transitory medium used to store information or instructions that can be accessed by a computer device.

[0224] The computer-executable instructions of embodiments of this disclosure can be organized into one or more computer-executable components or modules. Various aspects of this disclosure can be implemented with any number and combination of such components or modules. For example, aspects of this disclosure are not limited to the specific computer-executable instructions or particular components or modules shown in the drawings and described herein. Other embodiments may include different computer-executable instructions or components having more or fewer functions than those shown and described herein.

[0225] The above embodiments are merely exemplary embodiments of this disclosure and are not intended to limit this disclosure. The scope of protection of this disclosure is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this disclosure within its substance and scope, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this disclosure.

Claims

1. A user trajectory generation method based on generative adversarial imitation learning, comprising: Construct a policy function network and a reward function network based on generative adversarial imitation learning; The policy function network and reward function network are trained based on the user's historical movement trajectory to obtain the user trajectory generation model; The user's status information is input into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least a first user attribute status and a second user attribute status. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location. Prior to constructing the policy function network and reward function network based on generative adversarial imitation learning, the method further includes constructing a user mobility behavior model based on Markov decision processes, specifically including: Determine the user's state space, wherein the state space includes a time state, a first user attribute state, and a second user attribute state; Determine the user's action space, wherein the action space includes dwelling, residence return, preference return, and exploration; A user's state transition function is constructed based on the state space and the action space, wherein the user's state transition function represents the user's current state. Perform actions When, reach the next state The probability distribution; The construction of the policy function network includes: Based on the user's mobile behavior model, features are extracted from the user's historical mobile trajectory to obtain the user's target state features; The target state features are processed to obtain a dense representation vector; A nonlinear network is extracted from the dense representation vector using a self-attention mechanism; The nonlinear network is normalized using the softmax activation function, and the policy result is output, wherein the policy result includes the probability distribution of each action in the corresponding state.

2. The method according to claim 1, wherein, The user's current state includes the user's current location. At that time, a user's state transition function is constructed based on the state space and the action space, including: When the user's current action is staying or returning to their residence, the next state is... The probability of being in the current location is determined; When the user's current action is preference regression, the next state The location is determined based on the following probability distribution: in, Indicates the user's access location Number of times, This represents the total number of times a user visits all locations; When the user's current action is exploration, sort the locations the user has not visited according to their distance from the current location, and determine the next state. The probability distribution of visiting a new location, expressed as: in, For sorting size, the parameter This indicates the user's sensitivity to distance.

3. The method according to claim 1, wherein, The construction of the reward function network includes: Build an instant reward function ; Constructing a long-term reward function ; A multi-scale return function is constructed based on the immediate return function and the long-term return function, wherein the multi-scale return function is expressed as follows: in, To balance the impact of immediate returns and long-term returns, .

4. The method according to claim 3, wherein, The construction of the instant reward function includes: The user's state features and action features are encoded, and the state features and action features are mapped into feature vectors that characterize the state features and action features; The feature vector is input into a linear layer of the ReLU activation function to extract low-dimensional features; The low-dimensional features are processed using the Sigmoid activation function to obtain the instantaneous reward function; Wherein, the instant reward function Represented as: in, Indicates immediate identification results. Indicates the discriminator, This represents the state characteristics corresponding to a short-term action.

5. The method according to claim 3, wherein, The construction of the long-term reward function includes: Obtain the first and second high-level features from the user state features; Based on the mutual information between the first high-level feature and the second high-level feature, fit the first probability distribution of the first high-level feature under the second high-level feature; Calculate the KL divergence of the first probability distribution and the second probability distribution based on the first probability distribution and the second probability distribution of the first high-level feature; The long-term reward function is obtained based on the KL divergence and the discrimination result of the second high-level feature.

6. The method according to claim 5, wherein, The first high-rise feature includes the user's residential address. h The second high-level feature includes the user's historical movement trajectory. ; The long-term reward function is expressed as follows: in, For trajectory identification results, Based on h and The KL divergence after fitting the mutual information between them , is a hyperparameter used to adjust the influence of high-level user characteristics. To generate neural network fitting conditions for adversarial imitation learning in historical movement trajectories Distribution of user characteristics Given a number of time units, To transfer the user's residential address h The probability distribution obtained as input.

7. The method according to claim 1, wherein, The policy function network and reward function network are trained based on the user's historical movement trajectory to obtain a user trajectory generation model, including: The policy function network and reward function network are pre-trained using an optimization objective based on the negative log-likelihood function; The policy function network is trained by maximizing the total reward of the generated trajectory obtained by downsampling according to the policy function network decision, and the reward function network is trained by optimizing the adversarial optimization value of the reward function network.

8. A user trajectory generation device based on generative adversarial imitation learning, comprising: The building module is configured to construct a policy function network and a reward function network based on generative adversarial imitation learning; The training module is configured to train the policy function network and reward function network based on the user's historical movement trajectory to obtain a user trajectory generation model; The generation module is configured to input user status information into the user trajectory generation model to generate the user's movement trajectory. The status information includes at least a first user attribute status and a second user attribute status. The first user attribute status includes the user's location, and the second user attribute status includes user-related feature statuses other than the user's location. The building module is further configured to: before building the policy function network and reward function network based on generative adversarial imitation learning, construct a user mobile behavior model based on Markov decision processes, specifically including: Determine the user's state space, wherein the state space includes a time state, a first user attribute state, and a second user attribute state; Determine the user's action space, wherein the action space includes dwelling, residence return, preference return, and exploration; A user's state transition function is constructed based on the state space and the action space, wherein the user's state transition function represents the user's current state. Perform actions When, reach the next state The probability distribution; The specific configuration of the building module is to build a policy function network, which includes: Based on the user's mobile behavior model, features are extracted from the user's historical mobile trajectory to obtain the user's target state features; The target state features are processed to obtain a dense representation vector; A nonlinear network is extracted from the dense representation vector using a self-attention mechanism; The nonlinear network is normalized using the softmax activation function, and the policy result is output, wherein the policy result includes the probability distribution of each action in the corresponding state.