Method for generating airport terminal passenger boarding decision process in heterogeneous scene
Through a stratified reinforcement learning method, modeling the passenger boarding decision process in the terminal environment, solving the problem of predicting the passenger boarding decision process in the heterogeneous scenario, and achieving accurate prediction of traffic and effective prediction of congestion.
Patent Information
- Application Number
- CN202510327140.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
It is difficult for the prior art to accurately predict the decision process of passenger boarding in terminals in heterogeneous scenarios, especially in the case of changes in terminal facilities and flight schedule changes. Traditional methods such as time series prediction, regression prediction and simulation prediction have problems such as failure.
The hierarchical reinforcement learning method is used to model the passenger boarding decision process in the terminal environment. By evaluating the difference between the passenger flow in each area of the terminal and the actual area of the passenger flow, training the training data generation model of the return function, generating passenger trajectory data, and updating the training data generation model until the difference reaches the minimum value.
Accurate modeling of the passenger boarding decision process of the terminal in heterogeneous scenarios is achieved. The simulated traffic of people in various areas of the terminal is approaching the real situation, and it can effectively predict congestion under the situation of changing the scene.
Smart Images

Figure CN120197496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aviation information technology, and particularly to a method for generating a decision-making process for passengers to board a plane in a terminal building under heterogeneous scenarios. Background Art
[0002] Airports have a huge passenger throughput, and corresponding to the huge flow of people is a huge service pressure. In all aspects of airport services, such as security, security inspection, emergency response to emergencies, check-in, baggage tracking, etc., airport services hope to be able to predict the future passenger throughput and accordingly allocate human and material resources in advance to better serve passengers. However, due to changes in terminal building facilities, flight schedules, etc., the changes in the macroscopic scenarios essentially affect the changes in the microscopic behaviors of passengers. For example, the reduction of check-in counters or security inspection channels essentially affects when different passengers check in or undergo security inspection, and which check-in counter or security inspection channel to choose. Such microscopic decision-making behaviors lead to the following challenges in predicting the future passenger throughput in the terminal building under different scenarios:
[0003] Challenge 1: It is impossible to directly predict macroscopic indicators based on supervised learning through macroscopic indicators. Traditional airport passenger flow prediction directly uses time series prediction or regression prediction. For example, the passenger flow in the next time slice is predicted based on the passenger flow in the previous few time slices, or the passenger flow in the next time period is regressed based on the weather and airport passenger flow distribution in the previous time period.
[0004] Considering the following prediction requirements, in order to reduce airport management costs, several check-in counters and security inspection channels are closed, and the spatio-temporal passenger flow distribution in the terminal building after the change is predicted, and at the same time, the queue lengths in the check-in area and the security inspection area are predicted. Another situation is that the terminal building facilities remain unchanged, but the flight schedule has changed, especially during holidays or emergencies, the flight schedule will change significantly. In the above scenarios, due to the lack of training data after the scenario change, out-of-distribution generalization is required, which will make time series prediction or regression prediction ineffective.
[0005] Challenge 2: It is impossible to predict macroscopic indicators through microscopic simulation based on traditional simulation. To complete the prediction task under the condition of scenario change, simulation prediction can be used. First, set up the simulation environment of the terminal building and the behavior rules of passengers, and finally generate passengers based on flight schedule information. The macroscopic traffic flow distribution and queue data are obtained through the microscopic simulation of passenger behavior. The problem with this method is that it is impossible to determine whether the simulation results are correct only by designing the simulation process of passengers based on simple rules, and different simulation settings may obtain different simulation results.
[0006] In the traditional simulation process, the credibility of simulation results can be improved through simulation calibration. The main calibration is to correct the microscopic simulation parameters such as the behavior of passengers based on methods such as genetic algorithms. However, what is usually calibrated in practice are the spatio-temporal flows in each area of the terminal building, the queue lengths of each check-in counter, the queue lengths of each security checkpoint, and other macroscopic indicators. This process of calibrating macroscopic indicators may show good performance in deterministic, simple, low-dimensional single or group systems. However, when the cluster scale expands, especially when the cluster exhibits high-dimensional, complex, and strongly uncertain behavioral characteristics, the existing models or rule-based empirical knowledge are difficult to cover the entire solution space, and the applicability, stability, and robustness of traditional simulation calibration methods will be greatly reduced. Moreover, macroscopic parameters may also change when the scenario changes.
[0007] Challenge 3: It is difficult to directly simulate the large-scale boarding decisions of terminal passengers based on reinforcement learning.
[0008] There are two problems with this method. First, the goal of reinforcement learning is to maximize the reward for each individual or the whole, but we do not know the actual reward obtained by each individual in a specific scenario, and thus cannot design the corresponding reward function. And the reward function is an essential item in reinforcement learning, which leads to the inability to implement the entire learning process. Although the method of inverse reinforcement learning can be used to infer the reward function, in the problem of large-scale passenger decision-making in the terminal building, it is impossible to obtain the trajectory data of each passenger, so inverse reinforcement learning cannot be used to obtain the reward function. Second, when the terminal building contains a large number of passengers, the number of agents is very large, and the terminal building contains various behavioral decision-making environments, such as check-in, security check, shopping, dining, etc. Existing multi-agent reinforcement learning algorithms are difficult to achieve the collaborative learning of a large number of agents in the complex environment of the terminal building. Summary of the Invention
[0009] Aiming at the above deficiencies in the prior art, a method for generating the boarding decision-making process of terminal passengers in a heterogeneous scenario provided by the present invention solves the following problems: (1) In the original road network scenario, solve the behavioral strategy of the passenger boarding process, and the passenger makes a boarding process decision based on the strategy, so that the passenger flow in each area of the terminal building obtained by simulation approaches the real situation, that is, minimize ; (2) When the scenario of the terminal building changes, complete the passenger check-in process in the changed scenario based on the obtained behavioral strategy of the passenger boarding process, and obtain the congestion situation under the simulation.
[0010] To achieve the above objectives, the technical solution adopted by the present invention is: A method for generating the boarding decision-making process of terminal passengers in a heterogeneous scenario, including the following steps:
[0011] S1. Adopt hierarchical reinforcement learning to model the boarding decision-making process of passengers in the terminal environment;
[0012] S2. Based on the modeling results, evaluate the difference between the passenger flow in each area of the terminal building obtained through training and the actual passenger flow Q* of the terminal building area, and generate a model based on the training data of the reward function trained based on the difference;
[0013] S3. Generate training data based on the model generated from the training data, and update the training data generation model based on the generated training data, where the passenger trajectory is used as the training data;
[0014] S4. Determine whether the difference reaches the minimum value. If so, generate the behavior decision of the passenger boarding process in the heterogeneous scenario in the terminal building; otherwise, return to S1.
[0015] The beneficial effects of the present invention are as follows: The present invention proposes a hierarchical multi-agent reinforcement learning framework based on reward generation to solve the following problems: (1) In the original road network scenario, solve the behavior strategy of the passenger boarding process. Passengers make boarding process decisions based on the strategy, so that the passenger flow in each area of the terminal building obtained by simulation approaches the real situation, that is, minimize ; (2) When the terminal building scenario changes, complete the passenger check-in process in the changed scenario based on the obtained behavior strategy of the passenger boarding process, and obtain the congestion situation in the simulation.
[0016] Further, the specific content of S1 is as follows:
[0017] Perform task hierarchical division on the entire process of airport passenger boarding. Among them, the upper-level decision-making process is the task layer where passengers select actions, the specific decision actions are the action layers under the current tasks, the decision-making behaviors of the task layer are the non-leaf nodes of the decision tree, and the decision-making behaviors of the action layer are the leaf nodes of the decision tree;
[0018] Divide the passenger states into task layer states and action layer states. Among them, the task layer states are at the non-leaf nodes of the decision tree, and the action layer states are at the leaf nodes of the decision tree; when the passenger is in the check-in sub-task layer state, it includes two decisions: the decision to enter the check-in state and the decision to enter the service sub-process state of the check-in hall; when the passenger enters the action layer state, it is necessary to obtain the actual state in the current scenario and perform state transitions between different actual states according to the boarding process.
[0019] Under the hierarchical design, set the passenger actions.
[0020] Set the task layer node rewards and the original action layer node rewards respectively to complete the modeling of the passenger boarding decision-making process in the terminal building environment.
[0021] The beneficial effects of the above further solution are as follows: For passengers inside the terminal building, their behavior decision-making process is complex. Using traditional reinforcement learning algorithms may lead to the problem of dimensional space explosion. Therefore, hierarchical reinforcement learning is adopted to hierarchically divide the tasks in the whole process of airport passengers boarding the plane, converting complex problems into several relatively simple sub-problems.
[0022] Furthermore, the states of the passengers include:
[0023] Under the check-in sub-process, the time freedom freetime and the number of people at the check-in counter checkin together constitute the passenger state set;
[0024] Under the security check sub-process, the state determination of the passengers is determined by the time freedom freetime and the number of people in the security check area security;
[0025] Under the boarding sub-process, the state determination of the passengers is determined by freetime and the number of people in the passenger waiting area wait.
[0026] Furthermore, the expression of the reward for the task layer node is as follows:
[0027]
[0028] Where, represents the reward for the task layer node, represents the estimated value of the obtained benefit when taking action a in state s m of, represents executing sub-task a from state s m-1 to reach the lower-level sub-task a m and obtaining the return, represents the return obtained from executing sub-task a1 from state s to reach the lower-level sub-task a2, represents the return obtained from executing sub-task a0 from state s to reach the lower-level sub-task a1, represents the path node, represents the traversal path obtained through recursion under the current task, represents the state;
[0029] The expression of the reward for the original action layer node is as follows:
[0030]
[0031] Where, represents the reward for the original action layer node, represents the return obtained by the passenger when taking action a in state s.
[0032] The beneficial effects of the above further solution are as follows: The rewards of the upper-level subtasks are all obtained from the rewards of the lower-level primitive actions. Only by calculating the rewards of the primitive actions can the rewards of all subtasks be obtained.
[0033] Furthermore, the specific steps of S2 are as follows:
[0034] S201. Initialize the diffusion model f;
[0035] S202. Randomly generate an initial training set for the training data generation model , where the initial training set is a set of decision-making data during the boarding process of passengers, represents the state, represents the subtask, represents the reward obtained by executing the subtask in the state ;
[0036] S203. Based on the initial training set, generate a training data set , where represents the nth group of data Data;
[0037] S204. Using the method of random sampling without replacement, divide the training data set D into n groups and fit n groups of reward functions , where any training data is used for supervised learning to obtain any reward function , and the input of any reward function is the state behavior data of passengers, and the output of any reward function is the reward obtained by passengers, represents the nth group of reward functions;
[0038] S205. Respectively based on the n groups of reward functions obtained by fitting, perform hierarchical reinforcement learning for airport passengers to obtain the passenger flow Q in each area of the terminal building, and based on the passenger flow Q in each area of the terminal building, obtain several groups of differences between the passenger flow in each area of the terminal building and the actual passenger flow Q* of the terminal building area , where represents the simulated passenger flow in the terminal building area, represents the actual passenger flow in the terminal building area;
[0039] S206. When , a training data generation model based on the difference training reward function is obtained; otherwise, go to S207, where represents the nth group of differences, represents a positive number;
[0040] S207. Construct a data set , where , ;
[0041] S208. Use the data set to train the diffusion model f;
[0042] S209. Randomly generate a group of , where represents the abbreviation of , represents the last element in represents any element in represents a small positive number;
[0043] S2010. Based on the trained diffusion model f and , generate a group of , where represents the generated trajectory data, represents the training data generated by the diffusion model f under the condition below, represents the state value in each trajectory, represents the behavior in each trajectory, the return value in each trajectory represents;
[0044] S2011. Based on and construct the data set , and add the data set to the training data set D to complete the training of the training data generation model.
[0045] The beneficial effect of the above further solution is that the training data is usually called expert data and needs to be sampled in the actual environment or simulation environment, while the present invention can directly generate it only through the trained diffusion model f.
[0046] Furthermore, the use of the data set to train the diffusion model f is specifically as follows:
[0047] Based on the data set , sample a real data sample , where i represents the i-th data in the training data set , and n represents the total number of n groups of data in the training data set ;
[0048] Based on the real data sample , add noise at time step t to obtain a noisy data sample ;
[0049] Input the noisy data samples and the conditions into the neural network model ϵ θ to predict the noise;
[0050] Based on the noise predicted by the neural network model ϵ θ and the actually added noise, calculate the loss function and update the parameters of the diffusion model to complete the training of the diffusion model f.
[0051] The beneficial effect of the above further solution is that the trained diffusion model f can automatically generate new data with the same characteristics as the training data.
[0052] Furthermore, the expression of the loss function is as follows:
[0053]
[0054] where represents the loss function, represents the actual noise added to the original data sample in, represents the noise predicted by the diffusion model, represents the condition of the conditional diffusion model.
[0055] Furthermore, the generation process of each data in the generated trajectory data is as follows:
[0056] Sample noise samples from the standard normal distribution ;
[0057] Starting from time step T, gradually denoise until the time step is 0 to obtain the generated data sample. Description of the Drawings
[0058] Figure 1 is a schematic diagram of the training framework.
[0059] Figure 2 is a flowchart of the method of the present invention.
[0060] Figure 3 is a schematic diagram of the hierarchical decision-making process.
[0061] Figure 4 is a schematic diagram of the sub-task reward.
[0062] Figure 5 is a schematic diagram of the setting of the reward function.
[0063] Figure 6 is a schematic diagram of the passenger flow situation in the terminal building obtained by each model.
[0064] Figure 7 It is a diagram showing the passenger flow in each area of the terminal building obtained from the benchmark model M and the present invention in Scenario 1.
[0065] Figure 8 It is a diagram showing the passenger flow in each area of the terminal building obtained from the benchmark model M and the present invention in Scenario 2. Detailed implementation manners
[0066] The following describes the detailed implementation manners of the present invention to facilitate the understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0067] Embodiment
[0068] The present invention proposes a hierarchical multi-agent reinforcement learning framework based on return generation. First, the problems to be solved and the solution framework are introduced below, and then how this framework addresses the three challenges mentioned in the background art is described.
[0069] Problems to be solved:
[0070] Definition 1: True passenger flow in different areas of the terminal building , , , where represents the true passenger flow in the i-th area of the terminal building at time t, T represents the total number of time steps, i represents the i-th area of the terminal building, n represents the total number of areas, represents the true passenger flow in the n-th area, represents the true passenger flow in the i-th area of the terminal building at time T:
[0071]
[0072] Definition 2: Simulated passenger flow Q in different areas of the terminal building , :
[0073]
[0074] The problems that the present invention ultimately needs to solve include two parts: (1) In the original road network scenario, solve the behavior strategy for the passenger boarding process. Passengers make boarding process decisions based on the strategy to make the passenger flow in each area of the terminal building obtained by simulation approach the real situation, that is, minimize ; (2) When the terminal building scenario changes, based on the obtained behavior strategy for the passenger boarding process Complete the passenger registration process in the changing scenario, obtain the congestion situation in the simulation scenario, and attempt to verify its effectiveness.
[0075] To solve the above two problems, it is necessary to solve the behavioral strategy of the passenger boarding process , as described in Problem 1 below.
[0076] Problem 1: Passenger boarding process strategy problem , which aims to learn the passenger boarding strategy by maximizing the reward obtained during the boarding process , indicates that the passenger completes pathfinding based on this strategy, making the pedestrian flow situation in the final simulation area approximate the real pedestrian flow, represents the local state space observed by the passenger, represents the passenger action space, that is, terminal activities such as check-in, security check, and catering service, represents the reward obtained by the passenger when taking a certain action in a specific state.
[0077] Since the reward function of the passenger boarding process decision model in Problem 1 is unknown, the second problem to be solved by the present invention is to obtain . However, the existing learning method for the reward function based on inverse reinforcement learning requires passenger trajectories as training data, which cannot be obtained in the scenario of the present invention. Therefore, the problem of solving is transformed into the problem of how to generate training data. Once the correct training data is available, can be obtained.
[0078] Problem 2: Reward function of passenger decision-making solution. Let the reward function be a neural network, denoted as , the input of is the state s and the action a, and the output is the reward r obtained by the passenger in the case of (s,a). Since data training is required , so Problem 2 is actually transformed into the generation of training data for . Let the data generator be G, the input of the generator G is m groups of state and action data {(s1,a1),(s2,a),...,(s m ,a m )}, and the output of the generator G is m groups of reward data {r1,r2,...,r m}, where, represents the state, represents the subtask, represents executing the subtask in the state The obtained rewards can be used to construct a set of decision-making data \(\{(s_1, a_1, r_1), (s_2, a_2, r_2), \cdots, (s m , a m , r m )\}\) through the input and output of the generator \(G\). Let \(s i represent the current state of the passenger, \(a i represent the action taken by the passenger, and \(r i represent the reward obtained by the passenger when taking the action \(a i in the state \(s i . It can be obtained through supervised learning based on the decision-making data.
[0079] Solution framework:
[0080] Problems 1 and 2 require interactive training of the passenger boarding process decision-making model \(M\) of the passenger boarding process decision-making model \(M\), , such that the passenger flow in the terminal area obtained based on the passenger boarding process decision-making model is consistent with the actual situation in the real scenario, as shown in formula (1).
[0081] \(M, \Pi = Argmin(\Delta Q)
[0082]
[0083] where \(Argmin()\) represents finding \(M, \Pi\) to minimize \(\Delta Q\), and \(\Delta Q\) represents the difference between the passenger flow \(Q\) in each area of the terminal and the actual passenger flow \(Q^*\) in the terminal area, represents the \(F\)-norm of the matrix.
[0084] In summary, the present invention proposes a passenger boarding process decision-making model \(M\) and an interactive learning training framework, as Figure 1 shown in (b) of: (1) First, fix the reward function , and train the boarding decision-making model \(M\) of the passenger in the terminal environment; (2) Secondly, evaluate the difference \(\Delta Q = |Q - Q^*|\) between the passenger flow \(Q\) in each area of the trained terminal and the actual passenger flow \(Q^*\) in the terminal area; (3) Based on \(\Delta Q\), train the diffusion model of the training data of the reward function, and generate training data \(\{(s_1, a_1, r_1), (s_2, a_2, r_2), \cdots, (s m , a m , r m )\}\) based on the diffusion model , and finally update the reward model based on the training data; (4) Iterate steps (1) to (3) to minimize \(\Delta Q\). Figure 1The learning goal in (a) is to approximate the expert trajectory, and the expert trajectory data is supervised to obtain the reward function , and then perform reinforcement learning to obtain the pedestrian agent M, Figure 1 The learning goal of (b) is to minimize ΔQ, generate pedestrian trajectory data G for supervised learning, and obtain the reward function agent , and then perform reinforcement learning to obtain the pedestrian agent M.
[0085] The following is a brief introduction to how this framework overcomes the three challenges in 2:
[0086] In response to Challenge 1, the framework of the present invention is based on multi-agent reinforcement learning. After learning the behavioral strategy of the agent, simulation decisions can be made under different terminal environments and different flight schedule settings to obtain specific macro-simulation prediction indicators.
[0087] In response to Challenge 2, the traditional simulation simply sets passenger behavior rules, which makes it impossible to generate complex passenger decision-making behaviors. The framework of the present invention generates passenger boarding process behavior strategies based on multi-agent reinforcement learning. More importantly, in order to ensure that passenger behavior decisions are consistent with the actual scenario, the distribution of passenger flow in the real terminal is used as a fitting indicator for multi-agent reinforcement learning, and the learning of the agent decision-making process is guided based on the gap between this indicator and the corresponding indicator in the simulation environment, so as to achieve reverse guidance of micro-strategy learning through the alignment of macro indicators, and open up a closed-loop learning loop from micro to macro and then feedback to micro.
[0088] For challenge 3, the framework of the present invention can learn the reward function while learning the passenger behavior strategy, thus overcoming the problem that the reward function must be known before reinforcement learning can be performed. For problem 1, hierarchical reinforcement learning is used to model the passenger's boarding decision process, and multiple indicators that need to be calibrated are calibrated in layers, realizing the accuracy of the entire simulation process from coarse granularity to fine granularity. For problem 2, generative artificial intelligence technology is used to first generate training data for reward function learning, and then the reward function is trained based on the generated data set, realizing the learning of the reward function without expert data.
[0089] like Figure 2 As shown, the present invention provides a method for generating a passenger boarding decision process in a terminal building under a heterogeneous scenario, and the implementation method thereof is as follows:
[0090] S1. Hierarchical reinforcement learning is used to model the boarding decision process of passengers in the terminal environment, which is as follows:
[0091] The whole process of airport passenger boarding is divided into task layers. Among them, the upper-level decision-making process is the task layer for passengers to choose actions, the specific decision-making actions are the action layers under the current tasks, the decision-making behaviors of the task layer are the non-leaf nodes of the decision tree, and the decision-making behaviors of the action layer are the leaf nodes of the decision tree;
[0092] The passenger states are divided into task-layer states and action-layer states. Among them, the task-layer states are at the non-leaf nodes of the decision tree, and the action-layer states are at the leaf nodes of the decision tree; when the passenger is in the check-in sub-task layer state, it includes two decisions: the decision to enter the check-in state and the decision to enter the service sub-process state of the check-in hall; when the passenger enters the action-layer state, it is necessary to obtain the actual state in the current scenario and perform state transitions between different actual states according to the boarding process;
[0093] Under the hierarchical design, the passenger actions are set;
[0094] The rewards for the task-layer nodes and the original action-layer nodes are set respectively to complete the modeling of the passenger's boarding decision-making process in the terminal environment.
[0095] In this embodiment, the states of the passengers include:
[0096] Under the check-in sub-process, the time freedom freetime and the number of people at the check-in counter checkin together constitute the passenger state set;
[0097] Under the security check sub-process, the state determination of the passenger is determined by the time freedom freetime and the number of people in the security check area security;
[0098] Under the boarding sub-process, the state determination of the passenger is determined by freetime and the number of people in the passenger waiting area wait.
[0099] In this embodiment, for the passengers inside the terminal, their behavior decision-making process is complex, and using traditional reinforcement learning algorithms may cause the problem of dimensional space explosion. Therefore, hierarchical reinforcement learning is adopted to divide the whole process of airport passenger boarding into task layers, converting complex problems into several relatively simple sub-problems. The following explains how to obtain the behavior decision-making model of a large number of passengers during the boarding process in the terminal based on hierarchical reinforcement learning:
[0100] As Figure 3 The upper part of (a) is the task layer for passengers to choose actions, that is, the upper-level decision-making process; Figure 3 The lower part of (a) is the action layer under the current task, that is, the specific decision-making actions; Figure 3 (b) shows the topological relationship between the upper and lower layers, which is represented by a tree-shaped decision structure. The decision-making behaviors of the task layer are the non-leaf nodes of the decision tree, that is, Figure 3The red nodes in (b), such as the value machine sub-task, the security check sub-task, etc. The decision-making behavior of the action layer is the leaf node of the decision tree, that is Figure 3 The blue nodes in (b) include specific decision-making actions such as check-in, security check, shopping, dining, etc.
[0101] The above hierarchical reinforcement learning problem can be described as the following semi-Markov process. The semi-Markov decision process (SMDP) can be regarded as an extension of the Markov process and can be defined as , S represents the state passed through during the passenger boarding process, represents the set of actions for taking various decisions during the passenger boarding process, represents the probability transition function, R represents the reward obtained after executing action a in S, represents the probability that the passenger transfers to after N steps of executing action a in state s, represents the expected reward obtained by the agent for choosing action a in state s; the optimization goal of the agent is to maximize the reward during the action duration and obtain the optimal policy. Based on the semi-Markov decision process SMDP, the optimal Bellman equation of the value function is:
[0102]
[0103]
[0104] Among them, represents the estimate of the value in state s, represents the N-fold discount of the reward, represents that after executing action a in state s, it transfers to state , represents state value, represents the value obtained by taking action 0 in state s, represents the value obtained by taking action a in state s m obtained, Execute sub-task a from state s m-1 arrive at the lower-level sub-task a m obtained reward, represents the reward obtained by executing sub-task a1 from state s to reach the lower-level sub-task a2, represents the reward obtained by executing sub-task a0 from state s to reach the lower-level sub-task a1.
[0105] In this embodiment, the passenger state can be divided into a task layer state and an action layer state, as Figure 3 shown in (b).
[0106] First, the task layer state is at the non-leaf nodes of the decision tree. The red non-leaf nodes are the task layer states. When a passenger is in the check-in sub-task layer state, they can make two decisions. One is to go directly to check in and enter the check-in state, and the other is to enter the service sub-process state of the check-in hall. The transfer relationship between these states is fixed based on the passenger boarding process in the terminal building.
[0107] Secondly, the action layer state is at the leaf nodes of the decision tree. The blue non-leaf nodes are the action layer states. When a passenger makes a decision and enters the action layer, they need to obtain the actual state in the current scenario and perform state transfer between different actual states according to the boarding process. The specific states of passengers under different tasks are as follows:
[0108] (1) The specific states under the check-in sub-task = (freetime, checkin). Under the check-in sub-process, the time freedom freetime and the number of people at the check-in counter checkin together constitute the passenger state set.
[0109] The time freedom is defined as the difference between the current time and the passenger's flight departure time obtained through discretization. In the terminal building, the higher the passenger's time freedom, the longer the time they can freely dispose of. Conversely, they need to quickly complete all necessary processes.
[0110] The number of people at the check-in counter is a secondary factor affecting whether a passenger chooses to check in. Similarly, due to the large and uncertain numerical value of the number of people in the area, discretization processing is required during the training process.
[0111] (2) The specific states under the security check sub-task = (freetime, security). Under the security check sub-process, the state determination is determined by the time freedom freetime and the number of people in the security check area security. The number of people in the security check area refers to the number of people waiting or being inspected under the security check service.
[0112] (3) The specific states under the boarding sub-task = (freetime, wait). Under the boarding sub-process, the state determination is determined by freetime and the number of people in the passenger waiting area wait. The number of people in the passenger waiting area refers to the number of people in the area near the passenger's target boarding gate where passengers can sit and rest.
[0113] In this embodiment, as Figure 3 shown, under the hierarchical design, the passenger's action setting is affected by different stages, and the available action sets are different in different stages. The upper-layer state is a sub-task. The action setting of the upper-layer task can be regarded as the facilities or necessary process facilities that a passenger can choose to go to under a certain sub-task, which is an abstract action expression. The action setting of the lower-layer task includes actions such as check-in, dining, shopping, restroom, security check, and waiting for boarding. The execution of the original actions will directly act on the environment.Figure 3 Description (b) shows the different actions taken by the upper and lower layer tasks.
[0114] In this embodiment, the reward is designed as follows.
[0115] (1) Task layer node reward
[0116] Since there may be other tasks included under the task layer, the task reward is not only the reward value obtained by the passenger for selecting the current task, but also the cumulative reward obtained by the agent for selecting the lower layer tasks layer by layer until reaching the original action. Therefore, the reward obtained by the agent for completing the corresponding task is the sum of the rewards of all the subordinate tasks experienced by the agent for completing the task. It is necessary to perform recursive calculation on the expected return function That is, expand the possible tasks with the current subtask as the root node, and list all possible completion functions . If there are still tasks under the current subtask, continue the search and make comparisons. Finally, the task with the larger return value will be selected.
[0117] As shown in the following formula, assume that the traversal path obtained through recursion under the current task is {a0, a1,..., a m}}, then the last execution node a m will return and calculate the completion function . And so on, then the total reward obtained by this task is:
[0118]
[0119] Among them, represents the task layer node reward, represents the estimate of the benefit value obtained by taking action a m in state s, represents the return obtained by executing subtask a m-1 from state s to reach the lower layer subtask a m , represents the return obtained by executing subtask a1 from state s to reach the lower layer subtask a2, represents the return obtained by executing subtask a0 from state s to reach the lower layer subtask a1, represents the path node, represents the traversal path obtained through recursion under the current task, represents the state.
[0120] In this embodiment, take Figure 3Taking the reward obtained by the first subtask "check-in sub-task" as an example, calculate reward(check-in sub-task, s). Since the check-in sub-task includes service sub-tasks and the original action "check-in", and the service sub-task includes original actions "shopping", "catering", etc., the reward(check-in sub-task, s) needs to be calculated recursively. Here, it is assumed that in the service sub-task, the reward for choosing "shopping" is the largest. The following gives the specific recursive calculation process. Figure 4 The graphical description of the recursive process is given:
[0121]
[0122] To sum up, the node reward of the task layer is obtained by recursively adding the reward of the bottom layer actions. Therefore, as long as the reward of the original action layer can be obtained, the reinforcement learning training for airport passenger boarding decision-making can be completed.
[0123] In this embodiment, regarding the node reward of the original action layer, the present invention uses the reward function R to be solved in Problem 2 M to calculate the rewards in different states and behaviors, that is , where s and a respectively represent the state of the passenger and the action taken, s represents the current state of the passenger, represents the reward obtained when the passenger takes action a in state s.
[0124] In this embodiment, under the hierarchical condition, the passenger agent needs to evaluate all executable actions under the current subtask and select the optimal action in the current state. After selecting the action, the passenger agent executes the original action in the environment and obtains a reward, and repeats such a process until the overall process is completed. In the present invention, the purpose of agent optimization is to make the cumulative reward reach the maximum value at each stage. According to the formula is defined to make the final reward value maximized.
[0125] In this embodiment, the algorithm for predicting the trajectory of a single airport passenger based on hierarchical reinforcement learning is as follows:
[0126] First, initialize the airport environment and input the root task into the algorithm; initialize the queue information and obtain the observation value; determine whether the training is completed. If not, continue to execute downward; determine whether the training reaches the original action. If it is the original action, obtain the current state s t , execute action a, and get s t+1 , obtain the reward r for this time, and through ( represents the state value estimation at time t + 1, (denoting the estimated state value at time t) execute the action and obtain the reward. Otherwise, according to the current subtask M i and the exploration strategy, select the action a*; maximize the action a* , execute the action a* to obtain the subsequent subtask M j ; Let the subsequent subtask sequence of subtask M j be childseq: Update each state s in the subtask sequence childseq: , add the subtask sequence childseq to the front of the seq sequence (childseq represents the subtask sequence, and seq represents the task sequence), that is, continuously obtain the subsequent subtask sequence according to the exploration strategy; calculate the completion function for each state, and obtain the optimal action, then update all the completion functions involved. After completion, add the subsequent subtask sequence to the front of the action sequence. Loop the above process until the training task is completed, where, denotes the state execute the action to reach the subtask M i the reward obtained, denotes the estimated benefit value when the state s executes the action a* to reach the subtask M i , denotes the estimated reward value, denotes the estimated state value, denotes the reward obtained when the action a is executed in the state s to reach the subtask M i , denotes execute the action a* to reach the subtask M i the reward obtained, denotes the state execute the action a* to obtain the reward, denotes the reward obtained when the state s executes the action a to reach the subtask M j .
[0127] S2. Based on the modeling results, evaluate the difference between the passenger flow in each area of the terminal building obtained by training and the actual passenger flow Q* in the terminal building area, and generate a model based on the training data of the reward function for the difference;
[0128] In this embodiment, when the reward function is unknown, inverse reinforcement learning is usually used to learn the reward function. However, since inverse reinforcement learning requires expert trajectory data, which cannot be obtained in the scenario of the present invention. To overcome this problem, based on the generative conditional diffusion model, first generate passenger decision-making data that can train the reward function, and then learn the reward function based on this data. The implementation method is as follows:
[0129] S201. Initialize the diffusion model f, which is a neural network with a U-net architecture in the present invention;
[0130] S202. Randomly generate an initial training set for the training data generation model , where the initial training set is a set of decision-making data during the boarding process of passengers, represents the state, represents the subtask, represents executing the subtask in the state and obtaining the reward;
[0131] S203. Based on the initial training set, generate a training data set , where, represents the nth group of data Data, and any is randomly generated in S202;
[0132] S204. Using the method of random sampling without replacement, divide the training data set D into n groups and fit n groups of reward functions , where, using any training data to perform supervised learning to obtain any reward function , and any reward function takes the state behavior data of passengers as input, and any reward function outputs the reward obtained by the passengers, represents the nth group of reward functions;
[0133] S205. Respectively based on the n groups of reward functions obtained by fitting, perform hierarchical reinforcement learning for airport passengers to obtain the passenger flow Q in each area of the terminal building, and based on the passenger flow Q in each area of the terminal building, obtain several groups of differences between the passenger flow in each area of the terminal building and the actual area passenger flow Q* of the terminal building , where, represents the simulated passenger flow in the terminal building area, represents the actual passenger flow in the terminal building area;
[0134] S206. When , then obtain a training data generation model based on the difference training reward function, otherwise, enter S207, where, represents the nth group of differences, represents a positive number;
[0135] S207. Construct a data set , where, , ;
[0136] S208. Use the data set Train the diffusion model f;
[0137] S209. Randomly generate a set of , where represents the abbreviation of , represents the last element in represents any element in represents a small positive number;
[0138] S2010. Based on the trained diffusion model f and , generate a set of , where represents the generated trajectory data, represents the training data generated by the diffusion model f under the condition , represents the state value in each trajectory, represents the behavior in each trajectory, the return value in each trajectory is represented;
[0139] S2011. Based on and construct the data set , and add the data set to the training data set D to complete the training of the training data generation model.
[0140] In this embodiment, the diffusion model f is trained using the data set , and specifically:
[0141] Based on the data set , sample a real data sample , where i represents the i-th data in the training data set , and n represents the total number of n groups of data in the training data set ;
[0142] Add noise at time step t to obtain the noisy data sample ;
[0143] Input the noisy data sample and the condition into the neural network model ϵ θ to predict the noise;
[0144] Based on the neural network model ϵ θPredict the noise and the actually added noise, calculate the loss function, and update the parameters of the diffusion model to complete the training of the diffusion model f. Among them, the training objective of the conditional diffusion model f is to minimize the difference between the predicted noise and the actual noise, and the loss function is the mean square error (MSE):
[0145]
[0146] Among them, represents the loss function, represents the actual noise added to the original data sample ; represents the noise predicted by the diffusion model, represents the condition of the conditional diffusion model.
[0147] In this embodiment, the generation process of each data in the generated trajectory data is as follows:
[0148] Sample noise samples from the standard normal distribution ;
[0149] Starting from time step T, gradually denoise until the time step is 0 to obtain the generated data sample.
[0150] S3. Generate training data based on the training data generation model, and update the training data generation model based on the generated training data, where the passenger trajectory is used as the training data;
[0151] S4. Determine whether the difference reaches the minimum value. If so, generate the behavior decision of the passenger boarding process in the heterogeneous scenario; otherwise, return to S1.
[0152] The present invention will be further described below.
[0153] This experiment is based on the Anylogic simulation environment to verify the effectiveness of the algorithm. The experimental process is designed as follows:
[0154] (1) Build the airport environment: The data required to construct the terminal simulation environment includes the terminal CAD map and the service performance parameters of each facility. The service performance parameters of each facility include: the time interval for passengers to pass through this facility; the boarding time; the service time of the check-in counter, and use the space module of the Anylogic simulation software to build the corresponding facilities of the airport. The entire airport is divided into six large areas (Q1 to Q6), and several small areas (Q7 to Q 21 ) are divided under each large area, Q = {Q1, Q2,..., Q 21}.
[0155] (2) Passenger arrival data construction: The data required for passenger arrival construction includes all flight information within 6 hours, passenger speed, and flight information including the departure time of the target flight, check-in counter number, boarding gate number, and number of passengers on the flight. The number of passengers on the flight is generated by the intelligent agent. The passenger arrival distribution can be approximately regarded as a Poisson distribution.
[0156] (3) Construction of the passenger benchmark boarding decision-making process M: , where freetime and checkin represent the specific states under the check-in subtask (it can also be other subtasks, such as freetime and security), represents the degree of psychological willingness to choose a certain action (1-3, with 3 being the highest degree), w represents the weight. Based on existing research on the behavioral characteristics of airport passengers, the behaviors of most people are regular, and the influence weights of time, queuing, and psychology are approximately around 0.4, 0.3, and 0.7. Therefore, the calculation method of w is: , represents the normal distribution, represents the event The probability of .
[0157] (4) Construction of the terminal change scenario:
[0158] Plan 1: Close the check-in counters in area Q 11 The passengers who originally checked in in area Q 11 go to area Q 12 to check in.
[0159] Plan 2: Close the security inspection equipment in areas Q 15 and Q 16 .
[0160] (5) Benchmark data synthesis: For the original terminal scenario and the terminal scenarios changed by Plans 1 and 2, let passengers board the plane based on the behavioral strategy in (3) to obtain the benchmark data O1, O2, O3, as shown in Table 1. Among them, Q i,j represents the passenger flow in the jth area of the terminal in the ith scenario when passengers make decisions using the benchmark strategy, represents the passenger flow in the jth area of the terminal in the ith scenario when passengers make decisions using the strategy model of the present invention. Table 1 is a data description table.
[0161] Table 1
[0162]
[0163] (6) Comparison model
[0164] Since there is no similar algorithm framework in the existing work to complete the learning of the large-scale passenger boarding decision-making model, the following several comparison algorithms are designed based on the existing algorithms:
[0165] M1: Fixed reward function, directly learn the passenger behavior decision-making model based on the framework of S1. However, since the reward function is unknown, the reward function of the original actions in S1 is set manually.
[0166] In the airport simulation environment, all shopping, toilet, and catering service stores are divided into the check-in hall service area and the waiting hall service area according to the area. Each service facility can also be regarded as a service area within a certain area. One part of the terminal building is obtained by visualizing the density according to the real number of people, and the other part is obtained by visualizing the density in the simulation environment. Whether the densities match can be regarded as the number of people in the real scenario at a certain service facility being within a certain interval, and the number of people in the same interval in the simulation scenario is similar. Then it can be considered that the action selection conforms to the density. The original action reward returned to the agent in the current area can be expressed as:
[0167]
[0168] Among them, represents the area density, z represents the area, represents the number of agents selecting the area in the current time interval t. Taking the inverse is to make the density more matching and the reward value larger. The number of people in the area is divided into several intervals. In Figure 5 , the red dotted line area in the figure is the waiting area in front of the passenger boarding gate. Among them, the red dots represent the density (density > 5) being the highest, the blue dots represent the density being moderate (2.5 < density < 5), and the cyan dots represent the density being small (density < 2.5). If the number of people in the simulation is within the real interval during the simulation process, it means that the number of people in the area obtained by simulation at the current moment is approximately equal to the number of people in the area in the real situation, otherwise they are not equal. Among them, Figure 5 includes the pedestrian flow density map in the real scenario and the pedestrian flow density map in the simulation scenario. Figures 5 to 8 In
[0169] M2: Fix the strategy of the passengers and only learn the reward function.
[0170] M3: Change the MAX-Q maximizing Q-value hierarchical reinforcement learning algorithm in the framework of the present invention to the DQN deep Q-learning algorithm;
[0171] M4: Remove the discriminator that learns the reward function based on the generative adversarial mechanism in the framework of the present invention, and only retain the generator, that is, learn the reward function only by randomly generating samples.
[0172] In this embodiment, the regional pedestrian flow restoration experiments (RQ1, RQ2) under the original scenario.
[0173] To evaluate the difference between the congestion restoration result and the benchmark data, two evaluation indicators, MSE and RMSE, are proposed to be used.
[0174]
[0175] Among them, represents the pedestrian flow of the th area in the benchmark data, represents the pedestrian flow of the th area in the simulation result, represents the total number of areas. As shown in Table 2, Table 2 is the congestion restoration result table of different models on the benchmark data Q1.
[0176] Table 2
[0177]
[0178] In this embodiment, as Figure 6 shown, Figure 6 is the pedestrian flow situation of the terminal building obtained by each model. The ordinate is each area (1 - 21), and the abscissa is time (6000s - 9000s). Figure 6 It can be seen that compared with the benchmark model M, the human congestion situation in each key area of the present invention is generally in line. About 1 / 3 of the places of M1 and M2 do not conform to the benchmark model M, indicating that fixing either the reward function or the passenger strategy has a certain impact on the model. M3 is the same as the present invention and is more in line with the benchmark M, indicating the generality of the model framework, and replacing similar algorithms does not affect the model effect. About 2 / 3 of the places of M4 do not conform to the benchmark model M, indicating that the effect of randomly generating samples is poor.
[0179] In this embodiment, the regional pedestrian flow prediction experiment (RQ3) under the changed scenario. To verify the pedestrian flow prediction effect after the implementation of the scheme, the reward function and passenger strategy obtained in the restoration experiment are used in the implementation scenarios of Scheme 1 and Scheme 2 to complete the regional pedestrian flow prediction experiment. Tables 3 and 4 also use the two evaluation indicators, MSE and RMSE, to evaluate Q2 and 、Q3 and , Table 3 is the congestion prediction result table of different models on the benchmark data O2, and Table 4 is the congestion prediction result table of different models on the benchmark data O3.
[0180] Table 3
[0181]
[0182] Table 4
[0183]
[0184] From Figure 7 From the passenger flow data obtained from the benchmark model M, it can be seen that the change of the scenario, that is, the check-in counter in area 11 is closed, resulting in a sharp increase in the number of people checking in at area 12. Comparing Figure 7 the passenger flow data obtained by the present invention in with M, it can be seen that the present invention effectively predicts the congestion of the passenger flow in area 12 and other areas after the change of the scenario. Similarly, Figure 7 From the passenger flow data obtained from the benchmark model M, it can be seen that in scenario two, closing the security inspection equipment in areas 15 and 16 will cause an increase in the number of people in security inspection areas 13 and 14. The prediction result of the present invention also roughly conforms to the passenger flow data obtained from the benchmark model M.
[0185] In this embodiment, the time complexity analysis of different methods is shown in Table 5, and Table 5 is the training time table of each model on the data set.
[0186] Table 5
[0187]
[0188] In terms of time complexity, the present invention has a certain degree of lead over other models in terms of the time consumption per round and the total time consumption. M3 is close to the present invention, while M1 and M2 are worse, and M4 has the longest time consumption per round.
Claims
1. A method for generating passenger boarding decision process in a terminal under heterogeneous scenarios, characterized in that: The following steps are involved: S1. Hierarchical reinforcement learning is used to model passengers’ boarding decision-making process in a terminal environment. S2. Based on the modeling results, evaluate the difference between the passenger flow in each area of the terminal obtained by training and the passenger flow Q* in the actual area of the terminal, and generate a model based on the training data of the difference training reward function; S3, generating training data based on the training data generation model, and updating the training data generation model based on the generated training data, wherein the passenger trajectory is used as the training data; S4. Determine whether the difference reaches the minimum value. If so, generate the passenger boarding process behavior decision in the heterogeneous scenario. Otherwise, return to S1.
2. The method for generating the passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 1 is characterized in that: The S1 is specifically as follows: The whole process of airport passenger boarding is divided into tasks in layers, where the upper decision-making process is the task layer where passengers choose actions, the specific decision actions are the action layer under the current task, the decision behaviors of the task layer are the non-leaf nodes of the decision tree, and the decision behaviors of the action layer are the leaf nodes of the decision tree; The passenger status is divided into task layer status and action layer status, where the task layer status is at the non-leaf node of the decision tree, and the action layer status is at the leaf node of the decision tree; when the passenger is in the check-in sub-task layer status, it includes two decisions: entering the check-in state decision and entering the check-in hall service sub-process state decision; when the passenger enters the action layer state, it is necessary to obtain the actual state of the current scene and transfer the state between different actual states according to the boarding process; Under the hierarchical design, passenger actions are set; The task layer node rewards and the original action layer node rewards are set separately to complete the modeling of the passenger boarding decision process in the terminal environment.
3. The method for generating a passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 2 is characterized in that: The passenger's status includes: In the check-in sub-process, the time freedom freetime and the number of people at the check-in counter checkin together constitute the passenger state set; In the security check sub-process, the passenger status is determined by the freetime and the number of people in the security check area. In the boarding sub-process, the passenger status is determined by the free time and the number of passengers in the waiting area.
4. The method for generating a passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 2 is characterized in that: The expression of the task layer node reward is as follows: in, Indicates the task layer node reward, Indicates taking action a in state s m An estimate of the value of the benefit obtained, Indicates executing subtask a from state s m-1 Reach the lower level subtask a m The rewards obtained, represents the reward obtained from executing subtask a1 in state s to reach the lower subtask a2, represents the reward obtained from executing subtask a0 from state s to reach the lower subtask a1, represents a path node, Indicates the traversal path obtained by recursion under the current task, Indicates status; The expression of the original action layer node reward is as follows: in, represents the original action layer node reward, It represents the reward obtained by the passenger when taking action a in state s.
5. The method for generating the passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 1 is characterized in that: The S2 is specifically: S201, initializing diffusion model f; S202, randomly generate training data to generate the initial training set of the model , where the initial training set is a set of decision data during the boarding process of passengers. Indicates the status, Represents a subtask, Indicates in status Execute subtasks the rewards obtained; S203: Generate a training data set based on the initial training set ,in, Indicates the nth group of data Data; S204: Divide the training data set D into n groups by random sampling without replacement, and fit n groups of reward functions , where any training data is used Supervised learning for arbitrary reward functions , any reward function The input is the passenger's state behavior data, and any reward function The output is the return received by the passenger. represents the nth group reward function; S205, based on the n groups of reward functions obtained by fitting Perform airport passenger stratification reinforcement learning to obtain the passenger flow Q in each area of the terminal, and based on the passenger flow Q in each area of the terminal, obtain the differences between the passenger flow in each area of several groups of terminals and the actual passenger flow Q* in the terminal area. ,in, Indicates the simulated passenger flow in the terminal area. Indicates the actual flow of people in the terminal area; S206, for , then the training data generation model based on the difference training reward function is obtained, otherwise, enter S207, where represents the difference of the nth group, Indicates a positive number; S207. Build a data set ,in, , ; S208. Using Datasets Train diffusion model f; S209, randomly generate a group ,in, Express Abbreviation of express The last element in express Any element in Represents a small positive number; S2010, based on the trained diffusion model f and , generating a set ,in, represents the generated trajectory data, The diffusion model f is expressed under the condition The training data generated below is represents the state value in each trajectory, represents the behavior in each trajectory, The reward value in each trajectory is represented; S2011, based on and Building a dataset , and the dataset Add it to the training data set D to complete the training of the training data generation model.
6. The method for generating the passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 5 is characterized in that: The utilization data set The training diffusion model f is as follows: Based on the dataset , sample a real data sample , where i represents the training data set The i-th data in, n represents the training data set There are n sets of data in total; Based on real data samples , add noise at time step t to get noisy data samples ; The noisy data samples and conditions Input to the neural network model ϵ θ In the above example, the noise is predicted; Based on the neural network model ϵ θ The predicted noise and the actual added noise are used to calculate the loss function, and the parameters of the diffusion model are updated to complete the training of the diffusion model f.
7. The method for generating a passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 6 is characterized in that: The expression of the loss function is as follows: in, represents the loss function, Indicates adding to the original data sample The real noise in represents the noise predicted by the diffusion model, Represents the conditions of the conditional diffusion model.
8. The method for generating the passenger boarding decision process in a terminal under heterogeneous scenarios according to claim 5 is characterized in that: The generated trajectory data The generation process of each data in is as follows: Sampling noise samples from a standard normal distribution ; Starting from time step T, denoising is performed step by step until the time step is 0 to obtain the generated data sample.