Unmanned apron control method, system and medium based on reinforcement learning
By deploying collaborative cluster equipment on the apron, building a traffic topology model, and using the collaborative graph multi-agent reinforcement learning model to generate an unmanned vehicle trajectory planning strategy, the problems of insufficient rule constraints and conflict avoidance in traditional methods are solved, and the robot-vehicle-field road collaborative environment and efficient unmanned vehicle trajectory planning are realized.
Patent Information
- Application Number
- CN202510279883.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The traditional unmanned vehicle trajectory planning method in the apron does not fully consider the rules and constraints during the operation of the apron, and mainly focuses on the control of unmanned vehicle trajectory, ignoring the problem of conflict avoidance, resulting in large errors in actual applications.
Adopting a UAV control method based on reinforcement learning, a situation data is obtained through collaborative cluster equipment, a traffic topology model is constructed, a machine-vehicle-field road collaborative environment is established, and an unmanned vehicle trajectory planning strategy is generated based on the collaborative map multi-agent reinforcement learning model to achieve conflict avoidance.
It effectively avoided the conflict between the motorcycle and the vehicle, established a road rights control mechanism for the apron service service, realized a coordinated environment of "mobile-vehicle-field road", and improved the service level and safety level of the apron.
Smart Images

Figure CN119781504B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and in particular to an unmanned apron control method, system and medium based on reinforcement learning. Background Art
[0002] In recent years, artificial intelligence, vehicle-road collaboration, and unmanned driving technologies have become increasingly mature, and the aviation field is also comprehensively promoting the demonstration and application of airport unmanned driving technology. The airport apron area is a designated area within the airport for aircraft to board and disembark passengers, load and unload mail or cargo, refuel, park or repair. With the advancement of the apron unmanned flight support mode, the mixed operation of unmanned vehicles and manned aircraft on the apron has brought challenges to operational safety and efficiency. Among them, the aircraft target activities are intensive and the occupation time of field resources is relatively fixed; the total amount of unmanned vehicle support tasks is large and the support time is relatively loose, but there is a tight coupling relationship between each support operation. The vehicle-road collaboration technology in the apron unmanned operation scenario aims to achieve real-time perception and decision-making of traffic situations through the coordination of vehicles and field infrastructure, and improve the service and safety levels of the apron; how to reasonably plan the activity trajectories for many unmanned support vehicles in the context of "machine-vehicle-field" collaboration on the unmanned apron is a problem that needs to be studied.
[0003] Traditional methods for trajectory planning of unmanned vehicles on the apron include operations optimization and simulation optimization. Methods based on operations optimization include: first, the scheduling of a single type of ground support vehicle, which needs to consider the precise docking of a single type of vehicle with multiple parking spaces, and vehicle operations with flight support needs; second, the scheduling of multiple types of ground support vehicles, which is usually aimed at a single flight and solves the scheduling problem of multiple types of support vehicles. The trajectory planning method of unmanned vehicles on the apron, which is centered on simulation optimization, usually focuses on the safety issues of the trajectory planning process, and conducts rule-based discrete system modeling and simulation; conflict avoidance is achieved through effective operation regulations and multiple simulation verifications.
[0004] Looking at the above studies and related research in the field of road traffic, the shortcomings of traditional unmanned vehicle trajectory planning methods on the apron include: (1) Research centered on operations optimization mostly uses pure mathematical models to solve the unmanned vehicle trajectory planning problem; the rules and constraints during the apron operation are not fully considered, and the computational efficiency needs to be improved. (2) In the research on unmanned vehicle trajectory planning centered on simulation optimization, the main work focuses on improving models, conflict identification and classification, and few conflict avoidance mechanisms are proposed. (3) Although the research on unmanned vehicle trajectory planning under vehicle-road collaboration in the field of road traffic has reference value, it focuses on improving traffic signal timing mechanisms or trajectory control algorithms, and the research objects are limited to vehicles; for apron scenes, such models and algorithms are not applicable because they involve the interaction between unmanned vehicles and aircraft. Summary of the invention
[0005] The technical problem to be solved by the present invention is that the traditional unmanned vehicle trajectory planning method in the apron does not fully consider the rules and constraints in the apron operation process, mainly focuses on the unmanned vehicle trajectory control, ignores the conflict avoidance problem, and has large errors in practical applications; the purpose of the present invention is to provide an unmanned apron control method, system and medium based on reinforcement learning, to avoid locomotive-vehicle conflicts as the guide, to establish an apron service vehicle road right management mechanism, and to realize the "locomotive-vehicle-field road" collaborative environment; secondly, based on the collaborative graph multi-agent reinforcement learning model, an unmanned vehicle trajectory planning strategy is generated, which combines the anthropomorphic driving concept and the apron area operation rules to realize the unmanned vehicle trajectory planning for the apron scene; finally, a comprehensive simulation platform combining MATLAB software, VISSIM software, and Prepar3D software is developed to complete the visualization of the vehicle trajectory planning model in the unmanned apron of a large airport.
[0006] The present invention is achieved through the following technical solutions:
[0007] The unmanned apron control method based on reinforcement learning is characterized by comprising:
[0008] Deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron;
[0009] Constructing a traffic topology model of the unmanned apron based on the situation data;
[0010] Based on the traffic topology model, a machine-vehicle-field-track collaborative environment is constructed, and in the machine-vehicle-field-track collaborative environment, a trajectory planning strategy for the unmanned vehicle is generated based on a collaborative graph multi-agent reinforcement learning model;
[0011] According to the unmanned vehicle trajectory planning strategy, the unmanned vehicle in the unmanned apron is controlled through collaborative cluster equipment.
[0012] Working principle of this scheme: The traditional unmanned vehicle trajectory planning method in the apron does not fully consider the rules and constraints in the apron operation process, mainly focuses on the unmanned vehicle trajectory control, ignores the conflict avoidance problem, and has large errors in practical applications; the purpose of the present invention is to provide an unmanned apron control method, system and medium based on reinforcement learning, to avoid locomotive-vehicle conflicts as the guide, establish an apron service vehicle road right management mechanism, and realize the "locomotive-vehicle-field road" collaborative environment; secondly, based on the collaborative graph multi-agent reinforcement learning model, the unmanned vehicle trajectory planning strategy is generated, which combines the anthropomorphic driving concept and the apron area operation rules to realize the unmanned vehicle trajectory planning for the apron scene; finally, a comprehensive simulation platform combining MATLAB software, VISSIM software, and Prepar3D software is developed to complete the visualization of the vehicle trajectory planning model in the unmanned apron of a large airport.
[0013] A further optimization scheme is to deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron; including methods:
[0014] Sensing equipment is deployed on the track to sense the situation data of moving targets within the range;
[0015] Deploy control equipment at the moving target to control the unmanned vehicle to execute the unmanned vehicle trajectory planning strategy;
[0016] A signal conversion module is deployed on the traffic indication target to build a locomotive-vehicle-field-track collaborative environment.
[0017] A further optimization scheme is that the traffic topology model of the unmanned apron is constructed based on the situation data, including the method:
[0018] Construct service lanes, virtual channels, control nodes and physical nodes of the unmanned apron, and form a directed connected graph consisting of control nodes, physical nodes and the lines between control nodes and physical nodes;
[0019] The control node is used to manage activities between similar targets; the control node is arranged at the intersection of the peripheral service lane and the taxiway;
[0020] The physical nodes are used to manage activities between heterogeneous targets; the physical nodes are arranged at the turning position of the unmanned apron service lane and the intersection of the service lane and the virtual channel;
[0021] The virtual channel is used for aircraft to pass between parking stands; the virtual channel is set on the center line between the internal taxiways of two adjacent standard parking stands, and the width of the virtual channel is ≥ the maximum width of the unmanned vehicle in the unmanned apron.
[0022] A further optimization scheme is that the method further comprises:
[0023] On MATLAB software, a locomotive-vehicle-field-track collaborative environment is established in combination with a directed connectivity graph;
[0024] Establish an unmanned apron micro-model on VISSIM software;
[0025] Create a 3D model of the unmanned apron using the Prepar3D software.
[0026] A further optimization scheme is that the locomotive-vehicle-field-road collaborative environment is constructed based on the traffic topology model, including the following method:
[0027] Get the aircraft's off-block or on-block time t n , the time it takes for the aircraft to slide from the physical node on the taxiway to the parking position is t i, the time it takes for an aircraft to push out or taxi out of the parking position to the physical node of the apron taxiway t0;
[0028] For inbound aircraft, the parking position φ(f n ) The adjacent control node corresponds to the right-of-way signal state switching time window set R app is {[φ(f n )],[t n -t i , t n ]}; where the right-of-way signal state is (t n -t i ) is switched to the “red light” state, and the right-of-way signal state is at t n Always switch to the "green light" state;
[0029] For departing aircraft, the parking position φ(f n ) The adjacent control node corresponds to the right-of-way signal state switching time window R dep is {[φ(f n )],[t n , t o +t n ]}; Among them, the right-of-way signal state is at t n The right-of-way signal state is (t o + t n ) always switches to the “green light” state.
[0030] A further optimization scheme is that the collaborative graph-based multi-agent reinforcement learning model generates an unmanned vehicle trajectory planning strategy, including the following method:
[0031] Two agents are configured for each parking space, one as the entry agent to communicate with the OD pair of the unmanned vehicle entering the parking space, and the other as the exit agent to communicate with the OD pair of the unmanned vehicle leaving the parking space;
[0032] Using the undirected connected graph G p (V agent , E k,l ) represents a group of intelligent agents that have direct interaction relationships, and a collaborative graph is constructed; the node V of the undirected connected graph agent represents the agent, edge E k,l Represent the state and action interactions between agents;
[0033] Each intelligent agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each intelligent agent based on the congestion state space;
[0034] Generate unmanned vehicle trajectory planning strategy based on the optimal path.
[0035] A further optimization scheme is that the method for constructing the collaborative graph includes:
[0036] With the environment node as the central node, the upper and lower nodes of the collaboration graph represent entering and leaving the agent respectively; both entering and leaving the agent obtain reward signals from the environment node.
[0037] A further optimization scheme is that each intelligent agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each intelligent agent; including methods:
[0038] Each agent imitates the human driver to obtain all feasible behavior options and constructs the agent action set with all behavior options;
[0039] Congestion is determined using queue length as a criterion, and the current congestion state space is constructed: determine whether congestion occurs downstream of the path selected by each behavior;
[0040] Based on the current congestion state space, the optimal path is selected with the principle of avoiding congestion.
[0041] This solution also provides an unmanned apron control system based on reinforcement learning, which is used to implement the above-mentioned unmanned apron control method based on reinforcement learning; the system includes:
[0042] Cluster equipment module, used to deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron;
[0043] A model building module, used to build a traffic topology model of the unmanned apron based on the situation data;
[0044] A strategy generation module is used to construct a locomotive-vehicle-field-track collaborative environment based on the traffic topology model, and generate an unmanned vehicle trajectory planning strategy based on a collaborative graph multi-agent reinforcement learning model in the locomotive-vehicle-field-track collaborative environment;
[0045] The execution module is used to control the unmanned vehicle in the unmanned apron through collaborative cluster equipment according to the unmanned vehicle trajectory planning strategy.
[0046] The present solution also provides a computer-readable medium having a computer program stored thereon, and the computer program is executed by a processor to implement the unmanned apron control method based on reinforcement learning as described above.
[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0048] The present invention provides an unmanned apron control method, system and medium based on reinforcement learning; this solution improves the method on the basis of the traditional unmanned vehicle trajectory planning technology in the apron, and establishes an apron service vehicle road right management mechanism guided by avoiding vehicle-vehicle conflicts to realize the "machine-vehicle-field road" collaborative environment; secondly, the unmanned vehicle trajectory planning strategy is generated based on the collaborative graph multi-agent reinforcement learning model, which combines the anthropomorphic driving concept and the apron area operation rules to realize the unmanned vehicle trajectory planning for the apron scene; finally, a comprehensive simulation platform combining MATLAB software, VISSIM software and Prepar3D software is developed to complete the visualization of the vehicle trajectory planning model in the unmanned apron of a large airport. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:
[0050] Figure 1 This is a flow chart of the unmanned apron control method based on reinforcement learning;
[0051] Figure 2 Schematic diagram of deploying collaborative cluster equipment for unmanned aprons;
[0052] Figure 3 is the directed connected graph of the unmanned apron;
[0053] Figure 4 A schematic diagram of the process of constructing the machine-vehicle-field-track collaborative environment and the process of generating the trajectory planning strategy for the unmanned vehicle;
[0054] Figure 5 This is a schematic diagram of the locomotive-vehicle-field-track coordination mechanism;
[0055] Figure 6 This is the communication interaction diagram of the unmanned apron intelligent agent;
[0056] Figure 7 It is a schematic diagram of the collaboration diagram;
[0057] Figure 8 It is the data flow graph of MATLAB-VISSIM joint simulation;
[0058] Fig. 9 Data flow diagram for VISSIM-Prepar3D joint simulation;
[0059] Fig.10 Gantt chart for agent path selection and virtual traffic signal phase;
[0060] Fig.11 It is a schematic diagram for comparing indicators under different planning methods in the examples of the present invention. DETAILED DESCRIPTION
[0061] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0062] The traditional unmanned vehicle trajectory planning method on the apron does not fully consider the rule constraints during the apron operation, mainly focuses on the unmanned vehicle trajectory control, ignores the conflict avoidance problem, and has large errors in actual application. In view of this, this solution provides the following embodiments to solve the above technical problems.
[0063] Embodiment 1: This embodiment provides an unmanned apron control method based on reinforcement learning, such as Figure 1 As shown, including:
[0064] Step 1: deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron; this step specifically includes the following methods:
[0065] like Figure 2 As shown in the figure, sensing equipment (edge computing equipment) is deployed on the runway to sense the situation data of active targets within the range; including laser radar, millimeter wave radar, etc., which are distributed near each parking position and are responsible for the situation perception of active targets (aircraft, unmanned vehicles) within the sensing range; after the active targets on the apron are paired with the edge computing equipment near each parking position, the global operation situation of the apron traffic system is further formed and uploaded to the centralized dispatching and control platform for solution.
[0066] Control equipment is arranged at the activity target to control the unmanned vehicle to execute the unmanned vehicle trajectory planning strategy; the control equipment mainly includes an on-board automatic control unit and an automatic control unit, the on-board automatic control unit is distributed on each unmanned vehicle and is responsible for receiving the vehicle trajectory planning result, and in the time window corresponding to the vehicle trajectory, the automatic control unit makes the unmanned vehicle drive along the planned path;
[0067] A signal conversion module is deployed on the traffic indication target to build an aircraft-vehicle-airfield-road collaborative environment; traffic control equipment is distributed at each aircraft position and is responsible for receiving aircraft activity information, converting right-of-way signals in the time window corresponding to the aircraft activity, and realizing unmanned vehicle activity control facing the apron outer service lane.
[0068] It also includes the 5G AeroMACS network for drone-to-pad communications; a new generation of aviation broadband communications technology that applies the fifth-generation mobile communications technology (5G) to the civil aviation-specific network AeroMACS. The 5G AeroMACS network is used to transmit vehicle trajectory planning information, aircraft surface 4D trajectory information, and active target control instructions.
[0069] Step 2: constructing a traffic topology model of the unmanned apron based on the situation data; this step specifically includes the following method:
[0070] like Figure 3 As shown in the figure, the service lane, virtual channel, control node and physical node of the unmanned apron are constructed, and a directed connected graph is formed by the control node, physical node and the connection between the control node and the physical node; the directed connected graph represents the movement space range of the unmanned vehicle in the unmanned apron, which is convenient for simulation on the joint platform in the later stage;
[0071] The control node is used to manage activities between similar targets; the control node is arranged at the intersection of the peripheral service lane and the taxiway; specifically, it is arranged at the intersection before the peripheral service lane and the taxiway intersect; the control node is located in the apron peripheral service lane;
[0072] The physical nodes are used to manage activities between heterogeneous targets; the physical nodes are arranged at the turning position of the unmanned apron service lane and the intersection of the service lane and the virtual channel; the physical nodes form the spatial outline of the apron traffic system;
[0073] The virtual channel is used for aircraft to pass between parking stands; the virtual channel is set on the center line between the internal taxiways of two standard parking stands given in accordance with the "Technical Standards for Airfields", and the width of the virtual channel is ≥ the maximum width of the unmanned vehicle in the unmanned apron.
[0074] Step three, constructing a locomotive-vehicle-airfield-track collaborative environment based on the traffic topology model, and generating a trajectory planning strategy for the unmanned vehicle based on the collaborative graph multi-agent reinforcement learning model in the locomotive-vehicle-airfield-track collaborative environment; the locomotive-vehicle-airfield-track collaborative environment controls the time windows of heterogeneous activity targets in the unmanned apron to seamlessly connect the unmanned vehicle with the flight support operation needs, while avoiding activity conflicts between the unmanned vehicle and the aircraft; based on the collaborative graph multi-agent reinforcement learning model, each agent selects a suitable path for the unmanned vehicle based on the real-time unmanned apron congestion identification mechanism; the control system induces the unmanned vehicle to move and control speed. Figure 4 As shown in the figure, this step mainly includes two stages: target coordination stage and path allocation stage. The target coordination stage mainly builds the machine-vehicle-field-track coordination environment; the path allocation stage mainly generates the unmanned vehicle trajectory planning strategy;
[0075] The method of constructing a locomotive-vehicle-field-road collaborative environment based on the traffic topology model includes:
[0076] Get the aircraft's off-block or on-block time t n , the time it takes for the aircraft to slide from the physical node on the taxiway to the parking position is t i , the time taken for the aircraft to push out or taxi from the parking position to the physical node of the apron taxiway t o ; These data of aircraft can be obtained based on the flight schedule and the used seats. At the same time, the vehicle reserves and guarantee nodes of the unmanned vehicle need to be obtained.
[0077] Then, data processing is performed based on the above data, and machine-vehicle coordination and vehicle-field-track coordination are performed respectively to construct the right-of-way signal status and the unmanned vehicle departure time window. Finally, the signal conversion module performs signal switching according to the right-of-way signal status, and the corresponding unmanned vehicle is dispatched according to the unmanned vehicle departure time window.
[0078] like Figure 5 As shown, for inbound aircraft, the parking position φ(f n ) The adjacent control node corresponds to the right-of-way signal state switching time window set R app is {[φ(f n )],[t n -t i , t n ]}; where the right-of-way signal state is (t n -t i ) is switched to the “red light” state, and the right of way of the service lane outside the apron is closed accordingly, and the unmanned vehicle waits outside the stop line; the right of way signal state is at t n It is switched to the "green light" state at all times, and the road rights of the service lanes outside the apron are released accordingly, allowing unmanned vehicles to pass freely.
[0079] For departing aircraft, the parking position φ(f n ) The adjacent control node corresponds to the right-of-way signal state switching time window R dep is {[φ(f n )],[t n , t o +t n ]}; Among them, the right-of-way signal state is at t n The signal is switched to the "red light" state at any time, and the right of way of the service lane outside the apron is closed. The unmanned vehicle waits outside the stop line. The right of way signal state is (t o + t n) is switched to the "green light" state at all times, and the right of way of the apron peripheral service lane is released accordingly, and the unmanned vehicle can pass freely. Through the coordinated relationship between aircraft, vehicle and airport lane, the virtual traffic signal can accurately identify whether the aircraft on the apron is occupying the apron peripheral service lane, so as to close and release the right of way in time.
[0080] The method for generating the unmanned vehicle trajectory planning strategy based on the collaborative graph multi-agent reinforcement learning model includes:
[0081] S31, such as Figure 6 As shown, two agents are configured for each parking space, one as the entering agent to communicate with the OD pairs of the unmanned vehicle entering the parking space, and the other as the leaving agent to communicate with the OD pairs of the unmanned vehicle leaving the parking space. If the number of parking spaces is m, the total number of agents is 2m, and each agent is denoted as A(1)~A(2m). Among them, the entering agents A(1)~A(1m) match the m OD pairs entering the parking space; the leaving agents A(m+1)~A(2m) match the remaining m OD pairs leaving the parking space.
[0082] S32, using the undirected connected graph G p (V agent , E k,l ) represents a group of intelligent agents that have direct interaction relationships, and a collaborative graph is constructed; the node V of the undirected connected graph agent represents the agent, edge E k,l Represent the state and action interactions between agents;
[0083] The method for constructing the collaborative graph includes: Figure 7 As shown in the figure, with the environment node as the central node, the upper and lower nodes of the collaborative graph represent the entering agent and the leaving agent respectively; both the entering agent and the leaving agent receive reward signals from the environment node. Taking the collaborative graph G1 as an example, the upper nodes correspond to parking spaces 1~o adjacent to parking area 1, and the entering agents A(1)~ A(o) match the OD pairs of entering parking space o respectively; the lower nodes correspond to the leaving agents A(o+1)~ A(2o), and match the OD pairs of leaving parking space o respectively.
[0084] In view of the optimization goal of efficient operation of the apron service lane network and the need for normalization of indicators, the reward function for obtaining the reward signal in this embodiment is set to: the inverse of the average queue length of the apron traffic system at the current time. Considering that the lower bound of the queue length is 0, the independent variable of the reward function is translated:
[0085] ;
[0086] Among them, r t represents the immediate reward at time t; represents the average queue length of the apron traffic system at time t;
[0087] For a single simulation round, the cumulative reward r total for:
[0088]
[0089] in, represents the immediate reward at the current moment t*; t total represents all the moments of a single simulation round; It represents the average queue length of the apron traffic system at time t*.
[0090] S33, each intelligent agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each intelligent agent according to the congestion state space; this step specifically includes the following method:
[0091] S331, each intelligent agent imitates the human driver to obtain all feasible behavior options, and constructs an intelligent agent action set with all behavior options; Figure 7 In the traffic scenario shown, there are two feasible behavior choices: choosing the inner service lane and choosing the outer service lane; each agent has two behavior choices. For example, for agent A(1), there is an action set A1={ W 1 , W 2}; W 1 Indicates that the inner service lane is selected. W 2 Indicates the selection of the outer service lane;
[0092] S332, using queue length as a determination indicator to perform congestion determination, and constructing a current congestion state space: determining whether congestion occurs downstream of the path selected by each behavior;
[0093] Combined with the anthropomorphic driving requirements, the state space is the combination of the current moment and the congestion situation of the downstream intersections of each path. W 1 and W 2 Whether there is congestion downstream. There are four combinations of congestion conditions:
[0094] ① (F, F), indicating that there is no congestion downstream of the two optional paths;
[0095] ② (F, T), represents the path W 1 There was no congestion downstream. W2 Congestion occurred downstream;
[0096] ③ (T, F), indicating the path W 1 Congestion occurred downstream. W 2 There was no congestion downstream;
[0097] ④ (T, T), indicating that congestion occurs downstream of both optional paths;
[0098] Queue length is a common congestion determination indicator in the field of road traffic and is easy to read through the autonomous driving simulator. At the same time, considering the relativity of the concept of congestion, different from the previous congestion determination in which the threshold is set as a constant, this embodiment uses the average queue length as an indicator to determine the congestion condition. First, the set of queue lengths of all intersections downstream of path k is recorded as , take the largest element in the set; then, compare this value with the average queue length of the apron traffic system at the current time t For comparison; finally, if the value is greater than the average queue length, it is determined that the path is congested:
[0099] ;
[0100] In summary, (in M =1,2,3,…,2 m ), the state space is set as follows:
[0101] ;
[0102] ;
[0103] Among them, P M represents the state space of the apron traffic system, representing the instantaneous traffic situation (smoothness / congestion of each service lane intersection); x represents the set P M The first element in; y represents the set P M The second element in ;
[0104] S333, based on the current congestion state space, the optimal path is selected based on the principle of avoiding congestion.
[0105] S34, generating a trajectory planning strategy for the unmanned vehicle according to the optimal path.
[0106] Step 4: According to the unmanned vehicle trajectory planning strategy, the unmanned vehicle in the unmanned apron is controlled by collaborative cluster equipment.
[0107] Step 5: Figure 8 and Fig. 9As shown in the figure, on MATLAB software, a locomotive-vehicle-airfield-track collaborative environment is established in combination with a directed connected graph; an unmanned apron micro-model is established on VISSIM software, and parking lot 2 and parking lot 1 are constructed in the VISSIM simulation environment. Six parking spaces numbered 327 to 332 are set between the inner and outer paths of parking lots 2 and 1, with parking space 327 close to parking lot 1 and parking space 332 close to parking lot 2; the VISSIM simulation environment interacts with the MATLAB software through the VISSIM software and its COM interface, and the VISSIM simulation environment sends data to the VISSIM software and its C The OM interface sends the path induction action set, and the VISSIM software and its COM interface feed back the implementation traffic status set to the VISSIM simulation environment; the virtual traffic signal unit state switching time slot is input into the VISSIM software and its COM interface through the signal setting handle, and is read by the VISSIM simulation environment; the unmanned vehicle departure time slot is input into the VISSIM software and its COM interface through the vehicle input handle, and is read by the VISSIM simulation environment; a three-dimensional model of the unmanned apron is established on the Prepar3D software; a "one picture" of the coordinated operation of various activity targets is formed to complete the visualization of the unmanned apron scene of large airports.
[0108] Based on the Prepar3D SDK development kit and Airport Design Editor, Figure 3 The directed connected graph of the unmanned apron is imported into Prepar3D; secondly, the unmanned vehicle trajectory is translated through the data processing terminal, and the trajectory data output by VISSIM is projected into Prepar3D; finally, a visualization scene of the unmanned apron of a large airport is formed.
[0109] The unmanned vehicle trajectory translation is to export the .trj standard trajectory format in VISSIM and generate a .csv report through the SSAM trajectory decoding software; at this time, the trajectory is a seven-dimensional vector
[0110] The internal variables of the seven-dimensional vector represent the vehicle number, time period, two-dimensional rectangular coordinates of the front of the vehicle, two-dimensional rectangular coordinates of the rear of the vehicle, and the vehicle's ground speed in sequence.
[0111] Translate the 2D rectangular coordinates of the front and rear of the vehicle into the coordinates of the center point of the vehicle (X, Y), where
[0112] ;
[0113] The origin of the rectangular coordinates in VISSIM is (X0, Y0), which corresponds to the origin of the WGS84 geographic coordinates in Prepar3D. . Then calculate the geographic coordinates (Lon, Lat) of the vehicle center point in the Prepar3D environment: ;
[0114] At this point, the VISSIM unmanned vehicle trajectory is translated into the format required by Prepar3D:
[0115] .
[0116] Embodiment 2: This embodiment provides an unmanned apron control system based on reinforcement learning, which is used to implement the unmanned apron control method based on reinforcement learning in Embodiment 1; the system includes:
[0117] Cluster equipment module, used to deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron;
[0118] A model building module, used to build a traffic topology model of the unmanned apron based on the situation data;
[0119] A strategy generation module is used to construct a locomotive-vehicle-field-track collaborative environment based on the traffic topology model, and generate an unmanned vehicle trajectory planning strategy based on a collaborative graph multi-agent reinforcement learning model in the locomotive-vehicle-field-track collaborative environment;
[0120] The execution module is used to control the unmanned vehicle in the unmanned apron through collaborative cluster equipment according to the unmanned vehicle trajectory planning strategy.
[0121] Embodiment 3: This embodiment provides a computer-readable medium on which a computer program is stored. The computer program is executed by a processor to implement the unmanned apron control method based on reinforcement learning as described in Embodiment 1, and specifically performs the following steps:
[0122] Step 1: deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron;
[0123] Step 2: constructing a traffic topology model of the unmanned apron based on the situation data;
[0124] Step 3: construct a machine-vehicle-field-track collaborative environment based on the traffic topology model, and generate an unmanned vehicle trajectory planning strategy based on a collaborative graph multi-agent reinforcement learning model in the machine-vehicle-field-track collaborative environment;
[0125] Step 4: According to the unmanned vehicle trajectory planning strategy, the unmanned vehicle in the unmanned apron is controlled by collaborative cluster equipment.
[0126] Step 5. On the MATLAB software, establish the locomotive-vehicle-airfield-track collaborative environment in combination with the directed connectivity graph; establish the unmanned apron micromodel on the VISSIM software; and establish the unmanned apron three-dimensional model on the Prepar3D software.
[0127] like Fig.10As shown, this embodiment jointly compares the path selection of each intelligent agent and the time-varying characteristics of the virtual traffic signal. The results show that the two are tightly coupled, realizing the path planning of the unmanned vehicle under the vehicle-vehicle cooperative mechanism. As shown in the figure, the No. 331 aircraft position and its corresponding intelligent agents A(4) and A(8) entering and leaving the aircraft position, after the virtual traffic signal turns "red", the control node near the corresponding aircraft position starts road right control; the intelligent agents on the corresponding OD pair will focus on selecting paths in the future [200s, 1000s] and [4200s, 5000s] intervals. W 1 Agent Selection W 1 The corresponding red areas are relatively dense, indicating that the other time periods have similar linkage characteristics. However, the red areas corresponding to the other time periods are not concentrated into pieces. The main reason is that there are no vehicles leaving in some time periods, and the agent still retains the choice W 2 In summary, under the collaborative graph multi-agent reinforcement model, each agent can dynamically select an appropriate path according to the downstream queue situation of the path to achieve smooth travel within the apron.
[0128] This scheme also conducts MATLAB-VISSIM-Prepar3D joint simulation under the conditions of no planning, traditional multi-agent reinforcement learning (MARL) and collaborative graph multi-agent reinforcement learning of the embodiment of the present invention. The average vehicle speed, system calculation time and vehicle queue length indicators under the three methods are as follows: Fig.11 As shown. In terms of average vehicle speed, the collaborative graph multi-agent reinforcement learning method of the present invention is 28.53% and 11.60% higher than the unplanned and traditional MARL methods, respectively; in terms of system calculation time, the method of the present invention improves the calculation speed by 3.19% compared with the traditional MARL method. 1591.9s can plan the activity trajectory of the unmanned vehicle on the apron for the next 12000s, and the ratio of time to be planned and calculated is 7.54:1. Therefore, the method of the present invention provides a basis for the subsequent online planning of unmanned vehicles. In terms of average queue length, the collaborative graph multi-agent reinforcement learning method of the present invention is 68.72% and 32.34% lower than the unplanned and traditional MARL methods, respectively. In summary, the method of the present invention has a more significant optimization effect in improving the average vehicle speed and reducing the queue length; at the same time, it can reduce a certain amount of system calculation time.
[0129] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An unmanned apron control method based on reinforcement learning, characterized in that: include: Deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron; Constructing a traffic topology model of the unmanned apron based on the situation data; The method includes: constructing a service lane, a virtual channel, a control node and a physical node of an unmanned apron, and forming a directed connected graph by the control node, the physical node and the connection between the control node and the physical node; the control node is used to manage activities between similar targets; the control node is arranged at the intersection of the peripheral service lane and the taxiway; the physical node is used to manage activities between heterogeneous targets; the physical node is arranged at the turning position of the service lane of the unmanned apron and the intersection of the service lane and the virtual channel; the virtual channel is used for aircraft to pass between parking positions; the virtual channel is arranged on the center line between the taxiways inside two adjacent standard parking positions, and the width of the virtual channel is ≥ the maximum width of the unmanned vehicle in the unmanned apron; Based on the traffic topology model, a machine-vehicle-field-track collaborative environment is constructed, and in the machine-vehicle-field-track collaborative environment, a trajectory planning strategy for the unmanned vehicle is generated based on a collaborative graph multi-agent reinforcement learning model; The method of constructing an aircraft-vehicle-airport-road collaborative environment based on the traffic topology model includes: obtaining the wheel block removal or wheel block time t of the aircraft n , the time it takes for the aircraft to slide from the physical node on the taxiway to the parking position is t i , the time taken for the aircraft to push out or taxi from the parking position to the physical node of the apron taxiway t o ; For arriving aircraft, parking position φ (f n ) The adjacent control node corresponds to the right-of-way signal state switching time window set R app is {[φ(f n )],[t n -t i , t n ]}; Among them, the right-of-way signal state is (t n -t i ) is switched to the "red light" state, and the right-of-way signal state is at t n Always switch to "green light" status; for departing aircraft, the parking position φ (f n ) The adjacent control node corresponds to the right-of-way signal state switching time window R dep is {[φ(f n )],[t n , t o +t n ]}; the right-of-way signal status is at t n The right-of-way signal status is (t o + t n ) is always switched to the "green light" state; The collaborative graph-based multi-agent reinforcement learning model generates an unmanned vehicle trajectory planning strategy, including a method of configuring two agents for each parking space, one as an entry agent to communicate with the OD pair of the unmanned vehicle entering the parking space, and the other as a leaving agent to communicate with the OD pair of the unmanned vehicle leaving the parking space; using an undirected connected graph G p (V agent , E k,l ) represents a group of intelligent agents that have direct interaction relationships, and a collaborative graph is constructed; the node V of the undirected connected graph agent represents the agent, edge E k,l Represent the state and action interaction between agents; each agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each agent based on the congestion state space; generates the unmanned vehicle trajectory planning strategy based on the optimal path; The method for constructing the collaborative graph includes: using the environment node as the central node, the upper node and the lower node of the collaborative graph represent entering the intelligent agent and leaving the intelligent agent respectively; both entering the intelligent agent and leaving the intelligent agent obtain a reward signal from the environment node; According to the unmanned vehicle trajectory planning strategy, the unmanned vehicle in the unmanned apron is controlled through collaborative cluster equipment.
2. The unmanned apron control method based on reinforcement learning according to claim 1 is characterized in that: The method of deploying collaborative cluster equipment on the unmanned apron and obtaining situation data of the unmanned apron includes: Sensing equipment is deployed on the track to sense the situation data of moving targets within the range; Deploy control equipment at the moving target to control the unmanned vehicle to execute the unmanned vehicle trajectory planning strategy; A signal conversion module is deployed on the traffic indication target to build a locomotive-vehicle-field-track collaborative environment.
3. The unmanned apron control method based on reinforcement learning according to claim 1 is characterized in that: The method further comprises: On MATLAB software, a locomotive-vehicle-airfield-track collaborative environment is established in combination with a directed connectivity graph; an unmanned apron micromodel is established on VISSIM software; and an unmanned apron three-dimensional model is established on Prepar3D software.
4. The unmanned apron control method based on reinforcement learning according to claim 1 is characterized in that: Each intelligent agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each intelligent agent; including methods: Each agent imitates the human driver to obtain all feasible behavior options and constructs the agent action set with all behavior options; Congestion is determined using queue length as a criterion, and the current congestion state space is constructed: determine whether congestion occurs downstream of the path selected by each behavior; Based on the current congestion state space, the optimal path is selected with the principle of avoiding congestion.
5. The unmanned apron control system based on reinforcement learning is characterized by: Used to implement the unmanned apron control method based on reinforcement learning as described in any one of claims 1 to 4; the system comprises: Cluster equipment module, used to deploy collaborative cluster equipment on the unmanned apron and obtain situation data of the unmanned apron; A model building module is used to build a traffic topology model of an unmanned apron based on the situation data; specifically: construct service lanes, virtual channels, control nodes and physical nodes of the unmanned apron, and a directed connected graph is formed by control nodes, physical nodes and the lines connecting control nodes and physical nodes; the control nodes are used to manage activities between similar targets; the control nodes are arranged at the intersection of the outer service lane and the taxiway; the physical nodes are used to manage activities between heterogeneous targets; the physical nodes are arranged at the turning position of the service lane of the unmanned apron and the intersection of the service lane and the virtual channel; the virtual channel is used for aircraft to pass between parking positions; the virtual channel is set on the center line between the taxiways inside two adjacent standard parking positions, and the width of the virtual channel is ≥ the maximum width of the unmanned vehicle in the unmanned apron; The strategy generation module is used to construct a vehicle-vehicle-field-track collaborative environment based on the traffic topology model, and generate an unmanned vehicle trajectory planning strategy based on the collaborative graph multi-agent reinforcement learning model in the vehicle-vehicle-field-track collaborative environment; the construction of the vehicle-vehicle-field-track collaborative environment based on the traffic topology model specifically includes: obtaining the aircraft's wheel block withdrawal or wheel block time t n , the time it takes for the aircraft to slide from the physical node on the taxiway to the parking position is t i , the time taken for the aircraft to push out or taxi from the parking position to the physical node of the apron taxiway t o ; For arriving aircraft, parking position φ (f n ) The adjacent control node corresponds to the right-of-way signal state switching time window set R app is {[φ(f n )],[t n -t i , t n ]}; Among them, the right-of-way signal state is (t n -t i ) is switched to the "red light" state, and the right-of-way signal state is at t n Always switch to "green light" status; for departing aircraft, the parking position φ (f n ) The adjacent control node corresponds to the right-of-way signal state switching time window R dep is {[φ(f n )],[t n , t o + t n ]}; the right-of-way signal status is at t n The right-of-way signal status is (t o + t n ) is always switched to the "green light" state; The strategy for generating the trajectory planning of the unmanned vehicle based on the collaborative graph multi-agent reinforcement learning model specifically includes: configuring two agents for each parking space, one as an entry agent to communicate with the OD pair of the unmanned vehicle entering the parking space, and the other as a leaving agent to communicate with the OD pair of the unmanned vehicle leaving the parking space; using the undirected connected graph G p (V agent , E k,l ) represents a group of intelligent agents that have direct interaction relationships, and a collaborative graph is constructed; the node V of the undirected connected graph agent represents the agent, edge E k,l Represent the state and action interaction between agents; each agent considers imitating the path selection behavior of human drivers, constructs the current congestion state space, and performs optimal path selection for the unmanned vehicles on the communication OD pairs generated by each agent based on the congestion state space; generates the unmanned vehicle trajectory planning strategy based on the optimal path; The method for constructing the collaborative graph includes: using the environment node as the central node, the upper node and the lower node of the collaborative graph represent entering the intelligent agent and leaving the intelligent agent respectively; both entering the intelligent agent and leaving the intelligent agent obtain a reward signal from the environment node; The execution module is used to control the unmanned vehicle in the unmanned apron through collaborative cluster equipment according to the unmanned vehicle trajectory planning strategy.
6. A computer readable medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the unmanned apron control method based on reinforcement learning as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Airport surface variable slide-out time prediction method based on big data deep learning
WO2021082393A1
Unmanned ground vehicle-unmanned aerial vehicle collaborative autonomous tracking and landing method
WO2023097769A1