Method and apparatus for determining aircraft flight strategy

CN115470881BActive Publication Date: 2026-08-11HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-25
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

这不仅降低了飞行的安全性,而且给航空带来了巨大的经济损失

Benefits of technology

[0056]根据本申请实施例的方案,提供了一种用于空中交通流量优化的智能体训练方法和装置,通过使用多智能体强化学习来进行路线选择与速度调节的决策,使得飞行器平稳起飞或精准着陆,有效利用有限的起降场地,最大化空域利用率,进而优化终端区流量的分配。自动为飞行器选择最优进近路线或起飞路线,并调节飞行器在航线上的速度,能够提高航空运输效率,解决空中交通流量管理问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470881B_ABST
    Figure CN115470881B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for determining an aircraft's flight strategy. The method includes: acquiring a first model, which is obtained by training based on first training data, the first training data including flight state information of a first aircraft in a first time period, flight state information of at least one second aircraft in the first time period, and a target flight strategy of the first aircraft in the first time period; acquiring first parameters, the first parameters including flight state information of a third aircraft in a second time period and flight state information of at least one fourth aircraft in the second time period; and inputting the first parameters into the first model to obtain the flight strategy of the third aircraft. This method enables multiple agents to autonomously select routes for aircraft and adjust their speeds along flight paths through extensive training, thereby optimizing terminal area traffic allocation and improving air transport efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more specifically, to a method and apparatus for determining an aircraft flight strategy. Background Technology

[0002] Modern cities offer a diverse range of transportation options. With increasing population density and limited land resources in megacities, many cities are experiencing escalating traffic congestion and environmental pollution. Therefore, there is a need to expand available urban transportation modes in ways that can reduce surface traffic flow without requiring large amounts of land.

[0003] In recent years, my country's air transport industry has achieved rapid development. However, with the continuous growth in air traffic demand, the demand for airspace resources is also increasing, leading to increasingly prominent air traffic congestion. This not only reduces flight safety but also causes huge economic losses to aviation. Air traffic flow management is currently the most effective and economical means to solve air traffic congestion.

[0004] Artificial intelligence (AI) is a new technical science that studies the theories, methods, technologies, and application systems used to simulate, extend, and expand human intelligence. Reinforcement learning is a general method for achieving sequential decision-making. An agent learns through trial and error, using rewards obtained from interacting with the environment to guide its behavior, thereby maximizing the agent's reward.

[0005] Therefore, in urban air transport scenarios with extremely high traffic volume, how to enable intelligent agents to autonomously select flight routes and adjust speeds to improve air traffic management efficiency and aviation transport efficiency is an urgent problem to be solved. Summary of the Invention

[0006] This application provides a method and apparatus for determining the flight strategy of an aircraft, which enables an intelligent agent to autonomously select flight routes and adjust speeds in urban air transportation scenarios with extremely high traffic volume, thereby improving the efficiency of air traffic management and air transport.

[0007] In a first aspect, a method for determining an aircraft's flight strategy is provided, comprising: acquiring a first model, the first model being trained based on first training data, the first training data including flight state information of a first aircraft in a first time period, flight state information of at least one second aircraft in the first time period, and a target flight strategy of the first aircraft in the first time period, wherein the second aircraft is an aircraft located within a first range in the first time period, the first range being determined based on the position of the first aircraft; acquiring a first parameter, the first parameter including flight state information of a third aircraft in a second time period, and flight state information of at least one fourth aircraft in the second time period, wherein the fourth aircraft is an aircraft located within a second range in the second time period, the second range being determined based on the position of the third aircraft; and inputting the first parameter into the first model to obtain the flight strategy of the third aircraft.

[0008] According to the scheme provided in this application, through extensive training and learning of multiple agents, the agents acquire the ability to make autonomous choices, enabling aircraft to take off smoothly or land precisely, effectively utilize limited takeoff and landing sites, maximize airspace utilization, autonomously select the optimal approach or takeoff route for the aircraft, and adjust the aircraft's speed along the route, thereby optimizing the allocation of traffic in the terminal area and improving air transport efficiency.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the first model is trained based on the first training data.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the first training data is input into the original model to obtain the first flight strategy; the first flight strategy is sent to the server, which stores the target flight strategy; first reward data is obtained from the server, which is determined based on the relationship between the first flight strategy and the target flight strategy; the original model is adjusted based on the first reward data to determine the first model.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the first parameter satisfies:

[0012]

[0013] in, This indicates the flight status information of the third aircraft during the second time period, I (0) This indicates the distance and orientation between the current position and the landing position of the third aircraft, the current velocity and acceleration of the third aircraft, and the attitude information of the third aircraft. (i)This represents the distance and orientation between the current position and landing position of the i-th aircraft, the current velocity and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft. (i) LOS(o,i) represents the distance from the third aircraft to the i-th aircraft, LOS(o,i) represents the distance loss between the third aircraft and the i-th aircraft, n represents the number of the fourth aircraft, and i is a positive integer greater than or equal to 1 and less than or equal to n.

[0014] In conjunction with the first aspect, in some implementations of the first aspect, the action space of the third aircraft satisfies:

[0015]

[0016] in, This indicates the takeoff and landing route selected by the third aircraft, v min This indicates the minimum cruising speed of the third aircraft along the takeoff and landing route, v. t This represents the current speed of the third aircraft, v. max This indicates the maximum cruising speed of the third aircraft on the takeoff and landing route.

[0017] In conjunction with the first aspect, in some implementations of the first aspect, the termination criterion of the third aircraft satisfies:

[0018] N aircraft =0

[0019] Where, N aircraft This indicates the number of aircraft located within the terminal airspace of this second range.

[0020] In conjunction with the first aspect, in some implementations of the first aspect, the terminal area airspace includes multiple take-off and landing routes, and the starting and ending points of the approach routes of each of the multiple take-off and landing routes are configured with vertical take-off and landing airspace.

[0021] In conjunction with the first aspect, in some implementations of the first aspect, the vertical take-off and landing airspace is a stepped cylindrical airspace, which is configured with multiple approach points in different directions and sets multiple approach paths with different angles. The method further includes: when the third aircraft enters the stepped cylindrical airspace and encounters congestion, or when the stepped cylindrical airspace exceeds a first preset threshold, controlling the third aircraft to execute a waiting procedure, wherein the first preset threshold is the maximum number of aircraft that the stepped cylindrical airspace can accommodate.

[0022] In conjunction with the first aspect, in some implementations of the first aspect, the first reward data is calculated based on a reward function that satisfies:

[0023]

[0024] in, The distance between the third and fourth aircraft is represented by λ, α, and β, which are positive constants. t This indicates the total time for the takeoff and landing of the third aircraft.

[0025] In conjunction with the first aspect, in some implementations of the first aspect, when a conflict occurs between the third aircraft and the fourth aircraft, reward data corresponding to the conflict is obtained from the server. The conflict is used to indicate that the distance between the third aircraft and the fourth aircraft is less than a second preset threshold, which is the minimum safe distance between the third aircraft and the fourth aircraft.

[0026] In conjunction with the first aspect, in some implementations of the first aspect, the learning and training of the agent can be carried out according to the algorithm of the actor critic, such as: advantage actor critic (A2C), asynchronous advantage actor-critic (A3C) algorithm, soft actorcritic (SAC) algorithm, proximal policy optimization algorithm (PPO) algorithm, deep deterministic policy gradient (DDPG) algorithm, etc., and this application does not limit it.

[0027] In a second aspect, an apparatus for determining an aircraft flight strategy is provided, comprising: an acquisition unit for acquiring a first model, the first model being trained based on first training data, the first training data including flight status information of a first aircraft during a first time period, flight status information of at least one second aircraft during the first time period, and a target flight strategy of the first aircraft during the first time period, wherein the second aircraft is an aircraft located within a first range during the first time period, and the first range is determined based on the position of the first aircraft; the acquisition unit is further configured to acquire a first parameter, the first parameter including flight status information of a third aircraft during a second time period, and flight status information of at least one fourth aircraft during the second time period, wherein the fourth aircraft is an aircraft located within a second range during the second time period, and the second range is determined based on the position of the third aircraft; and a transceiver unit for inputting the first parameter into the first model to obtain the flight strategy of the third aircraft.

[0028] In conjunction with the second aspect, in some implementations of the second aspect, the apparatus further includes: a processing unit for training the first model based on the first training data.

[0029] In conjunction with the second aspect, in some implementations of the second aspect, the transceiver unit is further configured to: input the first training data into the original model to obtain a first flight strategy; send the first flight strategy to a server, the server storing the target flight strategy; the acquisition unit further includes acquiring first reward data from the server, the first reward data being determined based on the relationship between the first flight strategy and the target flight strategy; the processing unit further includes adjusting the original model based on the first reward data to determine the first model.

[0030] In conjunction with the second aspect, in some implementations of the second aspect, the first parameter satisfies:

[0031]

[0032] in, This indicates the flight status information of the third aircraft during the second time period, I (0) This indicates the distance and orientation between the current position and the landing position of the third aircraft, the current velocity and acceleration of the third aircraft, and the attitude information of the third aircraft. (i) This represents the distance and orientation between the current position and landing position of the i-th aircraft, the current velocity and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft. (i) LOS(o,i) represents the distance from the third aircraft to the i-th aircraft, LOS(o,i) represents the distance loss between the third aircraft and the i-th aircraft, n represents the number of the fourth aircraft, and i is a positive integer greater than or equal to 1 and less than or equal to n.

[0033] In conjunction with the second aspect, in some implementations of the second aspect, the action space of the third aircraft satisfies:

[0034]

[0035] in, This indicates the takeoff and landing route selected by the third aircraft, v min This indicates the minimum cruising speed of the third aircraft along the takeoff and landing route, v. t This represents the current speed of the third aircraft, v. max This indicates the maximum cruising speed of the third aircraft on the takeoff and landing route.

[0036] In conjunction with the second aspect, in some implementations of the second aspect, the termination criterion of the third aircraft satisfies:

[0037] N aircraft =0

[0038] Where, N aircraftThis indicates the number of aircraft located within the terminal airspace of this second range.

[0039] In conjunction with the second aspect, in some implementations of the second aspect, the terminal area airspace includes multiple take-off and landing routes, and the starting and ending points of the approach routes of each of the multiple take-off and landing routes are configured with vertical take-off and landing airspace.

[0040] In conjunction with the second aspect, in some implementations of the second aspect, the vertical take-off and landing airspace is a stepped cylindrical airspace. The stepped cylindrical airspace is configured with multiple approach points in different directions and multiple approach paths with different angles are set. The processing unit is also used to control the third aircraft to execute a waiting procedure when the third aircraft enters the stepped cylindrical airspace and causes congestion, or when the stepped cylindrical airspace exceeds a first preset threshold. The first preset threshold is the maximum number of aircraft that the stepped cylindrical airspace can accommodate.

[0041] In conjunction with the second aspect, in some implementations of the second aspect, the first reward data is calculated based on a reward function that satisfies:

[0042]

[0043] in, The distance between the third and fourth aircraft is represented by λ, α, and β, which are positive constants. t This indicates the total time for the takeoff and landing of the third aircraft.

[0044] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is further configured to obtain reward data corresponding to the conflict from the server when a conflict occurs between the third aircraft and the fourth aircraft. The conflict is used to indicate that the distance between the third aircraft and the fourth aircraft is less than a second preset threshold, which is the minimum safe distance between the third aircraft and the fourth aircraft.

[0045] In conjunction with the second aspect, in some implementations of the second aspect, the learning and training of the agent can be carried out according to the algorithm of the actor critic, such as: advantage actor critic (A2C), asynchronous advantage actor-critic (A3C) algorithm, soft actorcritic (SAC) algorithm, proximal policy optimization algorithm (PPO) algorithm, deep deterministic policy gradient (DDPG) algorithm, etc., and this application does not limit it.

[0046] Thirdly, an apparatus for determining an aircraft flight strategy is provided, the apparatus including a processor coupled to a memory for storing computer programs or instructions, the processor for executing the computer programs or instructions stored in the memory, such that the method in the first aspect or any possible implementation of the first aspect is executed.

[0047] Optionally, the device may include one or more processors and one or more memories.

[0048] Optionally, the device may include one or more memories.

[0049] Alternatively, the memory can be integrated with the processor, or the memory can be set separately from the processor.

[0050] Optionally, the device may also include a transceiver, which may specifically be a transmitter and a receiver.

[0051] Fourthly, an apparatus for determining an aircraft flight strategy is provided, comprising: various modules or units for implementing the method in the first aspect or any possible implementation of the first aspect.

[0052] Fifthly, a computer-readable storage medium is provided that stores a computer program or code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0053] In a sixth aspect, a chip is provided, including at least one processor coupled to a memory for storing a computer program, the processor for calling and running the computer program from the memory, such that a communication device having the chip system installed performs the method of the first aspect or any possible implementation thereof.

[0054] The chip may include input circuitry or interface for transmitting information and / or data, and output circuitry or interface for receiving information and / or data.

[0055] In a seventh aspect, a computer program product is provided, comprising: computer program code, which, when executed by a computer, causes the method in the first aspect or any possible implementation thereof to be implemented.

[0056] According to the embodiments of this application, an intelligent agent training method and apparatus for optimizing air traffic flow are provided. By using multi-agent reinforcement learning to make decisions on route selection and speed adjustment, aircraft can take off smoothly or land precisely, effectively utilize limited takeoff and landing sites, maximize airspace utilization, and thus optimize the allocation of traffic in the terminal area. Automatically selecting the optimal approach or takeoff route for aircraft and adjusting the aircraft's speed along the route can improve air transport efficiency and solve air traffic flow management problems. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of a multi-agent system to which this application applies.

[0058] Figure 2 This is a schematic diagram of a training process for reinforcement learning applicable to this application.

[0059] Figure 3 A schematic diagram of a multilayer perceptron to which this application applies;

[0060] Figure 4 A schematic diagram illustrating an example of loss function optimization applicable to this application;

[0061] Figure 5 A schematic diagram illustrating gradient backpropagation applicable to this application;

[0062] Figure 6 This is an example schematic diagram of a method for determining the flight strategy of an aircraft to which this application applies.

[0063] Figure 7 This is another schematic diagram illustrating the method for determining the flight strategy of an aircraft to which this application applies.

[0064] Figure 8This is a schematic diagram of an example of a device for determining the flight strategy of an aircraft to which this application applies.

[0065] Figure 9 This is another schematic diagram of the apparatus for determining the flight strategy of an aircraft to which this application applies. Detailed Implementation

[0066] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0067] Figure 1 This is a multi-agent system applicable to this application. Multi-agent collaboration is an application scenario in the field of artificial intelligence. For example, in a communication network containing multiple routers, each router can be regarded as an agent, and each router has its own traffic scheduling strategy. The traffic scheduling strategies of multiple routers need to be coordinated with each other in order to complete the traffic scheduling task with fewer resources.

[0068] Figure 1 In the diagram, A through F represent six routers, each with a neural network deployed on it. Therefore, one router is equivalent to one agent, and training an agent is equivalent to training the neural network deployed on that agent. The connections between routers represent communication lines. A through D represent four edge routers. The traffic between edge routers is called aggregated flow. For example, traffic from A to C is one aggregated flow, and traffic from C to A is another aggregated flow.

[0069] The aggregated flow between multiple routers can be generated by N B (N B -1) Determine, N B This represents the number of edge routers among the multiple routers. Figure 1 In the system shown, there are 4 edge routers, therefore, there are a total of 12 aggregated flows in the system.

[0070] For each aggregated flow, the multipath routing algorithm has already determined the available paths. Routers can determine the available paths based on routing table entries (S, D, Nexthop1, rate1%, Nexthop2, rate2%, Nexthop3, rate3%, ...). Here, S represents the originating router, D represents the destination router, Nexthop1, Nexthop2, and Nexthop3 represent different next hops, and rate1%, rate2%, and rate3% represent the percentage of traffic forwarded at each different next hop relative to the total forwarded traffic. The sum of all rates equals 100%.

[0071] A specific task of the above system is to determine the strategy for any one of the routers A through F to autonomously select routes and adjust speeds.

[0072] One method to accomplish the specific task mentioned above is to treat any one of the routers A through F as an agent, and train the agent to make appropriate route selection and speed adjustment strategies.

[0073] In order to describe the embodiments of this application, several terms involved in the embodiments of this application will be introduced first.

[0074] Air route: The path an aircraft takes is called an air traffic route, or simply an air route. An aircraft's air route not only determines the specific direction, origin, destination, and stopover points of its flight, but also, according to the needs of air traffic control, specifies the width and altitude of the air route to maintain air traffic order and ensure flight safety.

[0075] Airspace refers to the space occupied by flight. It is usually marked by prominent landmarks or navigation beacons. Like territory and territorial waters, airspace is within a nation's sovereign domain and is an important military and civilian aviation resource.

[0076] Terminal area: can be understood as a circular airspace extending upwards with a radius of about 10 kilometers from the airport.

[0077] Artificial Intelligence (AI): A branch of computer science that attempts to understand the nature of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Research in the field of AI includes robotics, speech recognition, image recognition, natural language processing, decision-making and reasoning, human-computer interaction, recommendation and search, etc.

[0078] Machine learning is the core of artificial intelligence. Industry experts define machine learning as: a process of progressively improving model performance P through training process E to achieve task T. For example, let's say a model is learning to distinguish between a cat and a dog in a picture (task T). To improve the model's accuracy (model performance P), we continuously provide it with pictures to help it learn the differences between cats and dogs (training process E). The final model obtained through this learning process is the product of machine learning. Ideally, the trained model should be able to recognize cats and dogs in pictures. This training process is the learning process in machine learning. Machine learning methods include reinforcement learning.

[0079] An intelligent agent is a software or hardware entity capable of autonomous action and decision-making, while the environment refers to the external conditions outside the intelligent agent. In communication systems, an intelligent agent can be software that makes decisions or a combination of software and hardware, while the environment is the sum of all external conditions other than the software or hardware entity. For example, an intelligent agent can be a computer system or part of a computer system within a specific environment. An intelligent agent can autonomously achieve its set goals within its environment by recognizing its own environment, following existing instructions, learning autonomously, and communicating and collaborating with other intelligent agents.

[0080] Intelligent vehicles mainly include stationary robots, such as robotic arms used in industrial production; and mobile robots, which move in the environment using wheels, legs or similar machines. Common examples include cargo robots (such as warehouse robots), aerial robots (such as drones), and automated vehicles (such as autonomous vehicles).

[0081] A policy function is a rule used by an agent to adopt actions in reinforcement learning. For example, during learning, an agent can output an action based on its state and use this action to explore the environment and update its state. The update of the policy function depends on the policy gradient (PG). The policy function is typically a neural network. For example, this neural network can include a multilayer perceptron. In practical applications, deep neural networks are commonly used as the agent's policy function. The agent observes the environment, obtains its current state, and decides on an action according to a certain rule (policy), feeding this action back to the environment. The environment then provides the agent with either a reward or penalty for performing the action. Through multiple iterations, the agent learns to make optimal decisions based on the environmental state.

[0082] To facilitate understanding of the technical solutions proposed in this application, we will first introduce decision models, reinforcement learning, and neural networks.

[0083] A decision model can be understood as a model for analyzing decision problems. The scheduling of wireless resources is a decision problem, and a decision model can be constructed for it.

[0084] Markov decision processes (MDPs) are common models in reinforcement learning. They are mathematical models for analyzing decision problems based on discrete-time stochastic control. They assume that the environment possesses the Markov property, meaning the conditional probability distribution of the environment's future states depends only on the current state. Decision-makers periodically observe the state of the environment and make decisions (also known as actions) based on the current state, interacting with the environment to obtain the next state and reward.

[0085] Wireless resource scheduling plays a crucial role in cellular networks. Essentially, it involves allocating available wireless spectrum and other resources based on factors such as channel quality and Quality of Service (QoS) requirements for each user. This application establishes the wireless resource scheduling process as a Multi-Agent Programming (MDP) process, employing reinforcement learning from artificial intelligence (AI) to solve it. Furthermore, it proposes an agent-based autonomous decision-making method, utilizing multi-agent reinforcement learning to make decisions on route selection and speed adjustment.

[0086] Reinforcement learning (RL) is a field within machine learning used to solve Markov decision processes. Also known as reward learning, evaluation learning, or reinforcement learning, it is a general method for implementing sequential decision-making. It describes and solves problems where an agent learns strategies to maximize rewards or achieve specific goals during interactions with its environment.

[0087] Generally, in the field of artificial intelligence, reinforcement learning is a process in which an intelligent agent learns through trial and error, using rewards obtained from interacting with the environment to guide its behavior. The goal is to maximize the reward obtained by the intelligent agent through reinforcement learning.

[0088] Reinforcement learning does not require a training dataset. In reinforcement learning, the reinforcement signals (i.e., rewards) provided by the environment evaluate the quality of an action, rather than telling the reinforcement learning system how to produce the correct action. Because the external environment provides very little information, the agent must learn through its own experience. In this way, the agent acquires knowledge in the action-evaluation (i.e., reward) environment and improves its action plans to adapt to the environment.

[0089] MDP-based reinforcement learning can be categorized into two types: environment-based state transition modeling and environment-free models. The former requires modeling the environment's state transitions, typically relying on empirical knowledge or data fitting. The latter, however, does not require environment-based state transition modeling but rather improves through continuous exploration and learning within the environment. Since the real-world environments that reinforcement learning focuses on are often more complex and unpredictable than the models built (e.g., robots, Go), environment-free reinforcement learning methods are often easier to implement and fine-tune.

[0090] Figure 2 This is a schematic diagram of a reinforcement learning training method. (For example...) Figure 2As shown, reinforcement learning mainly comprises four elements: agent, environment state, action, and reward. The agent's input is the state, and its output is the action. The agent 110 includes a decision policy (i.e., a policy function), which can be an algorithm represented by a formula or a neural network.

[0091] The current training process for an agent in reinforcement learning is as follows: the agent interacts with the environment multiple times to obtain the action, state, and reward for each interaction; these multiple sets of (action, state, reward) are used as training data to train the agent once. This process is repeated for the next round of training until the convergence condition is met.

[0092] The process of obtaining the action, state, and reward of an interaction is as follows: Figure 1 As shown, the current state s(t) 130 of the environment is input to the agent 110, and the action a(t) 140 output by the agent is obtained. Based on the relevant performance indicators of the environment 120 under the action a(t), the reward r(t) 160 of this interaction is calculated. Thus, the state s(t) 130, action a(t) 140, and reward r(t) 160 of this interaction are obtained. The state s(t) 130, action a(t) 140, and reward r(t) 160 of this interaction are recorded for subsequent training of the agent. The next state s(t+1) 150 of the environment under the action a(t) is also recorded to enable the next interaction between the agent and the environment.

[0093] In other words, at each time t, the state s(t) observed by the decision-maker will transition to the next state s(t+1) under the influence of the action a(t), and a reward r(t) will be given as feedback. Here, s(t) represents the state function, a(t) represents the action function, r(t) represents the reward value, and t represents time.

[0094] Specifically, the implementation steps of reinforcement learning training methods are as follows:

[0095] Step 1: Initialize the decision-making strategy of agent 110. This initialization refers to the initialization of the parameters in the neural network.

[0096] Step 2: Agent 110 acquires environmental state 130;

[0097] Step 3: Based on the input environmental state 130, the intelligent agent 110 uses the decision strategy π to obtain the decision action 140 and informs the environment 120 of the decision action 140.

[0098] Step 4: Environment 120 executes the decision action 140, and the environment state 130 transitions to the next environment state 150, while obtaining the reward 160 corresponding to the decision strategy π.

[0099] Step 5: Agent 110 obtains the reward 160 and the next environment state 150 corresponding to the decision strategy π, and updates the decision strategy based on the input environment state 130, decision action 140, reward 160 corresponding to the decision strategy π and the next environment state 150. The goal of the update is to maximize the reward or minimize the penalty.

[0100] Step 6: If the training termination condition is not met, return to Step 3; if the training termination condition is met, terminate the training.

[0101] It should be understood that the above training steps can be performed online or offline. If performed offline, the data from each iteration (e.g., the input environment state 130, decision action 140, reward 160 corresponding to the decision strategy, and the next environment state 150) are placed in the experience cache for training.

[0102] The training termination condition generally refers to the reward in step five of the agent's training process exceeding a certain preset threshold, or the penalty falling below a certain preset threshold. Alternatively, the number of training iterations can be pre-specified, terminating training after reaching the preset number of iterations. Training termination can also be controlled based on system performance, such as when system performance metrics (e.g., throughput, packet loss rate, latency, fairness, etc. in a communication system) reach a preset threshold.

[0103] Once the agent has completed training, it enters the inference phase and performs the following steps:

[0104] Step 1: The agent acquires the environmental state;

[0105] Step two: The agent uses a decision-making strategy to obtain a decision action based on the input environmental state, and then informs the environment of the decision action.

[0106] Step 3: The environment executes the decision action, and the environmental state transitions to the next environmental state;

[0107] Step four, return to step one.

[0108] As can be seen from the above, a well-trained agent no longer cares about the reward corresponding to the decision; it only needs to make decisions based on its own strategy according to the environmental state.

[0109] In actual use, the training and inference steps of the above-mentioned intelligent agent are carried out alternately. That is, after training for a period of time, inference begins after the training termination condition is reached. After inference for a period of time, the system environment changes, which may make the original trained strategy no longer applicable, and the training process needs to be restarted.

[0110] Combining reinforcement learning and deep learning yields deep reinforcement learning. Deep reinforcement learning still adheres to the framework of agent-environment interaction in reinforcement learning. The difference is that the agent uses deep neural networks for decision-making. The method of training the agent using deep reinforcement learning is also applicable to the technical solutions protected in the embodiments of this application.

[0111] A fully connected neural network, also known as a multilayer perceptron (MLP), consists of an input layer (left side), an output layer (right side), and multiple hidden layers (middle). Each layer contains several nodes, called neurons. Neurons in adjacent layers are connected pairwise, such as... Figure 3 As shown.

[0112] Considering neurons in two adjacent layers, the output h of a neuron in the next layer is the weighted sum of all neurons x connected to it in the previous layer, after passing through an activation function. This can be represented by a matrix as follows:

[0113] h = f(wx + b)

[0114] Where w is the weight matrix, b is the bias vector, and f is the activation function. The output of the neural network can then be recursively expressed as:

[0115] y = f n (w n f n-1 (…)+b n )

[0116] Simply put, a neural network can be understood as a mapping from an input data set to an output data set. Neural networks are typically initialized randomly, and the process of obtaining this mapping using existing data is called training the neural network.

[0117] The specific training method involves evaluating the output of the neural network using a loss function and backpropagating the error. Gradient descent is then used to iteratively optimize w and b until the loss function reaches its minimum. Figure 4 As shown.

[0118] The gradient descent process can be represented as

[0119]

[0120] Where θ is the parameter to be optimized (such as w and b), L is the loss function, and η is the learning rate, which controls the step size of gradient descent.

[0121] The backpropagation process utilizes the chain rule for partial derivatives, meaning the gradient of the parameters in the previous layer can be recursively calculated from the gradient of the parameters in the next layer, such as... Figure 5 As shown, the formula can be expressed as:

[0122]

[0123] Among them, w ij Let s be the weight of the connection between node j and node i. i The weighted sum of the inputs at node i.

[0124] Through reinforcement learning, an agent can continuously improve its parameter configuration by interacting with its environment (i.e., acquiring environmental states, making decisions, receiving rewards for those decisions, and obtaining the next environmental state), thus making increasingly better decisions. Furthermore, due to this environmental interaction and iterative self-improvement mechanism, the agent can track changes in the environment. In contrast, traditional decision-making algorithms do not receive rewards from the environment after making a decision, and therefore cannot improve themselves through interaction. Moreover, when the environmental state changes, the current decision-making algorithm becomes inapplicable, requiring a completely new mathematical model to be built.

[0125] This application proposes an agent training method for optimizing airspace traffic flow in urban air traffic airport terminal areas. This method trains multiple agents using reinforcement learning and then utilizes the trained agents for decision-making. This enables the agents to autonomously select flight routes and adjust flight speeds, thereby optimizing traffic allocation in the terminal area.

[0126] Modern cities offer diverse travel options, including walking, cycling, driving, public transportation, and ride-sharing services. Urban air traffic is extremely congested in the airspace near airports. Under the current air traffic management system, aircraft in the airport terminal area are assigned routes by tower controllers based on their experience. This type of aircraft route allocation and flow management is no longer suitable for urban air transport scenarios with extremely high traffic volumes.

[0127] To improve the efficiency of air traffic management, airport terminal control zones are established near urban air traffic airports. Once an aircraft enters this zone, its flight will be restricted, and the flight of aircraft entering this zone will be handed over to the urban air traffic airport collaborative decision-making system.

[0128] To address the aforementioned shortcomings, this application provides an agent training method and apparatus for optimizing airspace traffic flow in urban air traffic airport terminal areas. By using agent reinforcement learning to make decisions on route selection and speed adjustment, the method enables aircraft to take off smoothly or land precisely, effectively utilizing limited takeoff and landing space, maximizing airspace utilization, and thus optimizing traffic allocation in the terminal area. This method allows the agent to automatically select the optimal approach or takeoff route for the aircraft and adjust the aircraft's speed along the route, increasing airspace capacity in the terminal area and improving air transport efficiency.

[0129] Figure 6 This is a flowchart illustrating an example of a method for determining an aircraft flight strategy applicable to this application. The method 600 can be executed by a computer system, which includes an intelligent agent. The intelligent agent training method mainly includes steps such as: loading, perception, selection, and training, etc. Figure 6 As shown, the method 600 includes the following steps.

[0130] S610, Obtain a first model, which is obtained by training based on first training data. The first training data includes flight status information of a first aircraft in a first time period, flight status information of at least one second aircraft in the first time period, and target flight strategy of the first aircraft in the first time period. The second aircraft is an aircraft located within a first range in the first time period, and the first range is determined based on the position of the first aircraft.

[0131] For example, training the first model based on the first training data includes: inputting the first training data into an original model to obtain a first flight strategy; sending the first flight strategy to a server, where the server stores the target flight strategy; obtaining first reward data from the server, the first reward data being determined based on the relationship between the first flight strategy and the target flight strategy; and adjusting the original model based on the first reward data to determine the first model.

[0132] For example, the first parameter satisfies:

[0133]

[0134] in, This indicates the flight status information of the third aircraft during the second time period, I (0) This indicates the distance and orientation between the current position and the landing position of the third aircraft, the current velocity and acceleration of the third aircraft, and the attitude information of the third aircraft. (i) This represents the distance and orientation between the current position and landing position of the i-th aircraft, the current velocity and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft. (i)LOS(o,i) represents the distance from the third aircraft to the i-th aircraft, LOS(o,i) represents the distance loss between the third aircraft and the i-th aircraft, n represents the number of the fourth aircraft, and i is a positive integer greater than or equal to 1 and less than or equal to n.

[0135] For example, the operational space of this third aircraft satisfies:

[0136]

[0137] in, This indicates the takeoff and landing route selected by the third aircraft, v min This indicates the minimum cruising speed of the third aircraft along the takeoff and landing route, v. t This represents the current speed of the third aircraft, v. max This indicates the maximum cruising speed of the third aircraft on the takeoff and landing route.

[0138] For example, the termination criteria of the third aircraft satisfy:

[0139] N aircraft =0

[0140] Where, N aircraft This indicates the number of aircraft located within the terminal airspace of this second range.

[0141] For example, the first reward data is calculated based on a reward function that satisfies:

[0142]

[0143] in, The distance between the third and fourth aircraft is represented by λ, α, and β, which are positive constants. t This indicates the total time for the takeoff and landing of the third aircraft.

[0144] For example, when a conflict occurs between the third aircraft and the fourth aircraft, reward data corresponding to the conflict is obtained from the server. The conflict is used to indicate that the distance between the third aircraft and the fourth aircraft is less than a second preset threshold, which is the minimum safe distance between the third aircraft and the fourth aircraft.

[0145] It should be noted that the airport terminal area airspace in this application includes multiple takeoff and landing routes, and the starting and ending points of the approach routes for each of these routes are configured with vertical takeoff and landing (VTOL) airspace. This VTOL airspace is a stepped cylindrical airspace, with multiple approach points configured in different directions and multiple approach paths at different angles.

[0146] For example, the method further includes: when the third aircraft enters the stepped cylindrical airspace and causes congestion, or when the stepped cylindrical airspace exceeds a first preset threshold, controlling the target aircraft to execute a waiting procedure, wherein the first preset threshold is the maximum number of aircraft that the stepped cylindrical airspace can accommodate.

[0147] S620, Obtain a first parameter, the first parameter including flight status information of the third aircraft in the second time period and flight status information of at least one fourth aircraft in the second time period, the fourth aircraft being an aircraft located within a second range in the second time period, the second range being determined based on the position of the third aircraft;

[0148] S630, input the first parameter into the first model to obtain the flight strategy of the third aircraft.

[0149] It should be noted that this termination criterion is used to determine whether the agent should proceed with the next learning and training iteration. For example, when the number of aircraft in the airspace is zero, the learning and training process terminates. When the number of aircraft in the airspace is not zero, the learning and training process continues.

[0150] It should be understood that the algorithms described above for learning and training the agent are merely illustrative and should not be construed as limiting this application. Optionally, the learning and training of the agent can be performed using algorithms for actor critics, such as the asynchronous advantage actor-critic (A3C) algorithm, the softactor critic (SAC) algorithm, the proximal policy optimization algorithms (PPO) algorithm, the deep deterministic policy gradient (DDPG) algorithm, etc.

[0151] Figure 7 This is another example of a flowchart illustrating a method for determining an aircraft flight strategy applicable to this application. The method 700 can be executed by a computer system, which includes an intelligent agent, or by a dedicated neural network accelerator, a general-purpose processor, or other devices. The following description of the method 700 using an intelligent agent as the executing entity is exemplary and should not be construed as limiting the executing entity of the method 700. Figure 7 As shown, the method 700 includes the following steps:

[0152] The S710 trains an airborne collaborative decision-making system (agent) in a virtual environment. This agent is loaded with information such as state space, action space, and termination criteria.

[0153] It should be understood that this step is the initialization of the agent, including the policy function, state space, action space, and termination criteria. The target aircraft is any one of at least one aircraft entering the airport terminal area.

[0154] For example, the state space of the intelligent agent includes: the state of the target aircraft. And the state information of n intelligent agents surrounding the target aircraft. These n intelligent agents can be... Figure 1 Any two routers, A through F, can communicate directly with each other.

[0155] It should be noted that the airport terminal area has multiple takeoff and landing routes, with vertical takeoff and landing (VTOL) airspace designated at the beginning and end of the approach routes. Specifically, during takeoff, the intelligent agent coordinates target aircraft to reach different cruising altitudes within this airspace; and during landing, the intelligent agent coordinates target aircraft flying at different altitudes to select takeoff and landing routes. By establishing VTOL airspace, aircraft can take off smoothly or land precisely, achieving optimal utilization of the limited takeoff and landing space.

[0156] Furthermore, a stepped cylindrical airspace is set up within the vertical take-off and landing airspace. The stepped cylindrical airspace has multiple approach points in different directions and multiple approach paths with different angles to maximize airspace utilization.

[0157] It should be noted that the arrangement of approach and departure routes differs from the traditional single approach or takeoff route. The approach and departure routes of urban air traffic airport terminal areas are arranged in a distributed manner at certain angles, with multiple routes at different angles. If congestion occurs or the airport's maximum capacity is exceeded when entering the stepped cylindrical airspace (i.e., one example of the first preset threshold), a waiting procedure is executed. The optimal approach or takeoff route is automatically selected for the aircraft, and the speed of the aircraft on the route is adjusted. While ensuring a safe separation, the terminal area airspace resources are maximized, the airspace capacity within the terminal area is increased, the use of waiting procedures is reduced, and the efficiency of air transport is improved.

[0158] S720, determine the flight status information of the target aircraft in the target route (i.e., an example of the first flight status information).

[0159] It should be understood that the flight status includes the flight status information of the target aircraft, as well as the status information of n aircraft at a distance L from the target aircraft.

[0160] It should be noted that the target aircraft's flight status... The expression can be:

[0161]

[0162] Among them, I (0) This indicates the distance and azimuth between the current position of the target aircraft and the target position, the current velocity and acceleration of the target aircraft, and the attitude information of the target aircraft. (i) This represents the distance and orientation between the current position of the i-th aircraft and the target position, the current velocity and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft. (i) LOS(o,i) represents the distance from the target aircraft to the i-th aircraft, and LOS(o,i) represents the distance loss between the target aircraft and the i-th aircraft, where i is a positive integer greater than or equal to 1 and less than or equal to n.

[0163] It should be understood that the distance between the target aircraft and other aircraft should be less than a preset threshold, meaning that multiple aircraft should operate within a safe distance to maximize airspace utilization. If the distance between multiple aircraft is too large, it will be detrimental to improving air transport efficiency and may result in a loss, namely the loss of distance LOS(o,i) between the target aircraft o and other aircraft i.

[0164] In the S730, the intelligent agent controls the target aircraft to select actions according to the operational space and outputs the actions to the environment.

[0165] For example, the action space of this intelligent agent mainly includes route and speed, and can be defined as A. t Its expression is:

[0166]

[0167] in, Indicates the flight path i, v selected by target aircraft o. min v is the minimum permissible cruise speed (deceleration) for the target aircraft. t v represents the current (maintained) speed of the target aircraft. max This represents the maximum permissible cruising speed (acceleration) for the target aircraft. By introducing machine intelligence to eliminate the impact of human factors and errors on flight safety, the incidence of flight malfunctions, accidents, and disasters can be significantly reduced. Furthermore, flights using intelligent decision-making systems can autonomously select routes and adjust speeds in real time, handling extremely high flight density.

[0168] It should be noted that if a conflict occurs between two agents, both agents will be penalized. The remaining n-2 agents, since no conflict occurs, will not be penalized. The reward function for the target aircraft is r(t), and its expression is as follows:

[0169]

[0170] in, λ is the distance from the target aircraft to the nearest surrounding aircraft. λ can be a large constant used to penalize the target aircraft for collisions with other aircraft. α and β are small positive constants used to penalize losses due to takeoff and landing time and proximity collisions of the target aircraft. t It is the total cumulative time spent on takeoff and landing of the aircraft.

[0171] Therefore, the above-mentioned conflict means that the flight distance between any two aircraft is less than min (i.e., an example of the second preset threshold), where min is the minimum value set for the safety interval.

[0172] For example, the current regulations stipulate that if two aircraft are at the same altitude, the danger distance is 5 nautical miles (n mile); if the two aircraft are not at the same altitude, the danger distance is 3 nautical miles (n mile). This large distance is mainly because the exhaust of jet aircraft will generate turbulence, which can easily affect passing aircraft and cause the fuselage to shake violently.

[0173] In S740, the agent continuously learns and tries to obtain the maximum reward until it reaches a termination state, which is when the number of aircraft in the airspace is zero.

[0174] It should be noted that the agent obtains second flight state information and reward data from the environment in response to the decision action output in step S730.

[0175] For example, the number of aircraft in this airspace can be defined as N. aircraft When N aircraft When N = 0, the learning and training process terminates. It should be understood that the termination criterion of this agent satisfies N. aircraft =0.

[0176] After the S750 agent learns its decision-making capabilities, the trained collaborative decision-making system (agent) needs to be installed on the aircraft.

[0177] It should be understood that, based on the completion of the training and loading of the aforementioned intelligent agent, the intelligent agent is able to autonomously select flight routes and adjust speeds. That is, when the target aircraft enters the airport terminal control area designated near the urban air traffic airport, its flight route and speed are decided by the onboard collaborative decision-making system (intelligent agent).

[0178] In this embodiment of the application, the process of training the agent can be carried out by the "advantage actor critic" (A2C) algorithm.

[0179] This algorithm uses a merit function instead of the original reward in the critic network, which can serve as a metric for evaluating the quality of the selected action value and the average value of all actions.

[0180] The dominance function is:

[0181] A π (s,a)=Q π (s,a)-V π (s)

[0182] State value function V(s): represents the sum of the action value function corresponding to all possible actions in this state multiplied by the probability of taking that action.

[0183] Action value function Q(s,a): represents the value function corresponding to action a in this state.

[0184] Advantage function Q π (s,a)-V π (s): Indicates the advantage of the action value function over the current state value function.

[0185] It should be understood that if the advantage function is greater than zero, it means that the action is better than the average action; if the advantage function is less than zero, it means that the current action is not as good as the average action.

[0186] The specific implementation steps for training the A2C algorithm are as follows:

[0187] Step 1: Instantiate actor / critic and initialize hyperparameters;

[0188] Step 2: for Epochs:

[0189] for Steps:

[0190] (1) Using actor networks from action space A t (Route, speed) Select actions for the target aircraft requesting takeoff or landing;

[0191] (2) Step(state, action) receives a reward and transitions to a new state;

[0192] (3) Use the Q_Net of the critic network to calculate the value V(s) of the current state and the value V(s′) of the next state, and obtain the error: TD_error=r+γV(s′)-V(s);

[0193] Q_Network is trained using the mean squared error of TD_error;

[0194] (4) TD_error is fed back to the Actor, and the Policy Gradient formula is used to train the Actor;

[0195] (5) The current state is updated to the next state, state = next_state, which corresponds to Figure 1 In the middle, s(t) ← s(t+1).

[0196] Step 3: After the training and learning process is completed, the network parameters obtained from the training record the state-behavior information. By using the network parameters to input the corresponding state information, the aircraft can make intelligent decisions. Through extensive training and learning in a simulation environment, the intelligent agent can acquire the ability to make optimal autonomous choices, thereby optimizing the allocation of traffic in the terminal area.

[0197] It should be understood that the algorithms described above for learning and training the agent are merely illustrative and should not be construed as limiting this application. Optionally, the learning and training of the agent can be performed using algorithms for actor critics, such as the asynchronous advantage actor-critic (A3C) algorithm, the softactor critic (SAC) algorithm, the proximal policy optimization algorithms (PPO) algorithm, the deep deterministic policy gradient (DDPG) algorithm, etc.

[0198] In summary, this application provides an agent training method for optimizing air traffic flow, including: during takeoff, the agent coordinates target aircraft to enter different cruising altitudes within the terminal area airspace; during landing, the agent coordinates target aircraft flying at different altitudes to select takeoff and landing routes. By setting up vertical takeoff and landing airspace, target aircraft can take off smoothly or land accurately, achieving optimal utilization of limited takeoff and landing sites.

[0199] Furthermore, this application maximizes airspace utilization by setting multiple approach paths at different angles. The arrangement of approach and departure paths differs from traditional single approach or takeoff paths. In urban air traffic airport terminal areas, approach and departure paths are distributed at certain angles, inherently offering multiple paths at different angles. The trained agent can automatically select the optimal approach or takeoff route for the aircraft and adjust its speed along the route, maximizing the use of terminal area airspace resources while ensuring safe separation, reducing the use of waiting procedures, and improving air transport efficiency.

[0200] Furthermore, this application eliminates the impact of human factors and errors on flight safety by introducing machine intelligence, significantly reducing flight malfunctions, accidents, and disasters. Simultaneously, flights using the intelligent decision-making system can autonomously select routes and adjust speeds in real time, handling extremely high flight density. The intelligent agent, through extensive training in a simulated environment, acquires autonomous selection capabilities, using multi-agent reinforcement learning to make decisions on route selection and speed adjustment, thereby optimizing traffic distribution in the terminal area.

[0201] The above text combined Figure 6 and Figure 7 This application introduces an agent training method for air traffic flow optimization based on embodiments of the present application. The following will combine... Figure 8 and Figure 9 This application introduces an intelligent agent training device for optimizing air traffic flow, based on embodiments of the present application.

[0202] Figure 8 This is a schematic block diagram of an aircraft flight strategy determination device 800 according to an embodiment of this application. The device 800 can be used to execute the agent training method provided in the above embodiments; for simplicity, it will not be described in detail here. The device 800 can be a computer system, a chip or circuit within a computer system, or it can also be called an AI module. Figure 8 As shown, the device 800 includes:

[0203] The acquisition unit 810 is used to acquire a first model, which is obtained by training based on first training data. The first training data includes flight status information of a first aircraft in a first time period, flight status information of at least one second aircraft in the first time period, and target flight strategy of the first aircraft in the first time period. The second aircraft is an aircraft located within a first range in the first time period, and the first range is determined according to the position of the first aircraft.

[0204] The acquisition unit 810 is also used to acquire a first parameter, which includes flight status information of the third aircraft in the second time period and flight status information of at least one fourth aircraft in the second time period. The fourth aircraft is an aircraft located within a second range in the second time period, and the second range is determined based on the position of the third aircraft.

[0205] The transceiver unit 820 is used to input the first parameter into the first model to obtain the flight strategy of the third aircraft.

[0206] It should be understood that the specific manner in which the device 800 executes the agent training method and the beneficial effects produced can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0207] Figure 9 This is a schematic block diagram of an aircraft flight strategy determination device 900 according to an embodiment of this application. The device 900 can be used to execute the reinforcement learning method provided in the above embodiments, which will not be described in detail here for simplicity. The device 900 includes a processor 910, a memory 920, and a transceiver 930, which communicate with each other through internal interconnection paths to transmit control and / or data signals. In one possible design, the processor 910, transceiver 930, and memory 920 can be implemented using a chip. The processor 910 is coupled to the memory 920, which stores computer programs or instructions. The processor 910 executes the computer programs or instructions stored in the memory 920, causing the methods in the above method embodiments to be executed.

[0208] For example, processor 910 is used to acquire a first model, which is obtained by training based on first training data. The first training data includes flight status information of a first aircraft in a first time period, flight status information of at least one second aircraft in the first time period, and target flight strategy of the first aircraft in the first time period. The second aircraft is an aircraft located within a first range in the first time period, and the first range is determined according to the position of the first aircraft.

[0209] The processor 910 is also configured to acquire a first parameter, the first parameter including flight status information of the third aircraft during the second time period and flight status information of at least one fourth aircraft during the second time period, the fourth aircraft being an aircraft located within a second range during the second time period, the second range being determined based on the position of the third aircraft;

[0210] The transceiver 930 is used to input the first parameter into the first model to obtain the flight strategy of the third aircraft.

[0211] It should be understood that the specific manner in which the device 900 executes the agent training method and the beneficial effects produced can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0212] This application also provides a computer-readable storage medium storing a computer program or code that, when run on a computer, enables the computer to implement the methods described in the above embodiments.

[0213] This application also provides a computer program product, which includes computer program code, which, when executed by a computer, enables the method described in the above embodiments.

[0214] This application also provides a chip including at least one processor coupled to a memory for storing a computer program, and the processor for calling and running the computer program from the memory, so that a communication device equipped with the chip system performs the methods described above.

[0215] The chip may include input circuitry or interface for transmitting information and / or data, and output circuitry or interface for receiving information and / or data.

[0216] It should be understood that in the embodiments of this application, the processor can be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0217] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0218] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0219] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0220] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0221] It should also be understood that the terms "first" and "second" mentioned in this document are used only to distinguish the technical solutions of this application more clearly, and should not constitute any limitation on this application.

[0222] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0223] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0224] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0225] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0226] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0227] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0228] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0229] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. The scope of protection of this application shall be determined by the scope of the claims.

Claims

1. A method for determining an aircraft's flight strategy, characterized in that, include: A first model is obtained, which is trained based on first training data. The first training data includes flight status information of a first aircraft in a first time period, flight status information of at least one second aircraft in the first time period, and target flight strategy of the first aircraft in the first time period. The second aircraft is an aircraft located within a first range in the first time period, and the first range is determined based on the position of the first aircraft. Obtain a first parameter, which includes flight status information of a third aircraft during a second time period and flight status information of at least one fourth aircraft during the second time period. The fourth aircraft is an aircraft located within a second range during the second time period, and the second range is determined based on the position of the third aircraft. The first parameter is input into the first model to obtain the flight strategy of the third aircraft.

2. The method according to claim 1, characterized in that, The method further includes: The first model is trained based on the first training data.

3. The method according to claim 2, characterized in that, Training the first model based on the first training data includes: Input the first training data into the original model to obtain the first flight strategy; The first flight strategy is sent to the server, which stores the target flight strategy. The first reward data is obtained from the server, and the first reward data is determined based on the relationship between the first flight strategy and the target flight strategy; The original model is adjusted based on the first reward data to determine the first model.

4. The method according to any one of claims 1 to 3, characterized in that, The flight status information of the third aircraft during the second time period is associated with the following: the distance and orientation between the current position and the landing position of the third aircraft, the current speed and acceleration of the third aircraft, and the attitude information of the third aircraft; The distance and orientation between the current position and landing position of the i-th aircraft, the current speed and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft; the distance from the third aircraft to the i-th aircraft; the distance loss between the third aircraft and the i-th aircraft; The number n of the fourth aircraft, i is a positive integer greater than or equal to 1 and less than or equal to n.

5. The method according to any one of claims 1 to 3, characterized in that, The operational space of the third aircraft is associated with the following: the takeoff and landing route selected by the third aircraft, the minimum cruise speed of the third aircraft on the takeoff and landing route, the current speed of the third aircraft, and the maximum cruise speed of the third aircraft on the takeoff and landing route.

6. The method according to any one of claims 1 to 3, characterized in that, The termination criteria of the third aircraft satisfy: in, This indicates the number of aircraft located within the terminal area airspace of the second range.

7. The method according to claim 6, characterized in that, The terminal area airspace includes multiple take-off and landing routes, and the starting and ending points of the approach routes of each of the multiple take-off and landing routes are configured with vertical take-off and landing airspace.

8. The method according to claim 7, characterized in that, The vertical takeoff and landing airspace is a stepped cylindrical airspace, with multiple approach points configured in different directions and multiple approach paths at different angles. The method further includes: when the third aircraft enters the stepped cylindrical airspace and causes congestion, or when the stepped cylindrical airspace exceeds a first preset threshold, controlling the third aircraft to execute a waiting procedure, wherein the first preset threshold is the maximum number of aircraft that the stepped cylindrical airspace can accommodate.

9. The method according to claim 3, characterized in that, The first reward data is calculated based on a reward function, which satisfies: in, This indicates the distance between the third and fourth aircraft. , and It is a positive number. This indicates the total time for the takeoff and landing process of the third aircraft.

10. The method according to any one of claims 1 to 3, characterized in that, The method further includes: When a conflict occurs between the third aircraft and the fourth aircraft, reward data corresponding to the conflict is obtained from the server. The conflict is used to indicate that the distance between the third aircraft and the fourth aircraft is less than a second preset threshold, which is the minimum safe distance between the third aircraft and the fourth aircraft.

11. A device for determining the flight strategy of an aircraft, characterized in that: The acquisition unit is used to acquire a first model, which is obtained by training based on first training data. The first training data includes flight status information of a first aircraft in a first time period, flight status information of at least one second aircraft in the first time period, and target flight strategy of the first aircraft in the first time period. The second aircraft is an aircraft located within a first range in the first time period, and the first range is determined according to the position of the first aircraft. The acquisition unit is further configured to acquire a first parameter, the first parameter including flight status information of the third aircraft in the second time period and flight status information of at least one fourth aircraft in the second time period, wherein the fourth aircraft is an aircraft located within a second range in the second time period, and the second range is determined based on the position of the third aircraft; The transceiver unit is used to input the first parameter into the first model to obtain the flight strategy of the third aircraft.

12. The apparatus according to claim 11, characterized in that, The device further includes: The processing unit is used to train the first model based on the first training data.

13. The apparatus according to claim 12, characterized in that, The transceiver unit is further configured to: Input the first training data into the original model to obtain the first flight strategy; The first flight strategy is sent to the server, which stores the target flight strategy. The acquisition unit further includes acquiring first reward data from the server, wherein the first reward data is determined based on the relationship between the first flight strategy and the target flight strategy; The processing unit further includes adjusting the original model based on the first reward data to determine the first model.

14. The apparatus according to any one of claims 11 to 13, characterized in that, The flight status information of the third aircraft during the second time period is associated with the following: the distance and orientation between the current position and the landing position of the third aircraft, the current speed and acceleration of the third aircraft, and the attitude information of the third aircraft; The distance and orientation between the current position and landing position of the i-th aircraft, the current speed and acceleration of the i-th aircraft, and the attitude information of the i-th aircraft; the distance from the third aircraft to the i-th aircraft; the distance loss between the third aircraft and the i-th aircraft; The number n of the fourth aircraft, i is a positive integer greater than or equal to 1 and less than or equal to n.

15. The apparatus according to any one of claims 11 to 13, characterized in that, The operational space of the third aircraft is associated with the following: the takeoff and landing route selected by the third aircraft, the minimum cruise speed of the third aircraft on the takeoff and landing route, the current speed of the third aircraft, and the maximum cruise speed of the third aircraft on the takeoff and landing route.

16. The apparatus according to any one of claims 11 to 13, characterized in that, The termination criteria of the third aircraft satisfy: in, This indicates the number of aircraft located within the terminal area airspace of the second range.

17. The apparatus according to claim 16, characterized in that, The terminal area airspace includes multiple take-off and landing routes, and the starting and ending points of the approach routes of each of the multiple take-off and landing routes are configured with vertical take-off and landing airspace.

18. The apparatus according to claim 17, characterized in that, The vertical takeoff and landing airspace is a stepped cylindrical airspace, with multiple approach points configured in different directions and multiple approach paths at different angles. The processing unit is further configured to control the third aircraft to execute a waiting procedure when the third aircraft enters the stepped cylindrical airspace and causes congestion, or when the stepped cylindrical airspace exceeds a first preset threshold. The first preset threshold is the maximum number of aircraft that the stepped cylindrical airspace can accommodate.

19. The apparatus according to claim 13, characterized in that, The first reward data is calculated based on a reward function, which satisfies: in, This indicates the distance between the third and fourth aircraft. , and It is a positive number. This indicates the total time for the takeoff and landing process of the third aircraft.

20. The apparatus according to any one of claims 11 to 13, characterized in that, The processing unit is further configured to obtain reward data corresponding to the conflict from the server when a conflict occurs between the third aircraft and the fourth aircraft. The conflict is used to indicate that the distance between the third aircraft and the fourth aircraft is less than a second preset threshold, which is the minimum safe distance between the third aircraft and the fourth aircraft.

21. A communication device, characterized in that, include: Memory, used to store executable instructions; A processor for invoking and running the executable instructions in the memory to perform the method as described in any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed by a processor, perform the method as described in any one of claims 1 to 10.

23. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, performs the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle action decision-making method and device based on reinforcement learning

    CN111708355A