Aircraft route planning optimization method based on conditional diffusion model

By building a reinforced learning environment and conditional diffusion model for free route airspace track planning, aircraft track planning is optimized, and the problems of high energy consumption and low safety in free route airspace are solved, and trajectory planning with higher safety and lower energy consumption is achieved.

CN120333471APending Publication Date: 2025-07-18BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510466459.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art lacks effective optimization of aircraft flight costs and safety in free route airspace, resulting in high energy consumption and low safety.

Method used

Build a reinforced learning environment for free route airspace track planning, use conditional diffusion models, optimize aircraft track planning through the design of reward functions and state space, including training of diffusion sub-models and conditional function sub-models to generate tracks with higher safety and lower energy consumption.

Benefits of technology

It realizes safe and reliable track planning in free route airspace, reduces fuel consumption, improves flight safety and intelligence level, and is suitable for track optimization tasks of multiple sets of starting and ending points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120333471A_ABST
    Figure CN120333471A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of route planning, and discloses an aircraft route planning optimization method based on a conditional diffusion model. According to the method, a high-efficiency conditional diffusion model is established on the basis of a free airway airspace flight path planned by an A * algorithm and with the goals of safety and energy consumption reduction, so that optimization of free airway airspace flight path planning is realized. Comprising the following steps: (1) constructing a free airway airspace flight path planning reinforcement learning environment; (2) generating an expert trajectory data set and a trajectory reward value thereof; (3) establishing a conditional diffusion model and completing training, wherein the conditional diffusion model comprises a diffusion sub-model and a conditional function sub-model; and (4) providing an action decision by adopting a conditional diffusion model, and completing route planning optimization. Compared with the prior art, the method has the advantages that the airspace flight path of the free airway can be effectively optimized, and powerful support is provided for airspace operation of the free airway.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of route planning, and particularly relates to an optimization method for aircraft trajectory planning based on a conditional diffusion model. Background Art

[0002] With the increasing congestion of air traffic, airlines and passengers will encounter a series of problems, such as longer flight delays, higher operating costs, more environmental impacts, lower time predictability, and more severe safety challenges. To solve these problems, many aviation authorities are considering deviating from the traditional fixed-route operation mode. For example, the Single European Sky ATM Research (SESAR) project introduced the concept of Free Route Airspace (FRA) in Europe. In the free route airspace, users can freely choose the path between the entry and exit. Compared with the fixed-route operation mode, the free route airspace operation mode improves flexibility and reduces energy consumption. Since there are no fixed routes in the free route airspace, users need to achieve autonomous planning of flight paths.

[0003] However, current research on free route airspace trajectory planning often focuses on shorter paths, which may ignore flight costs related to the aircraft's own performance (such as fuel consumption costs, safety costs, etc.). There is an urgent need to propose a new planning method to enable aircraft to optimize their original planned trajectories in the free route airspace to improve flight safety and energy efficiency. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the present invention proposes an optimization method for aircraft trajectory planning based on a conditional diffusion model. The optimization method constructs a reinforcement learning environment for free route airspace trajectory planning, and uses the conditional diffusion model to optimize the planned free route airspace trajectory to make its energy consumption lower and safety higher.

[0005] The technical solution of the present invention is specifically as follows:

[0006] An optimization method for aircraft trajectory planning based on a conditional diffusion model, comprising the following steps:

[0007] Step S1: Construct a reinforcement learning environment for free route airspace trajectory planning, including a state space S, an action space A, and a reward function r, where the reward function r includes a penalty for flying into unavailable airspace, a penalty for the distance to the end point, a reward for reaching the end point, a penalty for fuel consumption, and a penalty for safety costs;

[0008] Step S2: Generate an expert trajectory dataset and its trajectory reward values;

[0009] Step S3: Establish a conditional diffusion model and complete training, where the conditional diffusion model includes a diffusion sub-model and a conditional function sub-model;

[0010] Step S4: Use a conditional diffusion model to provide action decisions and complete the optimization of trajectory planning.

[0011] Preferably, the state space S in the step S1 is:

[0012] S = [x now , y now , heading, x start , y start , x goal , y goal

[0013] Wherein, x now and y now are respectively the x-axis coordinate and y-axis coordinate of the current position of the aircraft, heading is the current heading angle of the aircraft, x start and y start are respectively the x-axis coordinate and y-axis coordinate of the starting position of the current trajectory planning task of the aircraft, x goal and y goal are respectively the x-axis coordinate and y-axis coordinate of the ending position of the current trajectory planning task of the aircraft.

[0014] Preferably, the action space A in the step S1 is:

[0015] A = [v x , v y

[0016] Wherein, v x and v y are respectively the x-axis component and y-axis component of the current horizontal speed of the aircraft.

[0017] Preferably, the reward function r in the step S1 is:

[0018] r = r invasion + r goal_distance + r goal + r fuel + r turn_rate

[0019] Wherein, r invasion is the penalty for flying into unavailable airspace, r goal_distance is the penalty for the distance to the end point, r goal is the reward for reaching the end point, r fuel is the penalty for fuel consumption, r turn_rate is the penalty for safety cost.

[0020] Preferably, the expression of r invasion is:

[0021] ​​

[0022] Preferably, the expression of r goal_distance is as follows:

[0023] r goal_distance = -||position now -positon goal || * 0.02

[0024] where position now and positon goal are the current position and the end position of the aircraft respectively, and ||·|| is the Euclidean norm.

[0025] Preferably, the expression of r goal is as follows:

[0026]

[0027] Preferably, the expression of r fuel is as follows:

[0028] r fuel = -f cr * dt * 0.1

[0029] where f cr is the cruise fuel flow rate of the aircraft, and dt is the planned step time interval.

[0030] Preferably, the expression of r turn_rate is as follows:

[0031]

[0032] Preferably, step S2 specifically includes:

[0033] Taking the entrance of the free route airspace as the starting point, the exit as the end point, and the unavailable airspace as obstacles, using the A* algorithm to generate an expert trajectory dataset, and interacting with the free route airspace trajectory planning reinforcement learning environment to obtain the corresponding trajectory reward value.

[0034] Preferably, step S3 specifically includes:

[0035] Step S3-1: Establish a diffusion sub-model, and train the diffusion sub-model using the trajectory data in the expert trajectory dataset;

[0036] Step S3-2: Establish a conditional function sub-model, and train the conditional function sub-model using the trajectory data and its trajectory reward value in the expert trajectory dataset.

[0037] Preferably, the diffusion sub-model in the step S3-1 adopts a U-Net network structure, which consists of three encoders and three decoders. Each encoder or decoder is composed of two temporal residual blocks. In each temporal residual block, the input data processed by the convolutional module is added to the temporal embedding processed by the temporal embedding fully connected layer, and then quadratic convolution is performed through the convolutional module. The obtained data is added to the input data to form a residual connection. The temporal embedding fully connected layer contains a Mish activation function and a linear layer.

[0038] Preferably, the trajectory data τ in the step S3-1 is expressed as:

[0039]

[0040] where s1, s2, and s T are the states of the aircraft at the 1st planning step length, the 2nd planning step length, and the Tth planning step length respectively, and a1, a2, and a T are the actions of the aircraft at the 1st planning step length, the 2nd planning step length, and the Tth planning step length respectively.

[0041] The loss function used for training is:

[0042]

[0043] where θ is the parameter of the diffusion sub-model, i is the diffusion step number, ε is the real noise, and ε θ (τ i , i) is the predicted noise, τ 0 is the initial trajectory without added noise, is the joint mathematical expectation of i, ε, and τ 0 .

[0044] Preferably, the conditional function sub-model in the step S3-2 consists of three encoders, two temporal residual blocks, and a fully connected layer. Each encoder is composed of two temporal residual blocks. In each temporal residual block, the input data processed by the convolutional module is added to the temporal embedding processed by the temporal embedding fully connected layer, and then quadratic convolution is performed through the convolutional module. The obtained data is added to the input data to form a residual connection. The temporal embedding fully connected layer contains a sine position embedding layer, two linear layers, and a Mish activation function.

[0045] Preferably, the conditional function in the step S3-2 is expressed as:

[0046]

[0047] where J(s0, a 0:T ) is the total trajectory reward value, T is the total planning step length, and r(st , a t ) is the single-step reward value at the t-th planning step length. is a series of actions corresponding to the maximum total trajectory reward value. s0 is the initial state value, a 0:T is a series of actions from the 0-T planning step lengths;

[0048] The loss function used in training is:

[0049]

[0050] Among them, φ is the parameter of the conditional function sub-model, J is the true trajectory reward value, J φ (τ i , i) is the predicted trajectory reward value, is the joint mathematical expectation of i, J, τ 0 .

[0051] Preferably, the step S4 specifically includes:

[0052] Step S4-1: Input: diffusion sub-model parameter θ, conditional function sub-model parameter φ, optimization rate α, total diffusion steps N, covariance ∑ i ;

[0053] Step S4-2: Repeat steps S4-3, S4-4, S4-5, S4-9 until the aircraft reaches the exit;

[0054] Step S4-3: Observe the current state of the aircraft;

[0055] Step S4-4: Initialize a pure noise trajectory τ that follows a standard Gaussian distribution N ∼ Ν(0, I);

[0056] Step S4-5: For the diffusion steps i = N,..., 1, loop and execute steps S4-6, S4-7, S4-8;

[0057] Step S4-6: Obtain the mean μ of the inverse diffusion according to the diffusion sub-model parameter θ and the trajectory τ at the current diffusion step i ;

[0058] Step S4-7: Combine the gradient of the conditional function sub-model parameter φ to optimize the mean μ of the inverse diffusion to obtain a more optimized trajectory at the previous diffusion step, that is, τ i-1 ∼ Ν(μ + α∑▽J φ (μ), ∑ i );

[0059] Step S4-8: Set the state at the starting planning step of the trajectory at the previous diffusion step as the current state of the aircraft;

[0060] Step S4-9: The aircraft performs the action of the initial planning step length of the clear trajectory

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0062] 1. An aircraft trajectory planning optimization method based on a conditional diffusion model proposed by the present invention optimizes the free airway airspace trajectory by using the conditional diffusion model, aiming to utilize the combinability of the reward function, the flexibility of the state space, and the safety of offline learning to optimize the free airway airspace trajectory, enabling the aircraft to fly safely and reliably in the free airway airspace, and providing strong technical support for the operation of the free airway airspace.

[0063] 2. An aircraft trajectory planning optimization method based on a conditional diffusion model proposed by the present invention generates an optimized trajectory based on the actions of the aircraft itself and is applicable to the free airway airspace that does not rely on fixed waypoints. It faces the multi-objective trajectory optimization task of the starting and ending points determined by multiple groups of entrances and exits, and adds the position information of the starting and ending points to the state space, enabling the model to learn the specific information of the task.

[0064] 3. An aircraft trajectory planning optimization method based on a conditional diffusion model proposed by the present invention uses the conditional diffusion model to learn the historical trajectory information, does not require the aircraft to explore in the environment, and improves the safety of the learning process. By using the conditional diffusion model to optimize the free airway airspace trajectory planning, it can provide a safer and more energy-efficient trajectory, significantly improving the intelligent level of the free airway airspace operation.

[0065] 4. An aircraft trajectory planning optimization method based on a conditional diffusion model proposed by the present invention adds penalties for flying into unavailable airspace, end-point distance penalty, arrival at end-point reward, fuel consumption penalty, and safety cost penalty to the reward function to achieve multi-objective trajectory optimization such as flying around unavailable airspace, enhancing safety, and reducing fuel consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention can be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0067] Figure 1 It is a flowchart of an aircraft trajectory planning optimization method based on a conditional diffusion model of the present invention.

[0068] Figure 2 It is a schematic diagram of the network structure of the diffusion sub-model of the present invention.

[0069] Figure 3 It is a schematic diagram of the structure of the time residual block of the present invention.

[0070] Figure 4 It is a schematic diagram of the structure of the time embedding fully connected layer in the diffusion sub-model of the present invention.

[0071] Figure 5 It is a schematic diagram of the structure of the convolution module of the present invention.

[0072] Figure 6 It is a schematic diagram of the network structure of the conditional function sub-model of the present invention.

[0073] Figure 7 It is a schematic diagram of the structure of the time embedding fully connected layer in the conditional function sub-model of the present invention.

[0074] Figure 8 It is a schematic diagram of the structure of the last connection layer of the conditional function sub-model of the present invention. Detailed implementation manners

[0075] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0076] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0077] The present invention proposes an aircraft flight path planning optimization method based on a conditional diffusion model. The optimization method constructs a free flight path airspace flight path planning reinforcement learning environment according to the simulated free flight path airspace environment, and uses the conditional diffusion model to optimize the planned free flight path airspace flight path to make its energy consumption lower and safety higher.

[0078] Specifically, first, according to the free route airspace information, the state space, action space, reward function, etc. are added to construct a reinforcement learning environment for free route airspace trajectory planning. Then, the A* algorithm is used to generate expert trajectories in the original free route airspace environment, and the expert trajectory dataset formed by them is interacted with the reinforcement learning environment to obtain trajectory reward values. Then, the expert trajectory dataset is used to train the diffusion sub-model, and together with the trajectory reward values, it is used to train the conditional function sub-model. The diffusion sub-model is used to fit the trajectory distribution characteristics, and the conditional function sub-model is used to fit the trajectory reward characteristics. Finally, guided sampling is performed by combining the trained diffusion sub-model and conditional function sub-model, and the conditional function is used to guide the aircraft to generate action decisions with higher reward values.

[0079] Embodiment 1

[0080] Set the following conditions:

[0081] (1) Do not consider severe weather such as thunderstorms in the free route airspace.

[0082] (2) Do not consider the impact of wind on the operating performance of the aircraft.

[0083] (3) Do not consider the effective time of the unavailable airspace and assume that the unavailable airspace always exists.

[0084] (4) Since the upper limit of the unavailable airspace is large, it is assumed that the aircraft can only bypass the unavailable airspace in the horizontal dimension. Therefore, for the trajectory optimization during the cruise phase of the aircraft, the change in the aircraft altitude is not considered.

[0085] (5) Do not consider the change in the aircraft mass due to fuel consumption during the flight.

[0086] As Figure 1 shown, the specific implementation details of the optimization method for free route airspace trajectory planning are as follows:

[0087] I. Step 1: Construct a reinforcement learning environment for free route airspace trajectory planning

[0088] Set up a simulated free route airspace environment to achieve the optimization of aircraft trajectory planning. Considering that the main influencing factor during the operation of the aircraft in the free route airspace environment is the unavailable airspace, and the most basic requirement is also to avoid the unavailable airspace, the unavailable airspace information is used as the main information for the free route airspace environment.

[0089] The information related to the free route airspace environment and reinforcement learning is constructed as follows:

[0090] (1) State space S: The starting and ending points of the aircraft trajectory optimization are determined by multiple pairs of entrances and exits of free airway airspace. Therefore, the state space should not only include the current position and attitude information of the aircraft itself, but also the position information of the starting and ending points, so that the model can learn the specific information of the aircraft's current trajectory planning task. The state space is specifically represented as follows:

[0091] S = [x now , y now , heading, x start , y start , x goal , y goal

[0092] Among them, x now and y now are the x-axis coordinate and y-axis coordinate of the current position of the aircraft respectively, heading is the current heading angle of the aircraft, x start and y start are the x-axis coordinate and y-axis coordinate of the starting point position of the aircraft's current trajectory planning task respectively, x goal and y goal are the x-axis coordinate and y-axis coordinate of the ending point position of the aircraft's current trajectory planning task respectively.

[0093] (2) Action space A: In order to enable the model to better fit the trajectory and trajectory reward value information, in this embodiment, simple two-dimensional velocity is used as the action space. The action space is specifically represented as follows:

[0094] A = [v x , v y

[0095] Among them, v x and v y are the x-axis component and y-axis component of the current horizontal velocity of the aircraft respectively.

[0096] (3) Reward function r: This embodiment needs to achieve multi-objective trajectory optimization such as flying around unavailable airspace, improving safety, and reducing fuel consumption. Therefore, the reward function needs to include the above-mentioned multi-faceted information:

[0097] a. Penalty for flying into unavailable airspace: The primary requirement for the aircraft to operate in the free airway airspace environment is to fly around unavailable airspace. The penalty for flying into unavailable airspace is specifically represented as follows:

[0098]

[0099] Among them, r invasion is the reward value of the aircraft for the penalty of flying into unavailable airspace.

[0100] ​​b. End - point distance penalty: The fundamental requirement for an aircraft to complete a mission is to fly from the starting point to the end - point. The end - point distance penalty is used to guide the aircraft to fly to the end - point. The end - point distance penalty is specifically expressed as follows:

[0101] r goal_distance = - ||position now - positon goal || * 0.02

[0102] Among them, r goal_distance is the reward value of the aircraft regarding the end - point distance penalty, position now and positon goal are the current position and the end - point position of the aircraft respectively, and ||·|| is the Euclidean norm.

[0103] c. End - point arrival reward: To achieve the same goal as in b, an end - point arrival reward is added. The end - point arrival reward is specifically expressed as follows:

[0104]

[0105] Among them, r goal is the reward value of the aircraft regarding the end - point arrival reward.

[0106] d. Fuel - consumption penalty: In this embodiment, the trajectory optimization needs to be based on the goal of reducing fuel consumption, and the fuel - consumption penalty is used to reduce the fuel consumption of the aircraft. A fuel - consumption model is constructed for a specific aircraft model based on the BADA dataset publicly disclosed by EUROCONTROL. The fuel - consumption penalty is specifically expressed as follows:

[0107] r fuel = - f cr * dt * 0.1

[0108] Among them, r fuel is the reward value regarding the fuel - consumption penalty, f cr is the cruise fuel flow rate of the aircraft, and dt is the planned time - step interval.

[0109] The calculation method of the cruise fuel flow rate f cr is as follows:

[0110] f cr =η * Thr * C fcr

[0111] Among them, η is the thrust - specific fuel consumption of the aircraft, Thr is the thrust of the aircraft, and C fcr is the cruise fuel - flow correction coefficient (this coefficient can be queried in the BADA dataset).

[0112] Taking a jet aircraft as an example, the calculation method of the thrust - specific fuel consumption η is as follows:

[0113]

[0114] Among them, C f1 is the first specific fuel consumption coefficient of thrust (this coefficient can be queried in the BADA dataset), C f2 is the second specific fuel consumption coefficient of thrust (this coefficient can be queried in the BADA dataset), and v is the true airspeed of the aircraft.

[0115] The thrust Thr of the aircraft is calculated as follows:

[0116]

[0117] Among them, m is the mass of the aircraft, g is the acceleration due to gravity, is the rate of climb or descent, is the acceleration, and D is the drag of the aircraft.

[0118] In this embodiment, since the change in aircraft altitude is not considered, it can be simplified to:

[0119]

[0120] The drag D of the aircraft is calculated as follows:

[0121]

[0122] Among them, C D is the drag coefficient, S w is the wing area (this coefficient can be queried in the BADA dataset), and ρ is the air density at the altitude of the aircraft (this coefficient can be queried in the BADA dataset).

[0123] For an aircraft in the cruise phase, the drag coefficient C D is calculated as follows:

[0124] C D = C D0,CR + C D2,CR * C L 2

[0125] Among them, C D0,CR is the parasite drag coefficient of the aircraft in the cruise phase (this coefficient can be queried in the BADA dataset), C D2,CR is the induced drag coefficient of the aircraft in the cruise phase (this coefficient can be queried in the BADA dataset), and C L is the lift coefficient of the aircraft.

[0126] The lift coefficient C L of the aircraft is calculated as follows:

[0127]

[0128] Among them, is the bank angle of the aircraft.

[0129] The bank angle of the aircraft is calculated as follows:

[0130]

[0131] Among them, is the turn rate of the aircraft.

[0132] e. Safety cost penalty: In this embodiment, the turn rate of the aircraft itself is considered during trajectory optimization. When the turn rate is too large, the risk factor will increase, and a certain penalty will be given. The safety cost penalty is specifically expressed as follows:

[0133]

[0134] Among them, r turn_rate is the reward value regarding the safety cost penalty.

[0135] Sum up the above reward value components to obtain the total reward value for each step:

[0136] r = r invasion + r goal_distance + r goal + r fuel + r turn_rate

[0137] Among them, r invasion is the penalty for flying into the unavailable airspace, r goal_distance is the penalty for the end - point distance, r goal is the reward for reaching the end - point, r fuel is the penalty for fuel consumption, r turn_rate is the safety cost penalty.

[0138] II. Step 2: Generate an expert trajectory dataset and its trajectory reward value.

[0139] In this embodiment, the A* algorithm is used to generate a trajectory for the aircraft in the free - route airspace. During this process, the entrance of the free - route airspace is used as the starting point, the exit is used as the end - point, and the unavailable airspace is regarded as an obstacle. The trajectory generated by the A* algorithm is the expert trajectory dataset. Interacting the trajectories in this dataset with the reinforcement learning environment constructed in Step 1 can obtain the reward value of the trajectory.

[0140] III. Step 3: Establish a conditional diffusion model and complete the training. The conditional diffusion model includes a diffusion sub - model and a conditional function sub - model.

[0141] Specifically, the diffusion sub-model and the conditional function sub-model need to be trained separately.

[0142] (1) Diffusion sub-model: The diffusion sub-model is a generative model. In the forward process, Gaussian noise is gradually added to the image, and finally a pure noise image is obtained; in the reverse diffusion process, according to the trained neural network, the pure noise image is gradually denoised, and finally a clear image is obtained. The diffusion sub-model can learn the distribution characteristics of the initial image, so that in the reverse diffusion process, the pure noise image can be restored to an image similar to the initial image.

[0143] The data set used for training the diffusion sub-model in this embodiment is trajectory data, which is expressed as follows:

[0144]

[0145] Among them, τ is the trajectory data used as the training set, s1, s2 and s T are the states of the first, second and Tth planning steps of the aircraft, respectively, a1, a2 and a T are the actions of the first planning step, the second planning step, and the Tth planning step of the aircraft respectively. The state and the action are combined to form a two-dimensional matrix. The representation of the trajectory data is the representation of the input and output of the diffusion sub-model. Similarly, the present invention uses τ i To represent the diffusion trajectory of the i-th step, To represent the state and action combination of the t-th planning step of the i-th diffusion trajectory, To represent the state of the t-th planning step of the i-th diffusion trajectory, To represent the action of the t-th planning step of the trajectory of the i-th diffusion step.

[0146] The overall network structure used for training adopts the U-Net network structure commonly used in the diffusion sub-model, which consists of three encoders and three decoders. The overall network structure diagram of the diffusion sub-model is shown in Figure 2 shown.

[0147] Each encoder (or decoder) consists of two time residual blocks. In each time residual block, the input data processed by the convolution module and the time embedding (i.e., the number of diffusion steps) processed by the time embedding fully connected layer are added, the summed data is processed again by the convolution module, and the data obtained by the secondary convolution is added to the input data to form a residual connection.

[0148] The schematic diagram of the temporal residual block structure is as follows Figure 3 shown.

[0149] The time embedding fully connected layer contains a Mish activation function and a linear layer, as Figure 4 shown.

[0150] The convolutional module for processing the input data contains a one-dimensional convolutional layer, a group normalization layer, and a Mish activation function, as Figure 5 shown.

[0151] During the training process, the network learns noise information and simplifies the loss function of the diffusion model for training. The loss function adopted in this embodiment is as follows:

[0152]

[0153] where θ is the diffusion sub-model parameter, i is the diffusion step, ε is the real noise, and ε θ (τ i , i) is the predicted noise, τ 0 is the initial trajectory without added noise, is the joint mathematical expectation of i, ε, and τ 0 .

[0154] (2) Conditional function sub-model: The conditional function fits the reward value characteristics of the trajectory and guides the generation of trajectories with higher reward values during the inverse sampling process. The conditional function can be expressed as follows:

[0155]

[0156] where J(s0, a 0:T ) is the total trajectory reward value, T is the total planning step length, and r(s t , a t ) is the single-step reward value at the t-th planning step.

[0157] The goal of the trajectory optimization in this embodiment is to find a series of actions to maximize the total trajectory reward value, that is:

[0158]

[0159] where is a series of actions corresponding to the maximum total trajectory reward value, s0 is the initial state value, and a 0:T is a series of actions from the 0-T planning step.

[0160] When training the conditional function, the network structure adopts the first half of the diffusion sub-model network structure and finally outputs a scalar through a linear layer. As Figure 6 shown. It includes three encoders, two time residual blocks, and a fully connected layer.

[0161] The input of the conditional function sub-model is the trajectory data, and the output is the total reward value corresponding to the trajectory data.

[0162] The structure of the encoder in the conditional function sub-model is the same as that in the diffusion model network. Each encoder consists of two temporal residual blocks (as Figure 3 shown). The two temporal residual blocks in the conditional function sub-model are also the same as those in the diffusion sub-model (as Figure 3 shown).

[0163] It should be noted that the temporal embedding fully connected layer for processing temporal embeddings in the conditional function sub-model is different from that of the diffusion sub-model. It consists of a sine positional embedding layer, a linear layer, a Mish activation function, and a second linear layer, and the structure is as Figure 7 shown.

[0164] The last fully connected layer in the conditional function sub-model consists of a linear layer, a Mish activation function, and a second linear layer, and the structure is as Figure 8 shown.

[0165] The loss function during the training of the conditional function sub-model is expressed as follows:

[0166]

[0167] where φ is the parameters of the conditional function sub-model, J is the true trajectory reward value, J φ (τ i , i) is the predicted trajectory reward value, is the joint mathematical expectation of i, J, and τ 0 .

[0168] IV. Step 4: Use the conditional diffusion model to provide action decisions to complete the optimization of trajectory planning.

[0169] After training the diffusion sub-model and the conditional function sub-model, clear trajectory data is gradually denoised from pure noise trajectory data according to the diffusion sub-model. During this process, the conditional function sub-model is combined to guide sampling to provide action decisions for the aircraft, thereby generating a trajectory with a higher reward value.

[0170] The algorithm for the diffusion sub-model to combine the conditional function sub-model for guided sampling is as follows:

[0171] The specific steps of step S4 include:

[0172] Step S4-1: Input: parameters θ of the diffusion sub-model, parameters φ of the conditional function sub-model, optimization rate α, total number of diffusion steps N, covariance ∑ i ;

[0173] Step S4-2: Repeat steps S4-3, S4-4, S4-5, and S4-9 until the aircraft reaches the exit;

[0174] Step S4-3: Observe the current state of the aircraft;

[0175] Step S4-4: Initialize a pure noise trajectory τ that follows a standard Gaussian distribution N ~Ν(0,I);

[0176] Step S4-5: For the diffusion step numbers i = N,..., 1, loop through and execute Steps S4-6, S4-7, and S4-8;

[0177] Step S4-6: Obtain the mean μ of the reverse diffusion based on the diffusion sub-model parameter θ and the trajectory τ at the current diffusion step i ;

[0178] Step S4-7: Optimize the mean μ of the reverse diffusion by combining the gradient of the conditional function sub-model parameter φ to obtain a more optimized trajectory for the previous diffusion step, i.e., τ i-1 ~Ν(μ+α∑▽J φ (μ),∑ i );

[0179] Step S4-8: Set the state of the starting planning step length of the trajectory of the previous diffusion step as the current state of the aircraft;

[0180] Step S4-9: The aircraft executes the actions of the initial planning step length of the clear trajectory

[0181] In the present invention, unless otherwise clearly specified and defined, terms such as "install", "connect", "join", "fix", etc. shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0182] In the present invention, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features therebetween. Moreover, the first feature being "above", "over", and "on top of" the second feature includes that the first feature is directly above and diagonally above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below", and "beneath" the second feature includes that the first feature is directly below and diagonally below the second feature, or merely indicates that the horizontal height of the first feature is lower than that of the second feature.

[0183] In the present invention, the terms "first", "second", "third", and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "plurality" refers to two or more, unless otherwise clearly defined.

[0184] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An optimization method for aircraft trajectory planning based on a conditional diffusion model, characterized in that, It includes the following steps: Step S1: Construct a reinforcement learning environment for free route airspace trajectory planning, including a state space S, an action space A, and a reward function r. The reward function r includes penalties for flying into unavailable airspace, penalties for the distance to the end point, rewards for reaching the end point, penalties for fuel consumption, and penalties for safety costs; Step S2: Generate an expert trajectory dataset and its trajectory reward values; Step S3: Establish a conditional diffusion model and complete training. The conditional diffusion model includes a diffusion sub-model and a conditional function sub-model; Step S4: Use the conditional diffusion model to provide action decisions and complete the optimization of trajectory planning.

2. The aircraft trajectory planning optimization method based on the conditional diffusion model according to claim 1, wherein, In step S1, the state space S is: S = [x now , y now , heading, x start , y start , x goal , y goal ​ where x now and y now are the x-axis coordinate and y-axis coordinate of the current position of the aircraft respectively, and heading is the current heading angle of the aircraft, x start and y start are the x-axis coordinate and y-axis coordinate of the starting position of the current flight path planning task of the aircraft respectively, x goal and y goal are the x-axis coordinate and y-axis coordinate of the ending position of the current flight path planning task of the aircraft respectively.

3. The aircraft flight path planning optimization method based on a conditional diffusion model according to claim 1, wherein In step S1, the action space A is: A = [v x , v y ​ where v x and v y are the x-axis component and y-axis component of the current horizontal speed of the aircraft, respectively.

4. The aircraft flight path planning optimization method based on a conditional diffusion model according to claim 1, characterized in that In step S1, the reward function r is: r=r invasion +r goal_distance +r goal +r fuel +r turn_rate Among them, r invasion is the penalty for flying into the no-fly zone, r goal_distance is the penalty for the end-point distance, r goal is the reward for reaching the end point, r fuel is the penalty for fuel consumption, r turn_rate is the penalty for safety cost.

5. The aircraft flight path planning optimization method based on a conditional diffusion model according to claim 1, characterized in that The said r invasion has the following expression: The said r goal_distance has the following expression: r goal_distance = -||position now -positon goal || * 0.02 where position now and positon goal are the current position and the destination position of the aircraft respectively, and ||·|| is the Euclidean norm; The said r goal has the following expression: The said r fuel has the following expression: r fuel = -f cr *dt*0.1 where f cr is the cruise fuel flow rate of the aircraft, and dt is the planned step time interval; The said r turn_rate has the expression:

6. The aircraft flight path planning optimization method based on the conditional diffusion model according to claim 1, wherein Step S2 specifically includes: Taking the entrance of the free route airspace as the starting point, the exit as the end point, and the unavailable airspace as obstacles, use the A* algorithm to generate an expert trajectory dataset, and interact with the reinforcement learning environment for free route airspace trajectory planning to obtain corresponding trajectory reward values.

7. The aircraft flight path planning optimization method based on a conditional diffusion model according to claim 1, characterized in that Step S3 specifically includes: Step S3-1: Establish a diffusion sub-model and train the diffusion sub-model using the trajectory data in the expert trajectory dataset; Step S3-2: Establish a conditional function sub-model and train the conditional function sub-model using the trajectory data and its trajectory reward values in the expert trajectory dataset.

8. The method for optimizing aircraft trajectory planning based on a conditional diffusion model according to claim 7, characterized in that The diffusion sub-model in step S3-1 adopts a U-Net network structure, which consists of three encoders and three decoders. Each encoder or decoder is composed of two temporal residual blocks. In each temporal residual block, the input data processed by the convolutional module and the temporal embedding processed by the temporal embedding fully connected layer are added together, and then secondary convolution is performed through the convolutional module. The obtained data is added to the input data to form a residual connection; the temporal embedding fully connected layer contains a Mish activation function and a linear layer; The conditional function sub-model in step S3-2 consists of three encoders, two temporal residual blocks, and a fully connected layer. Each encoder is composed of two temporal residual blocks. In each temporal residual block, the input data processed by the convolutional module and the temporal embedding processed by the temporal embedding fully connected layer are added together, and then secondary convolution is performed through the convolutional module. The obtained data is added to the input data to form a residual connection; the temporal embedding fully connected layer contains a sinusoidal position embedding layer, two linear layers, and a Mish activation function.

9. The method for optimizing aircraft trajectory planning based on a conditional diffusion model according to claim 7, characterized in that The trajectory data τ in step S3-1 is expressed as: Among them, s1, s2, and s T are the states of the aircraft at the first planning step length, the second planning step length, and the T-th planning step length respectively, and a1, a2, and a T are the actions of the aircraft at the first planning step length, the second planning step length, and the T-th planning step length respectively; The loss function used for training is: where θ is the diffusion sub-model parameter, i is the number of diffusion steps, ε is the true noise, and ε θ (τ i , i) is the predicted noise, τ 0 is the initial trajectory without added noise, is the joint mathematical expectation of i, ε, and τ 0 ; The conditional function in step S3-2 is expressed as: Among them, J(s0,a 0:T ) is the total trajectory reward value, T is the total number of planning steps, r(s t ,a t ) is the single-step reward value at the t-th planning step, is a series of actions corresponding to the maximum total trajectory reward value, s0 is the initial state value, a 0:T is a series of actions from the 0-T planning steps; The loss function used for training is: Among them, φ is the parameter of the conditional function sub-model, J is the true trajectory reward value, and J φ (τ i , i) is the predicted trajectory reward value, is the joint mathematical expectation of i, J, and τ 0 .

10. The aircraft flight path planning optimization method based on a conditional diffusion model according to claim 9, characterized in that, Step S4 specifically includes: Step S4-1: Input: Diffusion sub-model parameters θ, conditional function sub-model parameters φ, optimization rate α, total number of diffusion steps N, covariance ∑ i ; Step S4-2: Repeat steps S4-3, S4-4, S4-5, S4-9 until the aircraft reaches the exit; Step S4-3: Observe the current state of the aircraft; Step S4-4: Initialize a pure noise trajectory τ that follows a standard Gaussian distribution N ~Ν(0,I); Step S4-5: For diffusion step numbers i = N,..., 1, loop through and execute Steps S4-6, S4-7, and S4-8; Step S4-6: Obtain the mean μ of inverse diffusion according to the diffusion sub-model parameter θ and the trajectory τ at the current diffusion step i ; and obtain the mean μ of inverse diffusion Step S4-7: Optimize the mean μ of the inverse diffusion in combination with the gradient of the conditional function sub-model parameter φ to obtain a more optimized trajectory for the previous diffusion step, that is Step S4-8: Set the state of the starting planned step size of the trajectory of the previous diffusion step as the current state of the aircraft; Step S4-9: The aircraft performs the action of the initial planning step size for the clear trajectory

Citation Information

Cited By

  • Airfoil profile multi-objective constraint optimization method and system

    CN122197522A