Aircraft trajectory intelligent decision-making method, system, equipment and medium

Through deep reinforcement learning and artificial potential field methods guided by multiple experts, the efficiency and accuracy of trajectory planning for aircraft in complex environments have been improved, solving the problems of high training costs and insufficient strategy generalization in traditional methods, and achieving efficient and safe trajectory planning for aircraft formations.

CN120653001APending Publication Date: 2025-09-16UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802425.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional aircraft trajectory planning lacks flexibility and adaptability in complex environments. In addition, methods based on deep reinforcement learning have high training costs, and imitation learning is difficult to cover all scenarios, resulting in insufficient strategy generalization.

Method used

A multi-expert guided deep reinforcement learning method is adopted to divide the aircraft formation into a leading aircraft and a following aircraft. The leading aircraft plans its trajectory through multi-expert guided deep reinforcement learning, and the following aircraft adjusts its trajectory using the artificial potential field method. Combined with the dual-time scale architecture and linearization processing, the learning efficiency and strategy robustness are improved.

Benefits of technology

It improves the adaptability of the aircraft in complex environments and the accuracy of trajectory planning, reduces the time cost of model training, and achieves efficient trajectory planning and safe flight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653001A_ABST
    Figure CN120653001A_ABST
Patent Text Reader

Abstract

The invention discloses an aircraft trajectory intelligent decision-making method, system and equipment and a medium, provides an intelligent decision-making technology for multi-expert demonstration guidance, and aims to improve the intelligent decision-making ability of an aircraft when the aircraft deals with complex environments and multi-aspect performance requirements, and utilize demonstration data of experts in different fields to improve the intelligent decision-making ability of the aircraft when the aircraft deals with multi-aspect performance requirements. The intelligent agent is guided to become a comprehensive intelligent agent with capabilities of various fields, and even if a sub-optimal strategy is adopted by demonstration data, the intelligent agent can obtain a better strategy through training. Meanwhile, a layered framework (a pilot aircraft and a following aircraft) is designed, and cooperation of aircraft formation is achieved. The multi-expert guided intelligent decision-making combines the advantages of imitation learning and deep reinforcement learning, the learning efficiency and strategy robustness are improved, the comprehensive performance of the intelligent agent is improved, and a reliable solution is provided for autonomous and intelligent application of the aircraft in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-aircraft autonomous flight trajectory planning, and in particular to an aircraft trajectory intelligent decision-making method, system, equipment and medium. Background Art

[0002] With the development of 6G communication networks and the Internet of Things (IoT) technology, the number of network devices and users has increased dramatically, posing severe challenges to the coverage and capacity of ground networks. Unmanned aerial vehicles (UAVs) with communication capabilities can serve as aerial base stations, providing flexible airborne coverage and are considered a reliable solution for expanding coverage and alleviating communication needs. In recent years, UAVs have shown promising application prospects in logistics and distribution, agricultural monitoring, disaster relief, urban inspections, and other fields. In mountainous areas, disaster zones, or densely populated urban areas, UAVs can be quickly deployed as aerial base stations (BSs), providing flexible and low-cost communication services. Due to limited onboard energy resources, aircraft trajectories must be rationally planned to improve energy efficiency. Furthermore, in complex application environments, aircraft must consider obstacle avoidance to ensure flight safety. Traditional aircraft control relies primarily on preprogrammed rules, lacking flexibility and adaptability in navigating obstacles and in multi-task collaborative scenarios. Therefore, intelligent decision-making regarding aircraft trajectories in complex environments has become a critical issue that needs to be addressed.

[0003] Deep reinforcement learning (DRL) optimizes action strategies and makes autonomous decisions through continuous trial-and-error interactions between an intelligent agent and its environment, offering tremendous potential for aircraft trajectory planning. Existing research has validated the effectiveness of deep reinforcement learning in scenarios such as aircraft obstacle avoidance and energy optimization. However, deep reinforcement learning-based methods require extensive and lengthy iterative training to obtain usable agent strategies, which has limitations in real-world applications. Among existing improvements, imitation learning accelerates policy learning through expert demonstration data. However, high-quality expert data is difficult to obtain and covers all boundary scenarios. Therefore, an intelligent decision-making framework that combines expert prior knowledge, ensures safe exploration, and balances efficient learning is urgently needed.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, system, device and medium for intelligent decision-making of aircraft trajectories, which improves the adaptability of aircraft in unknown environments, reduces the time cost of model training, and greatly improves the accuracy and efficiency of trajectory planning.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] An intelligent decision-making method for an aircraft trajectory, comprising:

[0008] The aircraft formation is divided into a lead aircraft and a follower aircraft. The lead aircraft makes decisions and leads the aircraft formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions and maintains a formation relationship with the lead aircraft. The trajectory planning task optimization problem is constructed by combining the aircraft formation's coverage of all hot spots and the degree to which the follower aircraft maintains its formation during flight.

[0009] The optimization problem includes optimization problems related to the pilot aircraft and optimization problems related to the follower aircraft; an intelligent agent is configured on the pilot aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problem related to the follower aircraft; wherein, the intelligent agent of the pilot aircraft makes decisions in combination with a deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method. Combining the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problem related to the pilot aircraft, a constrained policy optimization problem is constructed and linearized, and the calculated intermediate parameters are used to update the parameters of the pilot aircraft intelligent agent.

[0010] An aircraft trajectory intelligent decision system, used to implement the aforementioned method, comprises:

[0011] The optimization problem construction unit for the trajectory planning task is used to divide the aircraft formation into a lead aircraft and a follower aircraft. The lead aircraft makes decisions and leads the aircraft formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions and maintains the formation relationship with the lead aircraft. The optimization problem of the trajectory planning task is constructed based on the coverage of all hot spots by the aircraft formation and the formation maintenance degree of the follower aircraft during flight.

[0012] A decision optimization unit based on deep reinforcement learning guided by multiple experts is used for decision optimization; wherein, the optimization problem includes optimization problems related to the pilot aircraft and optimization problems related to the follower aircraft; an intelligent agent is configured on the pilot aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problems related to the follower aircraft; the intelligent agent of the pilot aircraft makes decisions in combination with the deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method. Combining the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problems related to the pilot aircraft, a constrained policy optimization problem is constructed and linearized, and the calculated intermediate parameters are used to update the parameters of the pilot aircraft intelligent agent.

[0013] A processing device comprising: one or more processors; a memory for storing one or more programs;

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0015] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0016] It can be seen from the technical solution provided by the present invention that a multi-expert demonstration-guided intelligent decision-making technology is proposed, which aims to improve the intelligent decision-making ability of aircraft when dealing with complex environments and multi-faceted performance requirements. The demonstration data of experts in different fields are used to guide the intelligent agent to become a comprehensive intelligent agent with capabilities in various fields. Even if the demonstration data adopts a suboptimal strategy, the intelligent agent can obtain a better strategy through training. At the same time, the present invention designs a hierarchical framework (i.e., divided into a leading aircraft and a following aircraft) to achieve the coordination of aircraft formations (clusters). The intelligent decision-making guided by multiple experts combines the advantages of imitation learning and deep reinforcement learning, improves learning efficiency and strategy robustness, improves the comprehensive performance of the intelligent agent, and provides a reliable solution for the autonomous and intelligent application of aircraft in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of an aircraft trajectory intelligent decision-making method provided by an embodiment of the present invention;

[0019] Figure 2 An architectural diagram of an intelligent aircraft trajectory decision-making method provided by an embodiment of the present invention;

[0020] Figure 3 A dual-time-scale time slot division diagram provided by an embodiment of the present invention;

[0021] Figure 4 A schematic diagram of intelligent decision-making guided by multiple imperfect experts provided in an embodiment of the present invention;

[0022] Figure 5 A schematic diagram of the artificial potential field of a follower aircraft provided in an embodiment of the present invention;

[0023] Figure 6 A schematic diagram of an aircraft trajectory intelligent decision-making system provided by an embodiment of the present invention;

[0024] Figure 7 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] First, the following terms may be used in this article:

[0027] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.

[0028] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.

[0029] The following describes in detail the intelligent aircraft trajectory decision-making method, system, device, and medium provided by the present invention. Any information not described in detail in the examples of the present invention is prior art known to those skilled in the art. For any conditions not specified in the examples of the present invention, the procedures were performed according to conventional conditions in the art or the conditions recommended by the manufacturer. For any reagents or instruments used in the examples of the present invention without manufacturer identification, they are all commercially available conventional products.

[0030] Example 1

[0031] The embodiment of the present invention provides an intelligent decision-making method for aircraft trajectory, such as Figure 1 As shown, it mainly includes the following steps:

[0032] Step 1: Design a hierarchical framework for the aircraft formation and construct the optimization problem of the trajectory planning task.

[0033] In an embodiment of the present invention, an aircraft formation is divided into an upper-layer lead aircraft and a lower-layer follower aircraft. The lead aircraft makes decisions, leads the aircraft formation to explore the target mission area, and plans a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions, maintains a formation relationship with the lead aircraft; combined with the coverage of all hot spots by the aircraft formation and the degree of formation maintenance of the follower aircraft during flight, an optimization problem of the trajectory planning task is constructed.

[0034] Preferably, the coverage of all hotspot areas by the aircraft formation is calculated in the following manner:

[0035] All hotspot areas constitute the set in, is a single hotspot area, m=1,…,M, where M is the number of hotspot areas;

[0036] For a single hotspot area Let a m (t) represents the hot spot area The coverage situation in the tth time slot (corresponding to the time scale for the pilot aircraft to make decisions), a m (t) = 1 means that the user is covered by at least one aircraft (i.e., any user in the hotspot area is at least within the coverage of one aircraft), otherwise a m (t)=0, calculate the hot spot area The coverage rate c from the 1st time slot to the tth time slot m (t) indicates whether the time slot is covered by an aircraft, which is 1 if yes, and 0 if not. The aircraft here includes the following aircraft and the pilot aircraft.

[0037] By integrating the coverage rates of all hotspot areas, we can obtain the coverage degree of all hotspot areas by the aircraft formation, which is expressed as:

[0038]

[0039] Where C(t) is the coverage of all hotspot areas by the aircraft formation in the tth time slot.

[0040] Preferably, the formation maintenance degree of the following aircraft during the flight is calculated by the following formula:

[0041]

[0042] Where F(t) is the formation keeping factor of the t-th time slot, which is used to measure the formation keeping degree of the following aircraft during the flight of the t-th time slot, and N is the number of following aircraft; s n =1 means follow aircraft B n Within the formation range, otherwise s n =0; Indicates following aircraft B at the last step S of time slot t n Position relative to the pilot aircraft; Indicates the initial relative position (formation position), that is, initially following aircraft B n Relative to the position of the pilot aircraft, the time slot corresponds to the time scale for the pilot aircraft to make decisions, and the step corresponds to the time step for the follower aircraft to make decisions. Each time slot contains S steps.

[0043] Preferably, the optimization problem of the trajectory planning task is expressed as:

[0044]

[0045] stc1.0≤x n (t,s)≤L1,0≤y n (t,s)≤L2,h Lowest ≤z n (t,s)≤h Highest

[0046]

[0047] Where C(T′) is the coverage of all hotspot areas by the aircraft formation in the T′th time slot, and F(t) is the formation maintenance factor in the tth time slot, which is used to measure the formation maintenance degree of the following aircraft during the flight in the tth time slot. represents the position of the pilot aircraft in time slot t and time slot t+1; st is the constraint condition, c1~c4 are four constraints; x n (t,s),yn (t,s),z n (t,s) is the following aircraft B n The position of the sth step in the tth time slot corresponds to the coordinates of the x, y, and z axes, s = 1, ..., S, and the step corresponds to the time step of the follower aircraft's decision-making. Each time slot contains S steps; L1 and L2 are the number of map units in the horizontal and vertical directions of the target mission area, respectively, h Lowest With h Highest are the lower and upper limits of the target mission area height respectively; v max is the maximum flight speed, t′ is the length of a time slot; R cov is the maximum communication range of the pilot aircraft, R s The safety radius between following aircraft; Indicates the pilot aircraft B0 and the follower aircraft B n The spacing, To follow aircraft B n With follower aircraft B n′ spacing.

[0048] Step 2: Decision optimization based on deep reinforcement learning guided by multiple experts.

[0049] In an embodiment of the present invention, the optimization problem includes an optimization problem related to a lead aircraft and an optimization problem related to a follower aircraft; an intelligent agent is configured on the lead aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problem related to the follower aircraft; wherein, the intelligent agent of the lead aircraft makes decisions in combination with a deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method, and combines the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problem related to the lead aircraft to construct a constrained policy optimization problem and perform linearization processing, and use the calculated intermediate parameters to update the parameters of the intelligent agent of the lead aircraft.

[0050] Preferably, the deep reinforcement learning algorithm decision-making guided by multiple experts and the optimization problem related to the pilot aircraft are combined to construct a constrained policy optimization problem and perform linearization processing, which can be expressed as:

[0051]

[0052] in, is the agent strategy for the kth iteration, θ k ,θ k+1 The corresponding agent parameters are the kth and k+1th iterations; is the i-th expert strategy, θ eiis the parameter of the i-th expert strategy, η is the objective function, that is, the optimization problem related to the pilot aircraft; st is the constraint condition, D KL For the KL divergence calculation function, the first constraint is to limit the strategy function With the policy function The KL divergence is less than or equal to the set expert constraint constant d ki , the first constraint is to limit the policy function With the policy function The KL divergence is less than or equal to the set KL constraint constant δ; the strategy function Expert-based strategy The probability of taking action a in state s, the policy function Strategy-based In state s, take action a, the policy function Strategy-based Take action a in state s.

[0053] Preferably, the constrained strategy optimization problem is linearized into the following form and solved iteratively using a sequential quadratic programming method:

[0054] argmax g T Δθ

[0055]

[0056] Among them, b i ,c i ,H are all intermediate parameters, T is the transpose symbol; logp(θ k ), p(θ ei ) is the parameter θ of the i-th expert strategy ei The probability distribution under k ) is the agent parameter θ k The probability distribution under i =d ki -d θki ,parameter is the symbol of partial derivative, represents the partial derivative with respect to the agent parameters θ.

[0057] Preferably, the selecting of actions using the artificial potential field method includes: maintaining a formation relationship with the lead aircraft and maintaining a distance from other following aircraft through the selected actions.

[0058] There are two types of potential fields in the artificial potential field method, namely the attractive potential field and the repulsive potential field. The attractive potential field is used to maintain the formation relationship with the leading aircraft, and the repulsive potential field is used to maintain the distance with different following aircraft.

[0059] The attractive potential field U generated by the position of the aircraft formation att (p) is:

[0060]

[0061] The corresponding gravitational force is:

[0062] Among them, p represents the position vector of the following aircraft, p g represents the formation position vector of the following aircraft (the ideal formation position of the following aircraft, that is, the absolute position of the pilot aircraft plus the initial relative position of the following aircraft), k att is the attractive potential field coefficient; Denotes the attractive potential field U att The gradient of (p).

[0063] The repulsive potential field U generated by the obstacle rep (p) is:

[0064]

[0065] Among them, p o represents the position vector of the obstacle, k rep is the repulsive potential field coefficient generated by the obstacle, d o Indicates the range of repulsive force, U rep (p,p o ) is p o The repulsive potential field generated by the obstacle at position p on the aircraft at position p is denoted by F rep (p,p o );

[0066] The repulsive potential field between the following aircraft is:

[0067]

[0068] Among them, U drone (p n ,p n′ ) is the following aircraft B n With follower aircraft B n′ The repulsive potential field between them is denoted by F drone (p n ,p n′ ), R s k is the safety radius between following aircraft, droneis the repulsive potential field coefficient between following vehicles.

[0069] The total potential field is the superposition of the attractive potential field and the repulsive potential field:

[0070] U total (p) = U att (p)+U rep (p)+U drone (p)

[0071] Among them, U total (p) is the total potential field of the following vehicle;

[0072] Calculate the resultant force F corresponding to the total potential field total (p):

[0073]

[0074] The flight direction of the aircraft is determined according to the resultant force, that is, the direction of the resultant force is selected as the movement direction of the aircraft.

[0075] The solution provided by an embodiment of the present invention is a multi-expert-guided intelligent decision-making technology for addressing the trajectory planning problem of multiple aircraft. This solution cleverly combines deep reinforcement learning and imitation learning techniques to provide an efficient solution for optimizing the trajectory of drones in complex environments, improving the learning speed and initial performance of the intelligent agent. By guiding demonstrations from experts in multiple fields, the present invention overcomes the limitations of single expert data, enabling the intelligent agent to learn strategies with superior overall performance and enhancing the model's adaptability. At the same time, by converting expert demonstrations into constraints, it breaks away from the traditional imitation learning's reliance on perfect demonstration data and prevents the intelligent agent from converging to suboptimal strategies. Furthermore, the present invention adopts a hierarchical collaborative architecture and designs a dual-time-scale architecture to achieve collaboration between the pilot and follower aircraft. Through the training and optimization process of deep reinforcement learning, the drone can autonomously explore and converge to the optimal strategy, planning a route that covers all hotspots and is as energy-efficient as possible, while avoiding illegal behaviors such as collisions with obstacles.

[0076] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.

[0077] 1. Overall overview of the plan.

[0078] Considering the high training costs and low sample efficiency of traditional deep reinforcement learning methods, imitation learning leverages expert demonstrations to quickly acquire a good initial policy, reducing learning time. However, expert demonstrations often fail to cover all scenarios, resulting in insufficient policy generalization. When demonstration data is insufficient, a complete policy distribution cannot be trained. Furthermore, expert data on optimal policies is difficult to obtain, and imperfect demonstration data can also affect model performance.

[0079] To this end, the present invention proposes an intelligent decision-making technology guided by multi-expert demonstrations, aiming to improve the intelligent decision-making capabilities of aircraft in complex environments and facing multiple performance requirements. Demonstration data from experts in different fields is used to guide intelligent agents to become comprehensive intelligent agents with capabilities in various fields. Even if the demonstration data adopts a suboptimal strategy, the intelligent agent of the present invention can be trained to obtain a more optimal strategy. Furthermore, the present invention designs a hierarchical framework to achieve coordination among aircraft clusters: the upper-level leader UAV (LUAV) makes overall trajectory decisions for the formation through a deep reinforcement learning method guided by multiple experts, selecting an appropriate path to avoid obstacles and cover all hotspots. An artificial potential field-based method is applied to the lower-level follower UAVs (FUAVs) to adjust the trajectory and achieve overall control of the aircraft formation. Each follower aircraft designs an artificial potential field based on local observation information and instructions from the upper-level leader aircraft (specifically, the formation position within the attractive potential field generated by the follower aircraft's formation position) to perform trajectory planning, ensuring safe flight without obstacles. Multi-expert-guided intelligent decision-making combines the advantages of imitation learning and deep reinforcement learning, improves learning efficiency and strategy robustness, enhances the overall performance of the intelligent agent, and provides a reliable solution for the autonomous and intelligent application of aircraft in complex environments.

[0080] This paper regards the trajectory planning of aircraft serving ground hotspots as a multi-objective optimization problem and proposes a deep reinforcement learning algorithm based on the guidance of multiple imperfect experts, aiming to simultaneously improve the aircraft survival rate, ground hotspot area coverage and energy efficiency. It can be guided by expert demonstrations with different performance, improving the comprehensive performance of the intelligent agent strategy, realizing the intelligent and efficient aircraft trajectory planning, and is suitable for complex environments, resource-constrained scenarios and other scenarios, with significant technical advantages and application value.

[0081] 2. Detailed introduction of the plan.

[0082] 1. Introduction to application scenarios.

[0083] The present invention contemplates an aircraft-assisted communication network, such as Figure 2As shown in the figure, in a complex urban environment, there are multiple temporary hotspots. These hotspots (areas with dense user density within the target area to be explored) are densely populated with ground users and have high communication demands. To meet the communication needs of ground users, multiple aircraft are deployed in formation to establish a temporary communication network. The aircraft in the formation are connected via air-to-air (A2A) links.

[0084] The aircraft formation is divided into a pilot aircraft and a follower aircraft. The hierarchical aircraft decision-making adopts a dual time scale architecture. The pilot aircraft makes decisions on a large time scale (slot), leading the formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions on a small time scale (step), maintaining the formation relationship with the pilot aircraft and avoiding collisions between the follower aircraft. The dual time scale slot division is as follows: Figure 3 As shown, each slot consists of S steps.

[0085] Let the aircraft set be Where N is the number of follower aircraft, B0 is the pilot aircraft. Navigate in three-dimensional space, Indicates the position of the follower aircraft at the tth slot and sth step. Indicates following aircraft B n The relative position of the pilot aircraft.

[0086] The target mission area is discretized into L1×L2 map units. The present invention considers multiple hotspot areas. Where M represents the number of hotspots. In practice, hotspots with high user density typically face higher communication demands. Therefore, the focus is on covering these hotspots to effectively optimize user communication conditions.

[0087] 2. Mathematical modeling of multi-objective optimization problems.

[0088] For hot spots Let a m (t) = 1 means that the hotspot area is covered by at least one aircraft in the tth time slot, otherwise a m (t) = 0. In addition, let c m (t) indicates Coverage rate up to time slot t:

[0089]

[0090] To indicate the coverage of each user cluster (hotspot area), the present invention defines the coverage rate as:

[0091]

[0092] The present invention introduces a formation keeping factor to measure the degree of formation keeping of the following aircraft during the entire flight process:

[0093]

[0094] Among them, s n =1 means follow aircraft B n Within the formation, that is Otherwise n =0.

[0095] This paper assumes that the entire trajectory planning task is divided into T′ slots. Integrating the above performance indicators and giving the optimization problem of the entire trajectory planning task is as follows:

[0096]

[0097] stc1.0≤x n (t,s)≤L1,0≤y n (t,s)≤L2,h Lowest ≤z n (t,s)≤h Highest ,

[0098]

[0099] The above optimization problem aims to optimize the flight path length of the aircraft, improve the coverage of the hot spot area, and ensure the formation of the aircraft. In addition, the optimization problem incorporates necessary constraints. Constraint c1 limits the flight range of the aircraft to prevent them from exceeding the designated target area; constraint c2 limits the maximum flight speed of the pilot aircraft; constraint c3 ensures that the follower aircraft does not exceed the coverage range of the pilot aircraft; constraint c4 ensures that the necessary safety distance is maintained between aircraft to avoid collision. Existing research shows that the above optimization problem is NP-hard. Therefore, the present invention proposes a multi-expert guided deep reinforcement learning algorithm for aircraft trajectory planning, which determines the optimal flight trajectory of the aircraft by utilizing the interaction ability of the aircraft and the environment.

[0100] 3. Deep reinforcement learning algorithm guided by multiple experts.

[0101] This paper proposes a multi-expert-guided deep reinforcement learning algorithm, MEG-DRL (Multi-Imperfect Experts Guidance), which guides intelligent agents to learn from experts, enabling them to quickly achieve better overall performance. Furthermore, the present invention uses a decision-making method based on artificial potential fields on the follower aircraft.

[0102] (3.1) Multi-expert demonstration guidance.

[0103] Figure 4 The MEG-DRL algorithm proposed in the present invention is demonstrated, showing the guiding role of expert strategies on intelligent agents. In existing research, the reinforcement learning method with expert demonstrations (RLfD) requires that the expert demonstration adopts the most strategic and sufficient data. Imperfect expert demonstrations may cause the agent's strategy to converge to a suboptimal strategy. In order to deal with imperfect expert demonstrations, the present invention reformulates the RLfD task as a constrained policy optimization problem, in which the goal is specified by the original goal (i.e., the optimization problem related to the pilot aircraft in the aforementioned optimization problem), and the constraint limits the exploration area to a certain threshold of the demonstration (i.e., the set expert constraint limit constant d ki ). Penalizing imperfect demonstrations leads to noisy and misleading gradient updates, which hinder further performance improvement. This is not the case with soft constraints. The optimal agent policy is assumed to lie within a region of imperfect expert policies. Once the agent policy is within this region, its optimization is influenced only by interactions with the environment, not by demonstrations. Expert demonstrations adjust the agent's policy updates only when the policy falls outside the constrained region.

[0104] This paper uses KL divergence, or relative entropy, to measure the difference between two probability distributions. The larger the KL divergence, the greater the difference between the two distributions. By restricting the agent's strategy to a region near the expert's strategy, the strategy optimization problem can be formulated as:

[0105]

[0106] in, is the current agent strategy, is an expert demonstration strategy, d ki and δ are the expert constraint and KL constraint constants, respectively. The second constraint is the KL constraint (natural gradient), which is used to limit the step size of each update to correct the direction of the policy update and improve training stability.

[0107] Since it is difficult to find a feasible solution directly under this constraint, and the computational cost is too high for the dimension of the neural network parameters, the present invention updates the The solution is approximated by linearization.

[0108] argmax g T Δθ

[0109]

[0110] Problems with multiple constraints cannot be solved intuitively using the Lagrange multiplier method. Therefore, the sequential quadratic programming (SQP) method can be used to iteratively solve the problem.

[0111] In addition, considering that the meanings of the symbols appearing in the above formulas have been introduced in detail in the previous text, they will not be repeated here.

[0112] (3.2) Introduction to the key elements of deep reinforcement learning.

[0113] The state input of the pilot aircraft is its own observation. Its action space is recorded as:

[0114] a n (t)=(l n (t),θ n (t))

[0115] Among them, l n (t) and θ n (t) represents the distance and direction of the aircraft's horizontal flight, respectively. The present invention discretizes the continuous action space to reduce the dimension of the action space and simplify the training process. The aircraft's horizontal flight direction is divided into eight angles: j is the index of the angle. Meanwhile, the present invention sets the optional set of horizontal flight distances to l = {0, a′, 2a′, 3a′, 4a′}, where a′ is the unit length.

[0116] (3.3)Algorithm flow.

[0117] The reward function depends on multiple reward factors, including flight path length, hotspot area coverage, and formation maintenance. At the same time, it is necessary to ensure that the aircraft does not make illegal actions such as exceeding the boundary and colliding with obstacles, thereby improving the aircraft's survival rate. Using the original reward function to guide the aircraft to find the best action requires a lot of iterative training. Therefore, multiple experts in different fields will be introduced, each with good performance in one field. Figure 4 As shown, the present invention trains experts in different fields in parallel in the source domain to obtain converged imperfect experts. The source domain refers to the domain where the expert strategy is pre-trained / trained, while the domain to which the present invention is applied is the target domain. These experts focus on different sub-goals. For example, coverage expert E1 strives to provide coverage for more hotspots, energy efficiency expert E2 explores energy-saving flight strategies to shorten flight paths, and legal action expert E3 avoids illegal behavior. Notably, the present invention simplifies the state and reward spaces of the imperfect experts, enabling them to converge quickly. The specific algorithm is shown in Table 1.

[0118] Table 1: MEG-DRL algorithm

[0119]

[0120]

[0121] 4. Artificial potential field method.

[0122] This invention uses an artificial potential field (APF) approach to trajectory planning for follower aircraft, focusing on two key aspects: 1) maintaining formation with the lead aircraft: by designing an attractive potential field, the follower aircraft remains near the lead aircraft's predetermined position. 2) avoiding collisions: by designing a repulsive potential field, the follower aircraft experiences a repulsive force when approaching other aircraft or obstacles, maintaining a safe distance.

[0123] Therefore, the present invention defines the artificial potential field as consisting of two parts: the attractive potential field and the repulsive potential field. The first is the attractive potential field generated by following the position of the aircraft formation:

[0124]

[0125] The corresponding gravitational force is:

[0126] Among them, p represents the position vector of the aircraft, p g represents the formation position vector of the aircraft, k att is the attractive potential field coefficient, Denotes the attractive potential field U att The gradient of (p).

[0127] Define the repulsive potential field generated by the obstacle:

[0128]

[0129] Among them, p o represents the position vector of the obstacle, k rep is the repulsive potential field coefficient, d o Indicates the range of repulsive force, U rep (p,p o ) is p o The repulsive potential field generated by the obstacle at position p on the aircraft at position p.

[0130] Similarly, the repulsive potential field between following vehicles can be defined as:

[0131]

[0132] Among them, U drone (p n ,p n′ ) is the following aircraft Bn With follower aircraft B n′ The repulsive potential field between s k is the safety radius between following aircraft, drone is the repulsive potential field coefficient between following vehicles.

[0133] The repulsive potential field U rep (p,p o ), U drone (p n ,p n′ ) The corresponding repulsive force is denoted as F rep (p,p o ), F drone (p n ,p n′ ).

[0134] The total potential field is the superposition of the attractive potential field and the repulsive potential field:

[0135] U total (p) = U att (p)+U rep (p)+U drone (p)

[0136] Among them, U total (p) is the total potential field following the aircraft.

[0137] The net force on the aircraft is Figure 5 As shown, the resultant force is the superposition of attraction and repulsion, and the aircraft moves in the direction of the resultant force.

[0138]

[0139] Among them, F total (p) is the total potential field U total (p) The corresponding resultant force is used to determine the flight direction of the aircraft. That is, the direction of the resultant force is selected as the direction of movement of the aircraft.

[0140] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.

[0141] Example 2

[0142] The present invention also provides an aircraft trajectory intelligent decision system, which is mainly used to implement the method provided in the above embodiment, such as Figure 6 As shown, the system mainly includes:

[0143] The optimization problem construction unit for the trajectory planning task is used to divide the aircraft formation into a lead aircraft and a follower aircraft. The lead aircraft makes decisions and leads the aircraft formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions and maintains the formation relationship with the lead aircraft. The optimization problem of the trajectory planning task is constructed based on the coverage of all hot spots by the aircraft formation and the formation maintenance degree of the follower aircraft during flight.

[0144] A decision optimization unit based on deep reinforcement learning guided by multiple experts is used for decision optimization; wherein, the optimization problem includes optimization problems related to the pilot aircraft and optimization problems related to the follower aircraft; an intelligent agent is configured on the pilot aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problems related to the follower aircraft; the intelligent agent of the pilot aircraft makes decisions in combination with the deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method. Combining the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problems related to the pilot aircraft, a constrained policy optimization problem is constructed and linearized, and the calculated intermediate parameters are used to update the parameters of the pilot aircraft intelligent agent.

[0145] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0146] Example 3

[0147] The present invention also provides a processing device, such as Figure 7 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.

[0148] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0149] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:

[0150] The input device can be a touch screen, image acquisition device, physical button or mouse;

[0151] The output device may be a display terminal;

[0152] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.

[0153] Example 4

[0154] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0155] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0156] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.

Claims

1. An intelligent decision-making method for aircraft trajectories, characterized in that: include: The aircraft formation is divided into a lead aircraft and a follower aircraft. The lead aircraft makes decisions and leads the aircraft formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions and maintains a formation relationship with the lead aircraft. The trajectory planning task optimization problem is constructed by combining the aircraft formation's coverage of all hot spots and the degree to which the follower aircraft maintains its formation during flight. The optimization problem includes optimization problems related to the pilot aircraft and optimization problems related to the follower aircraft; an intelligent agent is configured on the pilot aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problem related to the follower aircraft; wherein, the intelligent agent of the pilot aircraft makes decisions in combination with a deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method. Combining the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problem related to the pilot aircraft, a constrained policy optimization problem is constructed and linearized, and the calculated intermediate parameters are used to update the parameters of the pilot aircraft intelligent agent.

2. The method for intelligent decision-making of aircraft trajectory according to claim 1, characterized in that: The coverage of all hotspot areas by the aircraft formation is calculated as follows: All hotspot areas constitute the set in, is a single hotspot area, m=1,…,M, where M is the number of hotspot areas; For a single hotspot area Let a m (t) represents the hot spot area In the case of coverage in the tth time slot, a m (t) = 1 means it is covered by at least one aircraft, otherwise a m (t)=0, calculate the hot spot area The coverage rate c from the 1st time slot to the tth time slot m (t) indicates whether the aircraft is covered by the aircraft as of time slot t, which is 1 if yes and 0 otherwise. The time slot corresponds to the time scale for the pilot aircraft to make decisions. By integrating the coverage rates of all hotspot areas, we can obtain the coverage degree of all hotspot areas by the aircraft formation, which is expressed as: Where C(t) is the coverage of all hotspot areas by the aircraft formation in the tth time slot.

3. The method for intelligent decision-making of aircraft trajectory according to claim 1, characterized in that: The formation keeping degree of the following aircraft during the flight is calculated by the following formula: Where F(t) is the formation keeping factor of the t-th time slot, which is used to measure the formation keeping degree of the following aircraft during the flight of the t-th time slot, and N is the number of following aircraft; s n =1 means follow aircraft B n Within the formation range, otherwise s n =0; Indicates following aircraft B at the last step S of time slot t n The position relative to the pilot aircraft, Indicates that the aircraft B is initially followed n Relative to the position of the pilot aircraft, the time slot corresponds to the time scale for the pilot aircraft to make decisions, and the step corresponds to the time step for the follower aircraft to make decisions. Each time slot contains S steps.

4. The method for intelligent aircraft trajectory decision-making according to claim 2 or 3, characterized in that: The optimization problem of the trajectory planning task is expressed as: s.t.c1.0≤x n (t,s)≤L1,0≤y n (t,s)≤L2,h Lowest ≤z n (t,s)≤h Highest c2. c3. c4. Where C(T′) is the coverage of all hotspot areas by the aircraft formation in the T′th time slot, and F(t) is the formation maintenance factor in the tth time slot, which is used to measure the formation maintenance degree of the following aircraft during the flight in the tth time slot. represents the position of the pilot aircraft in time slot t and time slot t+1; st is the constraint condition, c1~c4 are four constraints; x n (t,s),y n (t,s),z n (t,s) is the following aircraft B n The position of the sth step in the tth time slot corresponds to the coordinates of the x, y, and z axes, s = 1, ..., S, and the step corresponds to the time step of the follower aircraft's decision-making. Each time slot contains S steps; L1 and L2 are the number of map units in the horizontal and vertical directions of the target mission area, respectively, h Lowest With h Highest are the lower and upper limits of the target mission area height respectively; v max is the maximum flight speed, t′ is the length of a time slot; R cov is the maximum communication range of the pilot aircraft, R s The safety radius between following aircraft; Indicates the pilot aircraft B0 and the follower aircraft B n The spacing, To follow aircraft B n With follower aircraft B n′ spacing.

5. The method for intelligent decision-making of aircraft trajectory according to claim 1, characterized in that: The above-mentioned deep reinforcement learning algorithm decision-making guided by multiple experts and the optimization problem related to the pilot aircraft are combined to construct a constrained policy optimization problem and perform linearization processing, which can be expressed as: in, is the agent strategy for the kth iteration, θ k ,θ k+1 The corresponding agent parameters are the kth and k+1th iterations; is the i-th expert strategy, θ ei is the parameter of the i-th expert strategy, η is the objective function, that is, the optimization problem related to the pilot aircraft; st is the constraint condition, D KL For the KL divergence calculation function, the first constraint is to limit the strategy function With the policy function The KL divergence is less than or equal to the set expert constraint constant d ki , the first constraint is to limit the policy function With the policy function The KL divergence is less than or equal to the set KL constraint constant δ; the strategy function Expert-based strategy In state s, take action a, the policy function Strategy-based In state s, take action a, the policy function Strategy-based Take action a in state s.

6. The method for intelligent decision-making of aircraft trajectory according to claim 5, characterized in that: Also includes: The constrained policy optimization problem is linearized into the following form and solved iteratively using the sequential quadratic programming method: argmax g T Δθ Among them, b i ,c i ,H are all intermediate parameters, T is the transpose symbol; p(θ ei ) is the parameter θ of the i-th expert strategy ei The probability distribution under k ) is the agent parameter θ k The probability distribution under i =d ki -d θki ,parameter is the symbol of partial derivative, represents the partial derivative with respect to the agent parameters θ.

7. The method for intelligent decision-making of aircraft trajectory according to claim 1, characterized in that: The selecting of actions using the artificial potential field method includes: maintaining a formation relationship with the lead aircraft and maintaining a distance from other following aircraft through the selected actions; There are two kinds of potential fields in the artificial potential field method: attractive potential field and repulsive potential field; The attractive potential field U generated by the position of the aircraft formation att (p) is: The corresponding gravitational force is: Among them, p represents the position vector of the following aircraft, p g represents the formation position vector of the following aircraft, k att is the attractive potential field coefficient, Denotes the attractive potential field U att The gradient of (p); The repulsive potential field U generated by the obstacle rep (p,p o )for: Among them, p o represents the position vector of the obstacle, k rep is the repulsive potential field coefficient generated by the obstacle, d o Indicates the range of repulsive force, U rep (p,p o ) is p o The repulsive potential field generated by the obstacle at position p on the aircraft at position p is denoted by F rep (p,p o ); The repulsive potential field between the following aircraft is: Among them, U drone (p n ,p n′ ) is the following aircraft B n With follower aircraft B n′ The repulsive potential field between them is denoted by F drone (p n ,p n′ ), R s k is the safety radius between following aircraft, drone is the repulsive potential field coefficient between the following aircraft; The total potential field is the superposition of the attractive potential field and the repulsive potential field: IN total =U att +U rep +U drone Among them, U total is the total potential field; Calculate the resultant force F corresponding to the total potential field total (p): Among them, the flight direction of the aircraft is determined according to the resultant force.

8. An aircraft trajectory intelligent decision-making system, characterized in that: The method for implementing any one of claims 1 to 7 comprises: The optimization problem construction unit for the trajectory planning task is used to divide the aircraft formation into a lead aircraft and a follower aircraft. The lead aircraft makes decisions and leads the aircraft formation to explore the target mission area, planning a trajectory that passes through all hot spots in the target mission area and avoids obstacles. The follower aircraft makes decisions and maintains the formation relationship with the lead aircraft. The optimization problem of the trajectory planning task is constructed based on the coverage of all hot spots by the aircraft formation and the formation maintenance degree of the follower aircraft during flight. A decision optimization unit based on deep reinforcement learning guided by multiple experts is used for decision optimization; wherein, the optimization problem includes optimization problems related to the pilot aircraft and optimization problems related to the follower aircraft; an intelligent agent is configured on the pilot aircraft, and an artificial potential field method is deployed on the follower aircraft to solve the optimization problems related to the follower aircraft; the intelligent agent of the pilot aircraft makes decisions in combination with the deep reinforcement learning algorithm guided by multiple experts, and each follower aircraft selects and executes actions using the artificial potential field method. Combining the decision-making of the deep reinforcement learning algorithm guided by multiple experts and the optimization problems related to the pilot aircraft, a constrained policy optimization problem is constructed and linearized, and the calculated intermediate parameters are used to update the parameters of the pilot aircraft intelligent agent.

9. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.