A path planning method, system and electronic device for multiple unmanned aerial vehicles
Through formation matrix description and MADDPG algorithm optimization, the path planning problem of multiple UAV formations in complex environments is solved, safe and efficient flight in obstacle environments is achieved, and the adaptability and fault tolerance of formations are improved.
Patent Information
- Application Number
- CN202310377014.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-04-11
AI Technical Summary
In the existing multi-UAV formation path planning, classic control algorithms have poor adaptability in complex environments, insufficient fault tolerance, and the research on distributed path planning is not yet mature, making it difficult to achieve safe and efficient flight in obstacle environments.
The formation matrix is used to describe the geometric formation of the UAV cluster, and an adaptive formation database is established. The optimal transformation strategy is generated based on the MADDPG algorithm. Combined with environmental constraints and track constraints, the optimal transformation strategy is selected through the formation transformation evaluation function to achieve safe flight of multiple UAV formations.
It improves the accuracy and real-timeness of multi-UAV path planning, enhances fault tolerance in obstacle environments, and ensures safe flight of formations.
Smart Images

Figure CN116382339B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and in particular, to a path planning method, system and electronic device for multiple unmanned aerial vehicles. Background Art
[0002] A single unmanned aerial vehicle (UAV) is unable to perform large-scale tasks such as collaborative search and fuel replenishment due to its own resource limitations. However, a multi-UAV formation can not only execute large tasks that a single UAV cannot complete, but also increase redundancy and improve the task success rate by virtue of its numerical advantage. Therefore, the research on path planning for multi-UAV formations has very important practical significance.
[0003] The formation transformation and obstacle avoidance problems of multi-UAV formations are the key research points and difficulties in the field of multi-UAV control. In the face of the continuous changes in task requirements and environmental conditions, how to comprehensively consider the priorities of various performance parameters, reasonably and efficiently formulate the optimal formation transformation strategy according to local conditions, and achieve safe flight of the formation under obstacle avoidance is an important evaluation index for measuring the quality of UAV formation control technology. Richert of the University of California proposed a formation partition cooperation algorithm by minimizing the voyage cost during the formation's task execution to achieve the purpose of dynamic formation transformation of the UAV formation during flight. Xiamen University addressed the collision avoidance problem in formation transformation, used the PID algorithm to control the UAV formation to form a certain formation structure, and performed obstacle avoidance by scaling the distance between UAVs, thus ensuring the formation and transformation of the formation. Giacomin.P et al. proposed a control algorithm based on trajectory segmentation and completed the formation reorganization by calculating the UAV navigation trajectory in segments. Jiang Rongxin et al. constructed a formation transformation evaluation model based on the total voyage energy consumption and time cost during formation flight, transformed the complex constraint problem of multi-UAV formation transformation in a complex obstacle environment into a problem of solving the optimal solution of a function model, and achieved formation transformation.
[0004] The above research on formation transformation and obstacle avoidance has certain advantages, but there are still many problems to be further solved. The problem of UAV swarm consensus based on classical control algorithms has always been one of the main research directions. However, classical control theory is often based on strict mathematical derivations and requires setting many assumptions and ideal conditions according to experience. These ideal conditions are often difficult to meet in reality, so the environmental adaptability is poor, which limits its development. In contrast, the booming development of artificial intelligence, especially deep reinforcement learning, provides new ideas for the research of multi-agent systems. Different from classical control theory, deep reinforcement learning endows UAVs with stronger practicality and wide applicability when facing complex and unknown environments through continuous interaction between UAVs and the environment. However, most current research is still centralized control, and the fault tolerance of the system is poor. The research on distributed path planning of formations based on multi-agent deep reinforcement learning is not yet mature. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a multi-unmanned aerial vehicle path planning method, system and electronic device.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A multi-unmanned aerial vehicle path planning method, comprising:
[0008] Describing the geometric formation of the UAV swarm in the form of a formation matrix, and establishing an adaptive formation library based on the obtained geometric formation;
[0009] Determining the formation dynamic transformation in the case of obstacle avoidance based on the adaptive formation library;
[0010] Establishing a formation transformation evaluation criterion for environmental constraints and UAV swarm trajectory constraints based on the formation dynamic transformation;
[0011] Constructing a formation transformation evaluation function based on the formation transformation evaluation criterion;
[0012] Selecting the optimal transformation strategy of the formation in the case of obstacle avoidance based on the formation transformation evaluation function; the optimal transformation strategy is generated by the MADDPG algorithm; the MADDPG algorithm is an algorithm improved based on the Actor-Critic algorithm and the DDPG algorithm;
[0013] Realizing the safe flight of the multi-UAV formation in an obstacle environment based on the optimal transformation strategy.
[0014] Preferably, the formation matrix is F:
[0015]
[0016] Wherein, b is the formation expansion parameter, f is the geometric configuration parameter, is the angle parameter of the formation, and n is the total number of UAVs in the cluster.
[0017] Preferably, the formation transformation evaluation criterion for establishing environmental constraints and UAV cluster trajectory constraints based on the formation dynamic transformation specifically includes:
[0018] Construct constraint conditions;
[0019] Determine the formation structure difference degree based on the formation geometric parameters corresponding to the current formation matrix and the formation geometric parameters corresponding to the transformed formation matrix;
[0020] Determine the formation transformation convergence time and the formation transformation path cost;
[0021] Under the constraint conditions, determine the evaluation vector of a formation transformation based on the formation structure difference degree, the formation transformation convergence time, and the formation transformation path cost; use the evaluation vector of the formation transformation as the formation transformation evaluation criterion.
[0022] Preferably, the formation transformation evaluation function is:
[0023] R(F start ,F end )=H(F start ,F end )·α;
[0024] Wherein, R(F start ,F end ) is the evaluation result of a formation transformation, H(F start ,F end ) is the evaluation vector of a formation transformation, α is a column vector, F start is the current formation matrix, and F end is the transformed formation matrix.
[0025] Preferably, the method for selecting the optimal transformation strategy for the formation under obstacle avoidance based on the formation transformation evaluation function specifically includes:
[0026] Determine the formation transformation factor based on the maximum length of the obstacle area and the formation width of the current UAV cluster;
[0027] Select the optimal transformation strategy according to the relationship between the formation transformation factor and the preset transformation threshold range.
[0028] Preferably, the method for selecting the optimal transformation strategy according to the relationship between the formation transformation factor and the preset transformation threshold range specifically includes:
[0029] When the formation transformation factor is greater than the maximum preset transformation threshold, it is not necessary to perform formation transformation to pass through this obstacle area;
[0030] When the formation transformation factor is greater than the minimum preset transformation threshold and less than the maximum preset transformation threshold, the UAV cluster performs formation stretching transformation;
[0031] When the formation transformation factor is less than the minimum preset transformation threshold, the UAV cluster performs formation structural transformation.
[0032] Preferably, the process of the UAV cluster performing formation stretching transformation is as follows:
[0033] The leader in the cluster determines the formation matrix to be finally transformed according to the formation transformation factor, and then sends the formation matrix to be finally transformed to other UAVs for formation stretching transformation.
[0034] Preferably, the process of the UAV cluster performing formation structural transformation is as follows:
[0035] Traverse the formation library to determine the formation transformation evaluation function corresponding to each different geometric formation;
[0036] Select the geometric formation with the smallest formation transformation evaluation function as the formation after transformation, and perform formation structural transformation;
[0037] Among them, when the UAV cluster encounters an obstacle and cannot bypass from the same side, according to the segmentation strategy, the cluster is divided into two sub-formations to bypass from both sides of the obstacle, and then restored to the original formation after crossing the obstacle.
[0038] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0039] The multi-UAV path planning method provided by the present invention, first, under the formation description method based on the leader reference, effectively describes the geometric formation of the UAV cluster using the formation matrix, and establishes an extensible adaptive formation library, laying a foundation for the research of formation transformation; second, studies the dynamic transformation of the formation in the case of obstacle avoidance, establishes a formation transformation evaluation criterion related to environmental constraints and UAV cluster trajectory constraints, constructs a formation transformation evaluation function, and designs an optimal formation transformation strategy in the case of obstacle avoidance using the formation transformation evaluation function; finally, according to the selected formation transformation scheme, uses the MADDPG algorithm to realize the safe flight of the multi-UAV formation in the obstacle environment, which can improve the accuracy and real-time performance of the multi-UAV path planning, and has strong fault tolerance.
[0040] Corresponding to the above-provided multi-UAV path planning method, the present invention also provides the following implementation structure:
[0041] A multi-unmanned aerial vehicle path planning system, comprising:
[0042] A formation library construction module, configured to describe the geometric formation of a UAV cluster in the form of a formation matrix, and establish an adaptive formation library based on the described geometric formation;
[0043] A dynamic transformation determination module, configured to determine the formation dynamic transformation in the case of obstacle avoidance based on the adaptive formation library;
[0044] An evaluation criterion establishment module, configured to establish a formation transformation evaluation criterion for environmental constraints and UAV cluster trajectory constraints based on the formation dynamic transformation;
[0045] An evaluation function determination module, configured to construct a formation transformation evaluation function based on the formation transformation evaluation criterion;
[0046] An optimal transformation strategy selection module, configured to select an optimal transformation strategy for the formation in the case of obstacle avoidance based on the formation transformation evaluation function; the optimal transformation strategy is generated by using the MADDPG algorithm; the MADDPG algorithm is an algorithm obtained by improving the Actor-Critic algorithm and the DDPG algorithm;
[0047] A safe flight control module, configured to achieve the safe flight of the multi-UAV formation in an obstacle environment based on the optimal transformation strategy.
[0048] An electronic device, comprising:
[0049] A memory, configured to store a control program;
[0050] A processor, connected to the memory, configured to retrieve and execute the control program to implement the multi-unmanned aerial vehicle path planning method provided above.
[0051] Since the technical effects achieved by the two implementation structures provided by the present invention are the same as those achieved by the multi-unmanned aerial vehicle path planning method provided above by the present invention, no further description will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0053] Figure 1 It is a flowchart of the multi-unmanned aerial vehicle path planning method provided by the present invention;
[0054] Figure 2 Schematic diagram of common formations of UAV formations provided by an embodiment of the present invention; among them, Figure 2 (a) is a schematic diagram of a UAV in a single-file column; Figure 2 (b) is a schematic diagram of a UAV in a polygon formation;
[0055] Figure 2 (c) is a schematic diagram of a UAV in a V-shaped formation; Figure 2 (d) is a schematic diagram of a UAV in a trapezoidal formation;
[0056] Figure 2 (e) is a schematic diagram of a UAV in a single-file horizontal formation; Figure 2 (f) is a schematic diagram of a UAV in a snake-shaped formation;
[0057] Figure 3 Schematic diagram of formation transformation provided by an embodiment of the present invention;
[0058] Figure 4 Schematic diagram of the UAV formation segmentation strategy provided by an embodiment of the present invention;
[0059] Figure 5 MADDPG algorithm framework diagram provided by an embodiment of the present invention;
[0060] Figure 6 UAV formation flight trajectory diagram provided by an embodiment of the present invention;
[0061] Figure 7 Distance curve graph between UAVs provided by an embodiment of the present invention;
[0062] Figure 8 Distance curve graph between a UAV and obstacle 1 provided by an embodiment of the present invention;
[0063] Figure 9 Distance curve graph between a UAV and obstacle 2 provided by an embodiment of the present invention;
[0064] Figure 10 Distance curve graph between a UAV and obstacle 3 provided by an embodiment of the present invention. Specific implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] The object of the present invention is to provide a multi-unmanned aerial vehicle path planning method, system and electronic device, which can improve the accuracy and real-time performance of multi-unmanned aerial vehicle path planning and have strong fault tolerance.
[0067] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0068] As Figure 1 shown, the multi-unmanned aerial vehicle path planning method provided by the present invention includes:
[0069] S1. Describe the geometric formation of the UAV cluster in the form of a formation matrix, and establish an adaptive formation library based on the obtained geometric formation.
[0070] The UAV cluster formation flight can complete many different types of tasks, such as aerial cooperative observation, fuel replenishment, multi-aircraft cooperative strike, etc. When performing formation tasks, due to the different numbers of UAVs in the UAV cluster, application environments and tasks, there are few tasks that always maintain an inherent geometric formation during the process of the UAV cluster performing tasks. In order to maximize the advantages of the UAV formation, a good formation control algorithm needs to ensure that the formation can be flexibly and quickly changed.
[0071] The UAV formation transformation refers to that after the formation is formed, during the process of the UAV cluster formation flying to the task area, in order to balance the energy consumption of each UAV in the cluster or to smoothly pass through the obstacle area, it is necessary to comprehensively consider the flight constraints and environmental constraints of the UAVs for formation transformation. Efficient and reasonable formation transformation is beneficial to exert the best effect of the UAV formation.
[0072] Since most of the existing UAV flight operations are in the single-aircraft flight mode, the specific formation design can draw on the experience of manned aircraft formation flight. From the long-term combat experience of aircraft, it can be known that the formation can be selected according to different task requirements. When the aviation troops perform combat tasks, correctly selecting and applying the combat formation can give full play to the overall power of the air force and conduct close coordinated operations.
[0073] Common formations in formation control include single-file column, regular polygon formation, V-shaped formation, trapezoidal formation, single-file horizontal formation, snake-shaped formation, etc. As Figure 2As shown. Among them, the single-file column is often used in airdrop and interception missions. It has strong maneuverability and is easy to change formations. It is more capable of performing tasks in complex areas and has higher safety. The regular polygon formation is often used to protect the target, and due to the relatively short distance between each other, the ability to protect each other in case of emergencies is stronger. The V-shaped formation is often used for reconnaissance. The trapezoidal formation is often used for batch attacks. The single-file horizontal formation is often used for wide-front searches, with a relatively large coverage area. It has stronger target detection ability in open and empty areas with good visibility and higher task execution efficiency. The snake formation is often used for large formations to set sail.
[0074] Based on this, in order to establish or maintain a specific geometric formation, it is first necessary to establish a method for representing geometric formations. Currently, there is no unified formation representation method in related research to describe the formations of UAV clusters. The present invention proposes a formation matrix method in the above step S1 for representing UAV formations. In order to describe the position of each UAV in the formation, a formation structure model is established in combination with the formation description method with a leading reference. The formation structure model uses a 4-row and n-column formation matrix F to represent the geometric formation of UAVs, where n represents the number of UAVs in the cluster. In the formation matrix F, the first row is the UAV number, and the second, third, and fourth rows respectively represent the distances between the UAV and the reference UAV in the x, y, and z directions. In the representation method of the formation matrix, the entire formation of the UAV cluster is described by the matrix F, and the formation information of each UAV is represented by a column vector a T i to represent. For example, in Figure 2 the V-shaped formation shown in (c), assuming the lead aircraft is the reference UAV, the formation matrix F of the entire UAV cluster can be represented as:
[0075]
[0076] In the formula, b represents the formation expansion parameter corresponding to this formation, f is the geometric configuration parameter corresponding to this formation, represents the angle parameter of the formation.
[0077] The present invention sets up a formation library for common formations such as "V-shaped", "single-file", "polygon", and "trapezoid" in the field of formation control, which can support UAV clusters of any number, and is used to support the formation formation and formation transformation of UAV cluster formation flight, making the formation transformation more flexible.
[0078] S2. Determine the formation dynamic transformation in case of obstacle avoidance based on the adaptive formation library.
[0079] S3. Establish a formation transformation evaluation criterion based on the dynamic transformation of the formation, taking into account environmental constraints and the flight path constraints of the UAV swarm. When the UAV swarm formation encounters obstacles during the mission execution, the common approaches in existing literature are either to disrupt the original formation and let individual UAVs bypass the obstacles and then reorganize the formation, or to change the formation to a "linear" formation to pass through the obstacle area. None of these existing methods consider the environmental constraints during formation transformation or the formation transformation that does not take into account the flight range constraints and flight performance constraints of the UAV swarm. Neither of them is optimal in terms of time consumption or formation retention rate. Based on this, in order to optimize the efficiency of formation transformation in actual scenarios, the present invention establishes an extensible formation transformation evaluation criterion to measure the formation transformation efficiency from the initial formation F start to the end formation F end , so as to make the optimal deformation transformation selection. Based on this, the process of step S3 for establishing the formation transformation evaluation criterion considering environmental constraints and the flight path constraints of the UAV swarm specifically includes:
[0080] S3-1. UAV kinematic constraints
[0081] During the entire process of formation transformation, the heading angle and heading angular velocity of the UAV must vary within a certain range to meet the flight performance constraint J uav . The constraint conditions are as follows:
[0082]
[0083] In the formula, ψ min , ψ max are the minimum and maximum heading angles of the UAV respectively. are the minimum and maximum heading angular velocities of the UAV respectively.
[0084] S3-2. Formation structural difference
[0085] When the formation is performing a mission, it is often necessary to maintain a specific geometric shape to ensure the efficiency of mission completion. When it is necessary to maintain the lowest overall energy consumption of the UAV swarm, a V-shaped formation is usually maintained, while maintaining a regular polygon formation is beneficial to maximizing the observation range of the formation. Therefore, when encountering an obstacle area and having to perform a formation transformation, maintaining the formation transformation under the same geometric configuration is beneficial to better mission execution.
[0086] The formation structural difference describes the geometric configuration difference of the formation before and after the formation transformation, and is represented by H fsd (F start ,F end ). The calculation method is as follows:
[0087]
[0088] In the above formula, F start is the current formation matrix, and F end represents the formation matrix after transformation. n is the total number of UAVs in the cluster, and F start (k,i) describes the formation geometric parameters corresponding to UAV i in the cluster, where k represents the matrix. The formation structure difference degree H fsd (F start ,F end ) quantitatively describes the difference in geometric configurations corresponding to the two geometric formations before and after transformation from the perspective of formation geometric description, eliminates the influence of the absolute value of formation parameters on the measurement value, and has a clear geometric meaning.
[0089] S3-3. Formation transformation convergence time
[0090] In the UAV cluster formation tracking and observation mission, the UAV cluster involves three processes: formation formation, formation maintenance, and formation transformation until the mission requirements are finally completed. During the entire formation process, the formation convergence rate measures the quality of a formation control algorithm. And a formation transformation describes the process from the destruction of the original formation until it converges to a new formation. In this process, if the convergence time used for the formation transformation is shorter, the efficiency of the overall formation mission is higher.
[0091] The formation transformation convergence time describes the time elapsed from the start of the formation transformation by the UAV cluster until it converges to form a new formation, and the calculation method is as follows:
[0092]
[0093] In the above formula, and respectively represent the time when the i-th UAV that starts to transform the formation earliest in the cluster starts to transform the formation and the time when the i-th UAV that converges to the new formation latest in the cluster converges for the formation transformation.
[0094] S3-4. Formation transformation path cost
[0095] Due to the problems of short endurance time and insufficient navigation energy of UAVs, in order to measure the energy consumption degree of UAVs, researchers usually use the navigation path length of the UAVs in a mission to measure the energy consumption of UAVs.
[0096] To optimize the energy consumption of the UAV swarm during formation transformation, the formation transformation path cost describes the overall path cost of all UAVs in the UAV swarm from the start of the formation transformation to the convergence into a new formation. The calculation method is as follows:
[0097]
[0098] In the above formula, v i represents the real-time speed of the i-th UAV during the formation transformation, represents the total flight path of the i-th UAV during the formation transformation. The path cost traveled by the UAV during the formation transformation is measured by the sum of the absolute values of each segment of the path.
[0099] S3-5: Denote the above evaluation factors as the evaluation vector of the formation transformation. Then, the evaluation vector of a formation transformation can be expressed as:
[0100] H(F start ,F end )=[H fsd (F start ,F end ),H ftct (F start ,F end ),H ftpc (F start ,F end )].
[0101] S4: Construct a formation transformation evaluation function based on the formation transformation evaluation criterion. Based on the specific implementation process of step S3, introduce the column vector α = [α1, α2, α3], where α1, α2, and α3 represent the weight factors corresponding to the formation structure difference degree, the formation transformation convergence time, and the formation change path cost, respectively. Use α1 to represent the importance of maintaining the geometric configuration during the formation transformation, use α2 to represent the importance of maintaining the formation stability during the formation transformation, and use α3 to represent the importance of saving the energy of the UAV swarm. Then, the evaluation function for measuring a formation transformation can be expressed as:
[0102] R(F start ,F end )=H(F start ,F end )·α.
[0103] This formation evaluation function fully considers the environmental constraints of the formation transformation and the flight range constraints of the UAV swarm. Using the formation transformation evaluation vector for calculation enables new evaluation factors to be easily added to the evaluation vector, increasing the scalability of this formation transformation evaluation criterion.
[0104] S5. Select the optimal transformation strategy for the formation during obstacle avoidance based on the formation transformation evaluation function. The optimal transformation strategy is generated using the MADDPG algorithm. The MADDPG algorithm is an algorithm obtained by improving the Actor-Critic algorithm and the DDPG algorithm.
[0105] Among them, in order to fully consider the environmental constraints and the flight trajectory constraints of the UAV swarm during the formation transformation, the optimal formation transformation method is dynamically selected when the formation transformation is required. In this step, a formation transformation factor is introduced, and the dynamic formation transformation is divided into two methods: formation scaling transformation and formation structural transformation, so as to select the optimal formation transformation strategy.
[0106] The formation transformation factor refers to the parameter for realizing the formation scaling transformation in terms of the formation size while keeping the existing geometric configuration of the UAV swarm formation unchanged, and is represented by the symbol ω. Each formation geometric shape corresponds to a specific minimum value of the formation transformation factor, denoted as ω min , and this minimum value represents the minimum value of the formation transformation factor that ensures no collision occurs among the UAVs in the swarm. The calculation method of the formation transformation factor is as follows:
[0107]
[0108] In the formula, Z is the maximum length of the obstacle area, and D F is the formation width of the UAV swarm.
[0109] When performing the formation scaling transformation, there is:
[0110]
[0111] In the formula, ω start and ω end respectively represent the formation transformation factors corresponding before and after the formation transformation. Therefore, when performing the formation scaling transformation, the formation matrix F end after the formation transformation is obtained by performing elementary transformation on the matrix F start before the transformation, that is, F end and F start are similar matrices. Combining the above-obtained formation transformation evaluation function, when F end and F start are similar matrices, the value of the evaluation function is the smallest. Therefore, under the same conditions, the formation scaling transformation is the optimal transformation strategy relative to the formation structural transformation.
[0112] Furthermore, the optimal transformation strategy for dynamic formation based on obstacle avoidance is as follows:
[0113] When the UAV cluster compares its own position data with the known environmental information and finds that there are obstacles or threat areas within the safe distance of the UAV cluster, it first calculates the formation transformation factor of the current formation and makes the following judgments according to the size of the formation transformation factor. The schematic diagram is as Figure 3 shown:
[0114] (1) If the formation transformation factor ω > 1, from the formula it can be seen that the distance between obstacles is greater than the formation width of the UAV cluster at this time, and there is no need to perform formation transformation to pass through this obstacle area.
[0115] (2) If the formation transformation factor ω min <ω<1, it means that a formation stretching and shrinking transformation is performed at this time. By shortening the formation spacing of the UAVs and keeping the geometric configuration unchanged, the obstacle area can be passed through. At this time, the UAV cluster performs a formation stretching and shrinking transformation. The leader in the cluster determines the final formation matrix F end to be transformed according to the formation transformation factor, and then sends it to other UAVs for formation stretching and shrinking transformation. Since the formation width takes into account the minimum safe distance between the UAV cluster and the obstacle area, adjusting the formation width of the transformed formation to the maximum distance of the obstacle area can ensure the safe passage of the UAV cluster through the obstacle area. The formation matrix after transformation at this time is:
[0116]
[0117] In the formula, D Hstart is the formation width of the UAV cluster before the formation transformation.
[0118] (3) If the formation transformation factor ω < ω min , at this time, even if the formation spacing is shortened to the shortest safe spacing to maintain the current geometric configuration, it is still impossible to pass through the obstacle area. Therefore, a formation structural transformation is required. First, traverse the formation library, calculate the formation transformation evaluation function corresponding to each different geometric formation, and select the geometric formation with the smallest formation transformation evaluation function as the formation F end after transformation, and perform a formation structural transformation.
[0119] In particular, when the UAV cluster encounters some obstacles and cannot bypass from the same side, according to the splitting strategy, the cluster can be split into two sub-formations and bypass from both sides of the obstacle respectively, and then resume the original formation after crossing the obstacle.
[0120] Determine whether to split the formation according to the relative position relationship between the formation and the obstacle with the greatest threat. Figure 4 is a schematic diagram of the formation splitting strategy. The formation splitting mainly includes two determining factors. One factor is the current positions of all UAVs. According to the current positions of the UAVs, draw a line passing through the center p obsAlong the current desired relative velocity of the lead aircraft The formation is initially divided into two parts by a straight line in this direction. Another factor is the desired relative velocity of the UAVs after considering each UAV as the lead aircraft For the part of the UAVs that do not include the lead aircraft after the initial division, the following formula is used for further judgment:
[0121]
[0122] In the formula, is the distance between the wingman i and the center of the obstacle, is the distance between the lead aircraft and the center of the obstacle.
[0123] The UAVs that meet the conditions are finally separated from the current formation to form a sub - formation. The two sub - formations will respectively bypass from both sides of the two obstacles. When both sub - formations have crossed the current obstacle, they will re - converge into a new formation.
[0124] Furthermore, the optimal transformation strategy adopted in this step S5 is generated by the MADDPG algorithm. Specifically, in a multi - UAV environment, each UAV needs to continuously learn to obtain the optimal strategy, which causes the static environment to become a non - stationary environment due to the continuously changing strategies of the learning UAVs. In a multi - UAV environment, when UAVs use independent reinforcement learning algorithms to optimize strategies through local action - value functions or value functions, it may be difficult to converge the strategies due to non - stationarity. The MADDPG algorithm is a reinforcement learning algorithm based on a multi - agent environment proposed for such problems.
[0125] The MADDPG algorithm has made a series of improvements based on Actor - Critic and DDPG, and adopts the principle of centralized learning and distributed application, enabling it to be applicable to complex multi - agent environments that cannot be handled by traditional reinforcement learning algorithms. Traditional reinforcement learning algorithms must use the same information data during learning and application, while the MADDPG algorithm allows using some additional information (i.e., global information) during learning, but only uses local information when making application decisions. Compared with the traditional Actor - Critic algorithm, there are n agents in the MADDPG algorithm environment. The strategy of the i - th agent is represented by π i and its strategy parameter is θ i , then the joint strategy set of the n agents can be obtained as π = π1, π2, …, π n , and the strategy parameter set is θ = θ1, θ2, …, θ n . Its core idea is to find the optimal joint strategy through a framework of centralized training and distributed execution, which can solve the non - stationarity of the multi - agent reinforcement learning environment and the problem of the failure of the experience replay method.
[0126] The experience pool of the MADDPG algorithm is designed as where is the set of observations of all UAVs at time t, is the set of actions of all UAVs at time t, is the reward obtained by all UAVs after executing their respective actions at time t, is the set of observations of all UAVs at time t + 1.
[0127] The so-called "centralized training, decentralized execution" means centralized training and decentralized execution, that is, the optimal policy obtained through training and learning only needs to use the observation information of the UAV - local information to output the optimal action when applied. During centralized training, some additional information is superimposed on the basic DDPG algorithm to obtain a more accurate Q - value calculation, which is fed back to the Actor network. These values can be the states and actions of other UAVs. Each UAV evaluates the value of the action output by the current Actor network not only based on its own observations and actions but also based on the actions of other UAVs. The Q - value calculation is as follows:
[0128]
[0129] In the formula, θ Q is the parameter of the online critic network, and Q(s t , a1, a2, …, a n |θ Q ) is the centralized state - action function, which contains not only the state observed by itself and the actions it executes but also the actions of other UAVs (a1, a2, …, a n ). When executing the action t in the state s the reward from the environment will be superimposed on Q.
[0130] The input of the online critic network of each UAV is the same and is updated by minimizing the loss function, which is equivalent to establishing a centralized Critic network. Its loss function is:
[0131] where N is the number of training times.
[0132] is the centralized state - action function, which contains not only the state observed by itself and the actions it executes but also the actions of other UAVs Therefore, the Critic network of each UAV not only knows the changes of its own UAV but also the action strategies of all other UAVs. In such a situation, even if the strategy is constantly updated and changed, the environment can still be considered stable. Because even when π i ≠π′i When this is the case, the following equation still holds, which is the reason why MADDPG can solve the environmental instability.
[0133]
[0134] In the formula, P is the dynamic model of the multi-agent system, and π i ′ represents an arbitrary policy different from π i , and i = 1, 2,..., n.
[0135] The online Actor network update policy gradient is:
[0136]
[0137] In the formula, N is the number of training times, Q(s, a|θ μ ) is the online Actor network, θ μ is the parameter of the online Actor network, μ(s|θ μ ) is the target Actor network, represents gradient calculation, μ(s i ) is the online actor network in the state si, and J represents the policy gradient.
[0138] Decentralized execution means that after training is completed, each Actor can take appropriate actions according to its own observations without the actions of other UAVs. The Actor network and the Critic network in the MADDPG algorithm cooperate with each other. Each UAV uses its own Actor to output a definite action. However, in addition to its own observed state information and action information, the input of the Critic network also includes the action information of other UAVs. Each UAV corresponds to a centralized Critic network, which simultaneously accepts the data generated by the Actor networks of all UAVs.
[0139] The MADDPG algorithm framework is as Figure 5As shown, from the overall framework of the algorithm, for a single UAV, first its state is input into its own policy network, and after obtaining an action, it is output and applied to the environment. At this time, a new state and a reward value will be obtained. Finally, the state transition data is stored in the UAV's own experience pool. All UAVs will continuously interact with the environment, continuously generate data and store it in their respective experience pools. During the process of updating the network, a batch of data at the same moment is randomly taken out from the experience pool of each UAV, and they are concatenated to obtain new experience (S, A, S′, R). Among them, S and S′ are the observation values of all UAVs at the same moment, s is the state information observed by the agent from the environment at a certain moment, s′ is the state information observed by the agent from the environment after executing action a, A is the set of actions made by all UAVs at the same moment, and R selects the reward value of the i-th UAV. Finally, the observation value S′ is input into the target Actor network of the i-th UAV to obtain action A′, and then action A′ and observation value S′ are input into the target Critic network of the i-th UAV together to obtain the estimated target Q value for the next moment, and the target Q value at the current moment is calculated according to the following formula.
[0140] y i = r i + γQ′(s i+1 , μ′(s i+1 |θ μ′ )|θ Q′ ).
[0141] In the formula, μ′ = [μ′1, μ′2, …, μ′ n , represents the Actor network of the i-th UAV after executing action a, Q’ is the Q value after executing action a, θ μ′ is the target Actor network parameter, θ Q′ is the target Critic network parameter, γ is the discount factor, indicating the importance of the current feedback. The longer the time, the smaller the impact.
[0142] S6. Achieve the safe flight of multi-UAV formations in an obstacle environment based on the optimal transformation strategy. Based on the implementation process of step S5, in step S6, the reinforcement learning of UAVs completes the learning task under the guidance of the reward function. The reward function will guide the UAVs to interact with the environment according to the task requirements and indicate the UAVs to have a priority ranking for the learning perception of the environment. The selection of the reward function usually requires a large number of experiments and trials. Selecting an inappropriate form of the reward function may lead to unexpected problems, resulting in the UAVs learning bad solutions. The multi-UAV learning environment will face a more complex situation compared to the single-UAV learning environment. Therefore, in order to enable multi-UAVs to complete the task requirements, its reward function also needs to be more carefully tried and selected. Generally speaking, the reward function of multi-UAV reinforcement learning should consider both the behavior of individual UAVs and the cooperative behavior among UAVs, and this cooperative behavior also needs to be reflected through the reward function.
[0143] This invention conducts research on the path planning of UAV clusters based on a two-dimensional dynamic obstacle environment. There are n UAVs in the cluster, and the state S of each UAV itself uavi includes the velocity vector at the current moment and the position coordinates in the environment The environmental state S env includes the distance and velocity information (d, ψ d , θ d , v, ψ v , θ v ) of all j dynamic and static obstacles in the environment relative to the UAVs. Among them, d is the Euclidean distance of the obstacle relative to the UAV. ψ d is the relative distance heading angle. θ d is the relative distance climb angle. v is the movement speed of the obstacle relative to the UAV. ψ v is the relative movement speed heading angle. θ v is the relative movement speed climb angle.
[0144] In the MADDPG algorithm, the state of each UAV includes its own state, the states of other UAVs, and the environmental state. The state of UAV1 at time t can be defined as:
[0145]
[0146] Finally, the network input of each UAV is:
[0147]
[0148] The method of training different multi-UAV formation models mainly relies on the differences in the reward functions. The reward value function R consists of four parts. The first is to guide the UAVs to fly with a stable attitude and speed, using R singleIndication. Second, it is to guide how the UAV and other UAVs fly in coordination and maintain a certain formation distance, denoted by R form Indication. Third, it is to guide the UAV to avoid obstacles, denoted by R obstacle Indication. Fourth, it is to guide the UAV to determine the formation transformation method with the minimum cost, denoted by R(F start ,F end ), and the reward function is set as shown in the following formula:
[0149]
[0150] R form =-|d ij -d safe |.
[0151]
[0152] R(F start ,F end )=H(F start ,F end )·α.
[0153] R=α s R single +α f R form +α o R obstacle +α FF R(F start ,F end ).
[0154] In the formula, V = [v, ψ, θ], d ij represents the distance between the i-th UAV and the j-th UAV, d safe represents the safety distance of the formation, dist ij represents the distance between the i-th UAV and the j-th obstacle, d det represents the detection distance of the UAV, α s , α f , α o , α FF are constants, and α s +α f +α o +α FF =1.
[0155] In addition, during the process of the multi-unmanned aerial vehicle path planning method provided above in the present invention, it is also necessary to describe the UAV cluster problem, specifically including:
[0156] Step 1, task modeling
[0157] Suppose there is a multi-UAV formation consisting of n UAVs flying in a certain initial formation in a complex obstacle environment. During the flight, after receiving the instruction of formation transformation, the UAV formation completes the formation transformation in the shortest time and realizes the optimal position selection of each UAV. For the convenience of research, the following assumptions are given:
[0158] (1) In the entire UAV formation, each UAV can obtain the positions and headings of other UAVs in real time through sensors.
[0159] (2) When the formation transmits information, communication delay and packet loss problems are not considered.
[0160] (3) The aerodynamic influence between UAVs during the formation transformation process is not considered.
[0161] Step 2: Motion model
[0162] Regarding the UAVs in the formation transformation problem as a particle motion model, the acceleration and heading angle of the UAV are used to control the motion process of the UAV. The motion equation of the UAV can be expressed as:
[0163]
[0164] In the formula: i = 1, 2,..., n, where n is the number of UAVs, v i represents the velocity of the i-th UAV in the XOY plane, ψ is the heading angle of the UAV, and a represents the acceleration of the UAV. x i is the position of the i-th UAV in the x direction, y i is the position of the i-th UAV in the y direction, is the velocity of the i-th UAV in the x direction, is the velocity of the i-th UAV in the y direction.
[0165] Considering the saturation constraint of the control input, the acceleration a and heading angle ψ of the UAV satisfy the following conditions:
[0166]
[0167] In the formula, the specific constraint parameters of the acceleration depend on the UAV model and flight parameters.
[0168] Next, the obstacle avoidance performance of the proposed UAV cluster formation transformation strategy is tested in a simulation environment to illustrate the effectiveness of the multi-unmanned aerial vehicle path planning method provided above in the present invention.
[0169] The simulation experiment environment is Python 3.7, an Inter Core i5 processor with a main frequency of 2.42 GHz, and a Windows 10 operating system. The test scenario is as Figure 6As shown, the simulation environment range is set to 500m × 500m.
[0170] In the experiment, four UAVs start from the initial positions of (57.30m, 44.04m), (11.82m, 21.95m), (263.83m, 29.89m), and (91.10m, 12.34m) respectively. During the flight, they gradually form a target polygon formation and always maintain the formation flight until an obstacle area is detected within the safe range of the formation. Then, according to the optimal formation transformation strategy designed in step S5 above, they first perform a stretching transformation, and then bypass from both sides of the obstacle according to the segmentation strategy to complete the structural transformation. After safely passing through the obstacle area, they resume the original polygon formation flight until they reach the end point.
[0171] According to the formation transformation evaluation criteria established in step S3, in order to improve the stability of the formation maintenance as much as possible, the weights of the formation structure difference degree, the time cost of formation transformation, and the path cost of formation transformation are set to 0.4, 0.4, and 0.2 respectively. At the 80.4s of the simulation, UAV2 detects that the distance to obstacle 1 is only 26.1m at this time, while the set safe distance of the UAV formation is 30m. Therefore, the formation needs to adjust the formation accordingly. The coordinates of the three obstacles (i.e., obstacle 1, obstacle 2, and obstacle 3) set in the simulation environment are (84m, 180m), (200m, 220m), and (110m, 280m) respectively, and the diameters are 36m, 65m, and 60m respectively. The maximum passable width between obstacle 1 and obstacle 2 is 72.2m, while the formation width of the currently maintained quadrilateral formation is 109.9m, and the formation cannot directly pass through the obstacle area. After calculation, when the formation transformation factor takes the minimum formation transformation factor of the current formation, the formation width can be minimized to 46.9m. Therefore, at this time, only the formation expansion parameter b of the formation needs to be adjusted accordingly to safely pass through the obstacle area, and no formation structural transformation is required. However, since the maximum passable width between obstacle 1 and obstacle 3 is 55.3m, and the maximum passable width between obstacle 2 and obstacle 3 is only 45.7m, which is less than the minimum formation width under the current quadrilateral formation, a formation structural transformation is required to divide the original formation into two sub-formations and bypass from both sides of obstacle 3 respectively. When it moves to the 125.4s, the UAV formation exits the obstacle area and resumes the original polygon formation. The UAV flight trajectory is as Figure 6 shown.
[0172] The simulation uses 300ms as a sampling period, Figures 7 - 10 which is the position information of each UAV in the formation at each sampling moment, where Figure 7 is the distance curve between each UAV, Figures 8 - 10 is the distance curve between each UAV and the three obstacles respectively. From Figures 7 - 10It can be seen that the UAV formation initially formed a quadrilateral formation at time T1. After detecting the presence of an obstacle within a safe distance at time T2, the formation underwent a telescopic transformation, and the distance between the UAVs in the formation decreased relatively. Subsequently, the formation was divided during the T4 time period, and 4 UAVs passed through from both sides of the obstacle. After the UAV formation successfully passed through the obstacle area, the UAV formation regrouped and restored the original polygon formation.
[0173] In addition, corresponding to the above-provided multi-UAV path planning method, the present invention also provides the following implementation structure:
[0174] A multi-UAV path planning system for implementing the above-provided multi-UAV path planning method. The system includes:
[0175] A formation library construction module for describing the geometric formation of the UAV cluster in the form of a formation matrix and establishing an adaptive formation library based on the described geometric formation.
[0176] A dynamic transformation determination module for determining the formation dynamic transformation in the case of obstacle avoidance based on the adaptive formation library.
[0177] An evaluation criterion establishment module for establishing a formation transformation evaluation criterion for environmental constraints and UAV cluster trajectory constraints based on the formation dynamic transformation.
[0178] An evaluation function determination module for constructing a formation transformation evaluation function based on the formation transformation evaluation criterion.
[0179] An optimal transformation strategy selection module for selecting the optimal transformation strategy of the formation in the case of obstacle avoidance based on the formation transformation evaluation function. The optimal transformation strategy is generated using the MADDPG algorithm. The MADDPG algorithm is an algorithm obtained by improving the Actor-Critic algorithm and the DDPG algorithm.
[0180] A safe flight control module for realizing the safe flight of the multi-UAV formation in an obstacle environment based on the optimal transformation strategy.
[0181] An electronic device, including:
[0182] A memory for storing a control program.
[0183] A processor connected to the memory for retrieving and executing the control program to implement the above-provided multi-UAV path planning method.
[0184] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.
[0185] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A path planning method for multiple unmanned aerial vehicles, characterized in that Including: Describing the geometric formation of the UAV cluster in the form of a formation matrix, and establishing an adaptive formation library based on the described geometric formation; Determining the formation dynamic transformation under obstacle avoidance based on the adaptive formation library; Establishing a formation transformation evaluation criterion for environmental constraints and UAV cluster trajectory constraints based on the formation dynamic transformation, including: constructing constraint conditions; determining the formation structure difference degree based on the formation geometric parameters corresponding to the current formation matrix and the formation geometric parameters corresponding to the transformed formation matrix; determining the formation transformation convergence time and the formation transformation path cost; under the constraint conditions, determining an evaluation vector of a formation transformation based on the formation structure difference degree, the formation transformation convergence time, and the formation transformation path cost; using the evaluation vector of the formation transformation as the formation transformation evaluation criterion; Constructing a formation transformation evaluation function based on the formation transformation evaluation criterion; Selecting an optimal transformation strategy for the formation under obstacle avoidance based on the formation transformation evaluation function; the optimal transformation strategy is generated using the MADDPG algorithm; the MADDPG algorithm is an algorithm obtained by improving the Actor-Critic algorithm and the DDPG algorithm; wherein, determining a formation transformation factor based on the maximum length of the obstacle area and the formation width of the current UAV cluster; selecting an optimal transformation strategy according to the relationship between the formation transformation factor and a preset transformation threshold range; Implementing the safe flight of the multi-UAV formation in an obstacle environment based on the optimal transformation strategy; 2. The multi-unmanned aerial vehicle path planning method according to claim 1, wherein The formation matrix is F : ; wherein, is the formation expansion parameter, is the geometric configuration parameter, is the angle parameter of the formation, n is the total number of UAVs in the cluster.
3. The multi-unmanned aerial vehicle path planning method according to claim 1, wherein The formation transformation evaluation function is: ; In the formula, R ( F start , F end ) is the evaluation result of a formation transformation, H ( F start , F end ) is the evaluation vector of a formation transformation, α is a column vector, is the current formation matrix, is the formation matrix after transformation.
4. The multi-unmanned aerial vehicle path planning method according to claim 1, wherein The selecting of the optimal transformation strategy according to the relationship between the formation transformation factor and the preset transformation threshold range specifically includes: When the formation transformation factor is greater than the maximum preset transformation threshold, it is not necessary to perform a formation transformation to pass through this obstacle area; When the formation transformation factor is greater than the minimum preset transformation threshold and less than the maximum preset transformation threshold, the UAV cluster performs a formation stretching transformation; When the formation transformation factor is less than the minimum preset transformation threshold, the UAV cluster performs a formation structural transformation; 5. The multi-unmanned aerial vehicle path planning method according to claim 4, wherein The process of the UAV cluster performing a formation stretching transformation is: The leader in the cluster determines the formation matrix to be finally transformed according to the formation transformation factor, and then sends the formation matrix to be finally transformed to other UAVs for formation stretching transformation; 6. The multi-unmanned aerial vehicle path planning method according to claim 4, wherein The process of the UAV cluster performing a formation structural transformation is: Traversing the formation library to determine the formation transformation evaluation function corresponding to each different geometric formation; Selecting the geometric formation with the minimum formation transformation evaluation function as the formation after transformation, and performing a formation structural transformation; Wherein, when the UAV cluster encounters an obstacle and cannot bypass from the same side, according to the segmentation strategy, the cluster is segmented into two sub-formations to bypass from both sides of the obstacle respectively, and then restored to the original formation after crossing the obstacle; 7. A multi-unmanned aerial vehicle path planning system, characterized in that, For implementing the multi-unmanned aircraft path planning method according to any one of claims 1-6; the system includes: A formation library construction module, configured to describe the geometric formation of the UAV cluster in the form of a formation matrix, and establish an adaptive formation library based on the described geometric formation; A dynamic transformation determination module, configured to determine the formation dynamic transformation in the case of obstacle avoidance based on the adaptive formation formation library; An evaluation criterion establishment module, configured to establish a formation transformation evaluation criterion for environmental constraints and UAV swarm trajectory constraints based on the formation dynamic transformation; An evaluation function determination module, configured to construct a formation transformation evaluation function based on the formation transformation evaluation criterion; An optimal transformation strategy selection module, configured to select an optimal transformation strategy for the formation in the case of obstacle avoidance based on the formation transformation evaluation function; the optimal transformation strategy is generated by using the MADDPG algorithm; the MADDPG algorithm is an algorithm obtained by improving the Actor-Critic algorithm and the DDPG algorithm; A safe flight control module, configured to implement the safe flight of the multi-UAV formation in the obstacle environment based on the optimal transformation strategy.
8. An electronic device, characterized in that, It includes: A memory, configured to store a control program; A processor, connected to the memory, configured to retrieve and execute the control program to implement the multi-unmanned aerial vehicle path planning method according to any one of claims 1-6.
Citation Information
Patent Citations
UAV (unmanned aerial vehicle) cluster formation flight method based on behavior control
CN110502032A
Multi-unmanned aerial vehicle formation cluster control method based on multi-agent deep reinforcement learning
CN115755949A