A Multi-UAV Motion Planning Method and System in a Complex Environment
Through the combination of unit filtering and Q-mix network, the efficient collision-free problem of multi-UAV motion planning in complex environments is solved, and safe and efficient path planning is achieved in the case of limited field of vision and uncertain perception information.
Patent Information
- Application Number
- CN202310186719.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-01
AI Technical Summary
Existing multi-UAV motion planning methods are difficult to achieve efficient collision-free path planning in complex environments, especially when the field of view is limited and the perception information is uncertain.
The multi-unmanned aerial vehicle environment model is constructed using the crew filtering method, the belief status of each drone is estimated, and input it into the motion planning network based on the Q-mix network to obtain the actions of each drone to achieve efficient collision-free path planning.
Implement efficient collision-free path planning for multiple drones in complex environments, improving the accuracy and safety of motion planning, and is suitable for uncertain and partially observable environments.
Smart Images

Figure CN116360484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) motion planning, and particularly to a multi-UAV motion planning method and system in a complex environment. Background Art
[0002] In recent years, with the cross-disciplinary and permeating development of multiple disciplines such as artificial intelligence, computer science, and control science, multi-UAV systems integrating autonomous perception, intelligent decision-making, and automatic control have become a hot spot in the current research on multi-disciplinary cross-integration. Their application fields are gradually moving from video games and simulation to real-world applications such as intelligent unmanned driving, robot control, logistics warehousing, and UAV collaboration. And multi-UAV motion planning technology is an essential part of the application research of multi-UAV systems.
[0003] The multi-UAV motion planning problem is a type of problem that finds an optimal path set without conflicts for multiple UAVs to reach the target position from the starting position. How to make UAVs cooperate with other UAVs to avoid obstacles and safely and efficiently reach the designated area has become a major research problem. Currently, to solve the multi-UAV motion planning problem, various methods such as optimization algorithms based on control theory and search algorithms based on geometry have been proposed, which meet the requirements of multi-UAV motion planning to a certain extent. However, these methods are often prone to falling into local optima, difficult to quickly obtain numerical solutions, and cannot be used for large-scale collaborative tasks.
[0004] With the development of deep learning, neural networks have brought new vitality to reinforcement learning. And multi-UAV reinforcement learning optimizes the decision-making process through rewards and punishments, with the characteristics of autonomous learning and predictive learning. Representative distributed multi-UAV reinforcement learning algorithms such as the Q-mix algorithm and the Multi-Agent Deep Deterministic Policy Gradient (DDPG) algorithm can be applied to environments where UAVs can only perceive partial information similar to the real world, and have gradually become a research hot spot in multi-UAV motion planning.
[0005] However, the current application research of multi-UAV reinforcement learning algorithms in autonomous driving and robot control is still limited to simulation platforms, and there are not many application examples in the real environment. The key to the problem is that in the real environment, UAVs not only cannot perceive complete environmental information, but also the perceived environmental information is inaccurate due to factors such as sensor quality and environmental noise. And incorrect perceived information may cause decision-making errors of UAVs, which are likely to cause potential safety hazards in practical applications. Therefore, it has great practical significance to design an efficient collision-free path planning method that can adapt to complex environments for multi-mobile UAVs with limited vision and uncertain perceived information. Summary of the Invention
[0006] Based on this, an embodiment of the present invention provides a multi - UAV motion planning method and system in a complex environment, and plans an efficient collision - free path adapted to the complex environment for multiple UAVs.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A multi - UAV motion planning method in a complex environment, comprising:
[0009] Construct a multi - UAV environment model at time t; the multi - UAV environment model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model; t≥0;
[0010] According to the multi - UAV environment model at time t, use the set - theoretic filtering method to estimate the belief state of each UAV at time t.
[0011] Input the belief states of all UAVs at time t into the UAV decision network to obtain the actions of each UAV at time t; the UAV decision network is constructed based on the Q - mix network.
[0012] Optionally, the method for determining the UAV decision network is:
[0013] Construct a Q - mix network; the Q - mix network includes multiple Deep Recurrent Q - Learning Networks (DRQNs) and a mixing network; one Deep Recurrent Q - Learning Network corresponds to one UAV; each of the Deep Recurrent Q - Learning Networks is connected to the mixing network;
[0014] Use the reinforcement learning method to train and test the Q - mix network to obtain the UAV decision network.
[0015] Optionally, using the reinforcement learning method to train and test the Q - mix network to obtain the UAV decision network specifically includes:
[0016] Input the belief states of each UAV into the corresponding Deep Recurrent Q - Learning Network. Each of the Deep Recurrent Q - Learning Networks selects the actions of each UAV based on the ε - greedy policy and outputs the action values of each UAV. The action values of each UAV are used as the input of the mixing network, and the Q - mix network is jointly trained with the goal of minimizing the joint action loss to obtain a trained Q - mix network;
[0017] Test the trained Q - mix network and determine the tested Q - mix network as the UAV decision network;
[0018] Among them, the action value is used to determine the selected action; the joint action loss is determined according to the joint action value function; the joint action value function is determined according to the action values and weight values of each UAV; the weight value is obtained by processing the global belief state using a hypernetwork; the global belief state is determined according to the belief states of all UAVs.
[0019] Optionally, according to the multi-UAV environment model at time t, the membership filtering method is used to estimate the belief state of each UAV at time t, specifically including:
[0020] According to the multi-UAV environment model at time t, a belief estimation network is constructed using the membership filtering method; the belief estimation network includes a prediction network and an observation network;
[0021] The belief estimation network is used to determine the state estimation value of each UAV at time t;
[0022] According to the state estimation value at time t and the shape matrix at time t, the belief state of each UAV at time t is determined.
[0023] Optionally, the calculation formula for the state estimation value at time t is:
[0024]
[0025] Among them, is the state estimation value of UAV i at time t output by the prediction network; is the undetermined parameter of UAV i at time t; is the state observed by UAV i at time t; is the state estimation value of UAV i at time t output by the observation network; is the observation noise at time t; is the time-varying matrix of the observation noise of UAV i at time t; λ is an undetermined parameter, λ ∈ [0, 1]; is the state estimation value of UAV i at time t after filtering; is the true state of UAV i at time t; is the shape matrix of UAV i at time t; T is the transpose; ε is the symbol representing the region; represents the set range value corresponding to UAV i at time t; z represents the variable; represents the region containing the true state.
[0026] Optionally, the environment model is constructed based on the physical characteristics of the environment; the physical characteristics of the environment include the actual physical characteristics of the UAV, the actual physical characteristics of the obstacle, and the actual physical characteristics of the target.
[0027] The present invention also provides a multi - UAV motion planning system, including:
[0028] An environmental model construction module, configured to construct a multi - UAV environmental model at time t; the multi - UAV environmental model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model; t≥0;
[0029] A belief state estimation module, configured to estimate the belief state of each UAV at time t by using a set - membership filtering method according to the multi - UAV environmental model at time t;
[0030] An action planning module, configured to input the belief states of all UAVs at time t into a UAV decision network to obtain the actions of each UAV at time t; the UAV decision network is constructed based on the Q - mix network.
[0031] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0032] The embodiments of the present invention propose a multi - UAV motion planning method and system in a complex environment. The set - membership filtering method is used to estimate the belief state of each UAV; the estimated belief state is input into a motion planning network constructed based on the Q - mix network to obtain the actions of each UAV, realizing multi - UAV motion planning. The present invention combines the set - membership filtering method with the Q - mix network, and can achieve efficient collision - free path planning adapted to complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 It is a flowchart of the multi - UAV motion planning method provided by the embodiments of the present invention;
[0035] Figure 2 It is a schematic diagram of multi - UAV motion planning;
[0036] Figure 3 It is a training network framework diagram of a single UAV;
[0037] Figure 4 It is a joint training network framework diagram of multi - UAVs;
[0038] Figure 5 It is a structural diagram of the multi - UAV motion planning system provided by the embodiments of the present invention. Detailed implementation manners
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific implementation manners.
[0041] Embodiment 1
[0042] See Figure 1 , the multi-UAV motion planning method of this embodiment includes:
[0043] Step 101: Construct a multi-UAV environment model at time t; the multi-UAV environment model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model. Wherein, t≥0.
[0044] The environment model is constructed based on the physical characteristics of the environment; the physical characteristics of the environment include the actual physical characteristics of the UAVs, the actual physical characteristics of the obstacles, and the actual physical characteristics of the targets.
[0045] Step 102: According to the multi-UAV environment model at time t, use the set-membership filtering method to estimate the belief state of each UAV at time t + 1.
[0046] Step 103: Input the belief states of all UAVs at time t into the UAV decision network to obtain the actions of each UAV at time t; the motion planning network is constructed based on the Q-mix network.
[0047] In one example, the determination method of the motion planning network in step 103 is as follows:
[0048] 1) Construct a Q-mix network; the Q-mix network includes a plurality of deep recurrent Q-learning networks and a mixing network; one deep recurrent Q-learning network corresponds to one UAV; each of the deep recurrent Q-learning networks is connected to the mixing network. The deep recurrent Q-learning network includes a sequentially connected input fully connected layer, a recurrent neural network, and an output fully connected layer.
[0049] 2) Use the reinforcement learning method to train and test the Q-mix network to obtain the UAV decision network. Specifically:
[0050] Input the belief states of each drone into the corresponding deep recurrent Q-learning network. Each of the deep recurrent Q-learning networks selects the actions of each drone based on the ε-greedy policy and outputs the action values of each drone. The action values of each drone are used as the input to the hybrid network, and the Q-mix network is jointly trained with the goal of minimizing the joint action loss to obtain a trained Q-mix network.
[0051] Test the trained Q-mix network and determine the tested Q-mix network as the drone decision network.
[0052] Among them, the action value is used to determine the selected action; the joint action loss is determined according to the joint action value function; the joint action value function is determined according to the action values and weight values of each drone; the weight value is obtained by processing the global belief state using a hypernetwork; the global belief state is determined according to the belief states of all drones.
[0053] In one example, step 102 specifically includes:
[0054] 1) According to the multi-drone environment model at time t, use the set membership filtering method to construct a belief estimation network; the belief estimation network includes a prediction network and an observation network. The filtering process is the process of fusing prediction information and observation information.
[0055] 2) Use the belief estimation network to determine the state estimation values of each drone at time t.
[0056] The calculation formula for the state estimation value at time t is:
[0057]
[0058] Among them, is the state estimation value of drone i at time t output by the prediction network; is the undetermined parameter of drone i at time t; is the state observed by drone i at time t; is the state estimation value of drone i at time t output by the observation network; is the observation noise at time t; is the time-varying matrix of the observation noise of drone i at time t; λ is an undetermined parameter, λ ∈ [0, 1]; is the state estimation value of drone i at time t after filtering; is the true state of drone i at time t; P t i is the shape matrix of drone i at time t; T is the transpose; ε is the symbol representing the region; denote the set range value corresponding to UAV i at time t; z represents a variable; denote the area containing the true state.
[0059] 3) Determine the belief state of each UAV at time t according to the state estimation value at time t and the shape matrix at time t.
[0060] The purpose of the present invention is to provide a state estimation method based on set - membership belief for the multi - UAV motion planning problem in an uncertain environment, introducing a multi - UAV reinforcement learning algorithm to help UAVs make real - time planning decisions, so as to complete the tasks of efficient obstacle avoidance and path planning in a complex environment with maximum reward. The present invention can efficiently train a safe and stable motion planning strategy in a complex, uncertain and partially observable environment.
[0061] In practical applications, a more specific implementation process of the above - mentioned multi - UAV motion planning method is as follows:
[0062] Step 1: Establish a distributed partially observable Markov decision model (DEC - POMDPs) for multi - UAV motion planning.
[0063] Use DEC - POMDPs to model the environment, which is described by the seven - tuple <N, S, A, T, O, Z, R>, where N = {1, 2,..., n}, representing a set of n UAVs, S represents the joint state, A represents the joint action, T is the state transition model, O is the joint observation, Z is the observation model, and R represents the joint reward. Taking UAVs as an example, the motion planning of multi - UAVs is as Figure 2 shown.
[0064] Step 1 - 1: Create the physical environment for UAV flight.
[0065] Select a fixed point as the origin, establish a world coordinate system with the due - east direction as the positive direction, and construct physical models of UAVs, obstacles, targets, etc.
[0066] Step 1 - 2: Set the joint state of UAVs.
[0067] The joint state space of multi - UAVs is represented by S, which is the Cartesian product of the state spaces of all UAVs, that is, S = S 1 ×S 2 ×…×S n , where S n is the local state space of the i - th UAV.
[0068] At any time t, the state of UAV i is represented as where The own state information of the UAV, including the position, speed, angular velocity of the UAV, the angle between the forward direction of the UAV and the target direction, and the distances between the UAV and the target, the nearest teammate, and the nearest obstacle; The state information of the nearest teammate that the UAV can obtain through communication.
[0069] In the actual flight of the UAV, affected by factors such as sensor noise and environmental interference, it is generally impossible to obtain accurate UAV state values.
[0070] Step 1-3: Set the joint actions of the UAVs.
[0071] The joint action space of multiple UAVs is denoted as A, which is the Cartesian product of the action spaces of all UAVs, and A = A 1 ×A 2 ×…×A n , where A n represents the action space of the i-th UAV.
[0072] The action of UAV i at time t includes the set of velocities that it can allow in the continuous space, including the linear velocity and the angular velocity, that is
[0073] Step 1-4: Set the joint observations of the UAVs.
[0074] The observation space of the UAV is defined as the observation of the UAV state. The joint observation space of multiple UAVs is denoted as O, which is the Cartesian product of the observation spaces of all UAVs, that is O = O 1 ×O 2 ×…×O n , where O i is the local observation space of the i-th UAV.
[0075] The observation of UAV i at time t is denoted as It contains the real but inaccurate observation information such as the position, speed, direction of the UAV and the distances to obstacles / teammates obtained by the UAV position sensor (GPS), speed sensor (tachometer), direction sensor (gyroscope), and ranging sensor (radar) at the current moment.
[0076] Step 1-5: Set the state transition model of the UAVs.
[0077] The state transition model of multiple UAVs is f(·) represents the transition function for all UAVs to reach the next state by taking joint actions in the current state, It represents the disturbance noise during the movement of the unmanned aerial vehicle. The state transition model represents a functional relationship in which an agent transfers from state s to state s' after executing an action a, and there is noise in the state transition process.
[0078] Step 1-6: Set the observation model of the unmanned aerial vehicle.
[0079] The observation model of multiple unmanned aerial vehicles is which represents the relationship between the observation of the unmanned aerial vehicle and its potential state; g(·) represents the transition function for all unmanned aerial vehicles to reach the next state by taking joint actions in the current state, represents the observation noise of the unmanned aerial vehicle, and o represents the observation.
[0080] Step 1-7: Set the joint reward of multiple unmanned aerial vehicles.
[0081] Set the joint reward of multiple unmanned aerial vehicles as R, expressed as R = R 1 ×R 2 ×…×R n , that is, the Cartesian product of the reward functions of all unmanned aerial vehicles, where R i represents the reward value obtained by unmanned aerial vehicle i after interacting with the environment and realizing state transition.
[0082] The reward function of unmanned aerial vehicle i is set as follows:
[0083] Let the reward function of unmanned aerial vehicle i when it reaches the target be where, R i_goal represents the reward obtained by unmanned aerial vehicle i at the target point, represents the penalty value for the time consumed by the unmanned aerial vehicle to reach the target, W T represents the parameter value of the penalty degree, T i is the time actually consumed by the unmanned aerial vehicle to reach the target, represents the time consumed when the unmanned aerial vehicle moves from the starting point to the target position at a uniform speed along a straight line.
[0084] The reward function of unmanned aerial vehicle i when it collides is R i_collision .
[0085] When unmanned aerial vehicle i does not reach the target or does not collide, set 5 non-sparse reward functions, and the specific expressions are as follows:
[0086]
[0087] where, and are the distance values of unmanned aerial vehicle i to the target at the initial moment and at time t, respectively; represents the angle between the positive direction of unmanned aerial vehicle i and the target direction at time t; and respectively represent the distances between the UAV i and the obstacle k and the teammate j observed at time t; and respectively represent the distances from the agent i to the target point at time t-1 and time t; P t i represents the distance of the agent i at time t; P t j represents the distance of the agent j at time t; d ik and d ij are respectively the minimum safety distances between the UAV i and the obstacle k and the teammate j; then the reward function when the UAV i moves normally is expressed as R i = η1N(R1)+η2N(R2)+η3N(R3)+η4N(R4)+η5N(R5), where N(R) represents the normalization process of the reward function, η1, η2, η3, η4 represent the contribution rates of the 5 reward functions, and η1+η2+η3+η4+η5 = 1.
[0088] Step 2: Construct the motion planning network structure of each UAV under uncertain perception.
[0089] Based on the DEC-POMDPs model constructed in Step 1, an independent policy learning network is established for each UAV, which is jointly composed of a belief estimation network based on set membership filtering and a deep recurrent Q-learning network.
[0090] See Figure 3 , and the design of the belief estimation network based on set membership filtering is as follows:
[0091] First, fit the state transition function of the UAV and the observation function by a fully connected deep neural network, and name them the prediction network and the observation network respectively.
[0092] Assume that at time t-1, the true state satisfies the condition That is
[0093]
[0094] where represents the belief state of the UAV i at time t-1, represents the estimated state of the UAV i at time t-1, represents the shape matrix at time t-1. The process noise and the observation noise satisfy and are respectively known time-varying matrices.
[0095] The filter is designed as follows:
[0096]
[0097] where is the state estimation value of UAV i at time t obtained through the prediction network; is a parameter to be determined, is the state estimation value of UAV i at time t obtained through the observation network.
[0098] Furthermore, the state of UAV i at time t can be obtained:
[0099]
[0100] where λ ∈ [0, 1] is a parameter to be determined and is determined by calculating minTr(P t i ), where Tr(·) represents the trace of a matrix; denote the belief state P t i as the shape matrix of UAV i at time t; as the shape matrix of UAV i at time t - 1; represents the predicted value of the shape matrix of UAV i at time t.
[0101] Step 3: Train each UAV using the DRQN network.
[0102] The DRQN network is designed as follows:
[0103] The DRQN network consists of three parts: an input fully connected layer, a recurrent neural network, and an output fully connected layer; the input of the network is the belief state of the UAV obtained through the belief estimation network, and the output of the network is the current action value Q of the UAV. During the interaction with the environment, UAV i will formulate a strategy π with the goal of maximizing the expected discounted reward sum and select an action a according to the principle of the ε - greedy strategy t , and obtain a new observation from the environment, where γ is the discount factor, is the discounted reward of UAV i at time t. Subsequently, a new belief state is calculated through the belief estimation network and the experience is stored in the experience pool Ξ.
[0104] Step 4: Guide each UAV to make decisions using a mixed network.
[0105] First, the action value Q output by the DRQN network in each UAV is passed into the hybrid network. Additionally, the global belief state of the UAVs is processed by the hypernetwork to form weight values and input them into the hybrid network. Then, the hybrid network mixes the partial action value functions with weights into a joint action value function, establishes a loss function based on the joint action value function, and trains the UAV network by minimizing the loss function.
[0106] Step 3-1: Set the values of the training parameters, including the experience pool capacity M, the batch sampling quantity N, the target network update frequency F, the maximum number of training episodes E, and the maximum movement time T of each UAV in each episode max , and initialize the number of training episodes as e = 0.
[0107] Step 3-2: Initialize the initial positions, velocities, angular velocities, directions, and target positions of the UAVs and obstacles, initialize the iteration number k = 0 and the movement time t = 0; initialize the joint state of the UAVs.
[0108] Step 3-3: Set the loss function.
[0109] The loss function is as follows:
[0110]
[0111] where θ p is the evaluation network parameter of the DRQN network, B is the sampling batch for experience replay during training, Q tot represents the joint action value function, is the discounted cumulative return of the l-th batch, τ t is the historical record of the action-observation pairs.
[0112] Step 3-4: Update the evaluation network and target network parameters.
[0113] The update method is as follows:
[0114]
[0115] θ T ′ = βθ T + (1 - β)θ P ;
[0116] where θ p ' is the updated evaluation network parameter, is the learning rate, is the gradient operator; θ T ′ is the updated target network parameter, and β is the network replacement update rate, 0 ≤ β ≤ 1.
[0117] Step 5: Use the motion planning network trained in Step 3 to perform motion planning for multiple UAVs.
[0118] In this embodiment, an environmental space for multi-UAV motion planning is established, including the state, action, and observation spaces of each UAV, as well as the UAV state transition model, observation model, and reward model. A belief estimation network based on set membership filtering and a single-UAV decision network based on DRQN are established. The output of the policy network of each UAV is fed into a hybrid network, which mixes partial action value functions into a joint action value function representing the sum of the independent value functions of each UAV. A loss function is established based on the joint action value function, and the single-UAV policy network is trained by minimizing the loss function. This embodiment helps multi-UAVs perform real-time collaborative planning and action decision-making in an uncertain environment through the introduction of a multi-UAV reinforcement learning algorithm, so as to complete efficient obstacle avoidance and path planning tasks in a complex environment with maximum rewards.
[0119] Taking UAVs as an example, the multi-UAV motion planning method will be described below.
[0120] See Figure 2 , the method includes:
[0121] Step (1): Establish a DEC-POMDPs model for multi-UAV motion planning.
[0122] For the multi-UAV motion planning task, a sequential decision-making model based on DEC-POMDPs needs to be established, that is, the UAV needs to execute corresponding optimal control actions according to its own state at each moment to complete efficient obstacle avoidance and path planning tasks in a complex environment with maximum rewards. Therefore, the physical environment of UAV flight, UAV state, action, observation space, as well as UAV state transition model, observation model, and reward model are established in sequence, which will not be elaborated here.
[0123] Step (2): Design a set membership filter based on the task model established in step (1), and the details will not be elaborated here.
[0124] Step (3): Based on steps (1) and (2), design a multi-UAV joint motion planning network based on multi-UAV reinforcement learning. The decision network of the UAV takes the bounded belief state of the UAV as input, making the decision of the UAV more robust in a noisy environment.
[0125] As shown in part (b) of Figure 4 , the entire network is divided into two parts, namely a deep recurrent Q-learning network and a hybrid network. A deep recurrent Q-learning network serves as the decision network of a UAV, and the output of the decision network is the current action value Q of each UAV. The hybrid network outputs the overall final action value Q of the multi-UAV. total .
[0126] The decision network structure of each drone is as shown in part (c) of Figure 4 : The network includes a belief estimation network, a recurrent neural network, and a fully connected layer. The fully connected layer includes an input fully connected layer and an output fully connected layer. The input of the network is the current observation of the drone The output of the network is the action value Q of the drone at the current time i . When the input of the network is the current observation of the nth drone , the output of the network is the action value Q of the nth drone at the current time n (τ,a). The current action values of the first to the nth drones can be abbreviated as Q1...Q n . The belief state of the drone at time t is b t . During the interaction with the environment, the drone will select actions according to the ε-greedy policy principle and obtain new observations from the environment. Subsequently, a new belief state b is calculated through the belief estimation network t+1 , and the experience (b t ,a t ,r t ,b t+1 ) is stored in the experience pool Ξ.
[0127] The hybrid network is as shown in part (a) of Figure 4 : First, the action value Q output by the decision network of each drone is input into the hybrid network. In addition, the global belief state of the drone is processed by a hypernetwork to form weight values and input into the hybrid network. Then, the hybrid network mixes the partial action value functions with weights into a joint action value function, establishes a loss function based on the joint action value function, and trains the drone network by minimizing the loss function.
[0128] Step (4): Use the motion planning network trained in step (3) to perform motion planning for multiple drones.
[0129] The multi-drone motion planning method provided in this embodiment introduces a multi-drone reinforcement learning algorithm to help drones perform real-time planning decisions, so as to complete the tasks of efficient obstacle avoidance and path planning in complex environments with maximum rewards. When the state information observed by the drone has the characteristics of partial observability or uncertainty, that is, the drone can only observe limited information with bounded noise within a limited range, the above method is still effective.
[0130] Embodiment 2
[0131] To execute the method corresponding to the above Embodiment 1 to achieve the corresponding functions and technical effects, a multi-drone motion planning system is provided below.
[0132] See Figure 5 , the system includes:
[0133] The environmental model construction module 501 is used to construct a multi-UAV environmental model at time t; the multi-UAV environmental model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model; t≥0.
[0134] The belief state estimation module 502 is used to estimate the belief state of each UAV at time t by using the set membership filtering method according to the multi-UAV environmental model at time t.
[0135] The action planning module 503 is used to input the belief state of all UAVs at time t into the UAV decision network to obtain the actions of each UAV at time t; the UAV decision network is constructed based on the Q-mix network.
[0136] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0137] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-UAV motion planning method in a complex environment, characterized in that, including: Constructing a multi - UAV environment model at time t; The multi - UAV environment model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model; t≥0; According to the multi - UAV environment model at time t, using the set - membership filtering method to estimate the belief state of each UAV at time t; Inputting the belief states of all UAVs at time t into the UAV decision - making network to obtain the actions of each UAV at time t; the UAV decision - making network is constructed based on the Q - mix network; The step of using the set - membership filtering method to estimate the belief state of each UAV at time t according to the multi - UAV environment model at time t specifically includes: According to the multi - UAV environment model at time t, using the set - membership filtering method to construct a belief estimation network; the belief estimation network includes a prediction network and an observation network; Using the belief estimation network to determine the state estimation value of each UAV at time t; According to the state estimation value at time t and the shape matrix at time t, determining the belief state of each UAV at time t; The calculation formula for the state estimation value at time t is: ; Among them, is the state estimation value of UAV i at time t output by the prediction network; is the undetermined parameter of UAV i at time t; is the state observed by UAV i at time t; is the state estimation value of UAV i at time t output by the observation network; is the observation noise at time t; is the time-varying matrix of the observation noise of UAV i at time t; is the undetermined parameter, ; is the state estimation value of UAV i at time t after filtering; is the true state of UAV i at time t; is the shape matrix of UAV i at time t; T is the transpose; is the symbol representing the area; represents the set range value corresponding to UAV i at time t; z represents a variable; represents the area containing the true state.
2. The multi-UAV motion planning method in a complex environment according to claim 1, wherein, The method for determining the UAV decision - making network is: Constructing a Q - mix network; the Q - mix network includes multiple deep recurrent Q - learning networks and a mixing network; one deep recurrent Q - learning network corresponds to one UAV; each of the deep recurrent Q - learning networks is connected to the mixing network; Using the reinforcement learning method to train and test the Q - mix network to obtain the UAV decision - making network.
3. A multi-UAV motion planning method in a complex environment according to claim 2, characterized in that Using the reinforcement learning method to train and test the Q - mix network to obtain the UAV decision - making network specifically includes: Inputting the belief states of each UAV into the corresponding deep recurrent Q - learning network; each of the deep recurrent Q - learning networks selects the actions of each UAV based on the ε - greedy strategy and outputs the action values of each UAV; the action values of each UAV are used as the input of the mixing network, and the Q - mix network is jointly trained with the goal of minimizing the joint action loss to obtain a trained Q - mix network; Testing the trained Q - mix network and determining the tested Q - mix network as the UAV decision - making network; Among them, the action value is used to determine the selected action; the joint action loss is determined according to the joint action value function; the joint action value function is determined according to the action values and weight values of each UAV; the weight value is obtained by processing the global belief state using a hyper - network; the global belief state is determined according to the belief states of all UAVs.
4. A multi-UAV motion planning method in a complex environment according to claim 1, characterized in that, The environment model is constructed based on the physical characteristics of the environment; the physical characteristics of the environment include the actual physical characteristics of the UAVs, the actual physical characteristics of the obstacles, and the actual physical characteristics of the targets.
5. A multi-UAV motion planning system, characterized in that, including: An environment model construction module, configured to construct a multi - UAV environment model at time t; the multi - UAV environment model includes: the state space of each UAV, the action space of each UAV, the observation space of each UAV, a state transition model, an observation model, and a reward model; t≥0; The belief state estimation module is used to estimate the belief state of each unmanned aerial vehicle (UAV) at time t according to the multi-UAV environment model at time t by using the set membership filtering method; The step of estimating the belief state of each UAV at time t according to the multi-UAV environment model at time t by using the set membership filtering method specifically includes: According to the multi-UAV environment model at time t, a belief estimation network is constructed by using the set membership filtering method; the belief estimation network includes a prediction network and an observation network; The belief estimation network is used to determine the state estimation value of each UAV at time t; According to the state estimation value at time t and the shape matrix at time t, the belief state of each UAV at time t is determined; The calculation formula for the state estimation value at time t is: ; Among them, is the state estimation value of UAV i at time t output by the prediction network; are the undetermined parameters of UAV i at time t; is the state observed by UAV i at time t; is the state estimation value of UAV i at time t output by the observation network; is the observation noise at time t; is the time-varying matrix of the observation noise of UAV i at time t; are undetermined parameters, ; is the state estimation value of UAV i at time t after filtering; is the true state of UAV i at time t; is the shape matrix of UAV i at time t; T is the transpose; is the symbol representing the area; represents the set range value corresponding to UAV i at time t; z represents a variable; represents the area containing the true state; The action planning module is used to input the belief states of all UAVs at time t into the UAV decision network to obtain the actions of each UAV at time t; the UAV decision network is constructed based on the Q-mix network.