Unmanned aerial vehicle formation obstacle avoidance method and device based on reinforcement learning, and unmanned aerial vehicle
By converting the multi-objective reward function of the drone formation into a scalar reward function, and searching for the target weight vector in the simplified task environment, combining the simulation training environment with gradually increasing complexity and real machine data constraints, the problems of difficulty in multi-objective optimization, excessive search volume and large differences in simulation real machine in the UAV formation are solved, and efficient strategy training and migration are achieved.
Patent Information
- Application Number
- CN202510094497.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
In the field of drones, reinforcement learning methods have difficulties in the problems of multi-objective optimization, huge search space and large differences in simulation real machines, resulting in difficulty in converging algorithm training and degradation in strategy performance when deploying real machines.
By obtaining the multi-objective reward function of the drone formation, converting it into a scalar reward function, and searching for the target weight vector in a simplified task environment. A simulation training environment is built that gradually increases complexity, and real machine data is introduced as training constraints in simulation training, collective thrust and body angular velocity are used as action representations of reinforcement learning strategies, and finally the strategy is transferred to the real machine.
It effectively solves the problems of difficulty in multi-objective optimization, excessive search volume and large differences in simulation real machines, reduces training overhead, and reduces the migration error of the strategy from simulation to real machines.
Smart Images

Figure CN120029312A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of UAV technology, and in particular to a UAV formation obstacle avoidance method, device and UAV based on reinforcement learning. Background Art
[0002] In the field of drones, reinforcement learning, as an artificial intelligence method that enables the system to learn through trial and error and optimize behavior strategies through interaction with the environment, is gradually being applied. At present, the application of reinforcement learning in the field of drones has the following three main difficulties:
[0003] Multi-objective optimization is difficult. Reinforcement learning strategies need to balance the following objectives end-to-end: flying according to instructions, maintaining formation, avoiding dynamic and static obstacles, and avoiding collisions. These objectives often have complex interactions and may even conflict with each other in some cases, making it difficult to define the reward function for reinforcement learning.
[0004] Large search space. Compared with unmanned vehicles, the action space of drones is three-dimensional. The increased dimension, coupled with the extremely high complexity of the formation obstacle avoidance task itself, results in a huge search space for the algorithm, making it difficult for training to converge.
[0005] There is a big difference between simulation and real machine. Reinforcement learning requires the drone to interact with the environment, obtain trial and error data, and then adjust the behavior. Since the cost of trial and error on real machines is too high, drones are generally allowed to collect data in a simulation environment. However, there is a large gap between the simulation environment and the real machine environment, including the simulation modeling is not accurate enough, and the real environment has the center of mass offset, wind disturbance and other situations. Therefore, when the strategy is deployed to the real machine, it usually faces different degrees of performance loss. Summary of the invention
[0006] The present application provides a UAV formation obstacle avoidance method, device and UAV based on reinforcement learning to solve the problems of multi-objective optimization difficulty, excessive search volume and large difference between simulation and real machine in related technologies.
[0007] The first aspect of the present application provides a reinforcement learning-based obstacle avoidance method for a UAV formation, comprising the following steps: obtaining a multi-objective reward function for the UAV formation, converting the multi-objective reward function into a scalar reward function, and searching for a target weight vector of the scalar reward function in a simplified task environment; constructing a simulation training environment with gradually increasing complexity, and performing simulation training on a reinforcement learning strategy for the UAV formation in the simulation training environment, wherein the reinforcement learning strategy uses the scalar reward function corresponding to the target weight vector to solve the obstacle avoidance task; during the simulation training process, the real machine data of the UAV formation is used as a training constraint, and the collective thrust and body angular velocity are used as action representations of the reinforcement learning strategy, and after the simulation training is completed, the reinforcement learning strategy is migrated to the real machine of the UAV formation.
[0008] Optionally, the scalar reward function is a weighted combination of multi-objective reward vector functions:
[0009] R 总 = w T R
[0010] where w is a weight vector, w = (w 编队 , w 前进 , w 避障 , w 真机约束 ), and the weight vector will be obtained by algorithm search. Each w i represents the proportion of the corresponding sub-reward function R i in R 总 , and ∑w i = 1, w T is the transpose of w; R is a reward vector composed of four sub-reward functions, R = (R 编队 , R 前进 , R 避障 , R 真机约束 ), R 编队 = α 形状 × r r (e 形状 ) + α 最远距离 × r r (e 最远距离 ) + α 最近距离 × r i (e 最近距离 ), where e 形状 is the difference between the current UAV formation and the target formation measured using the Laplacian matrix, e 最远距离 is the distance between the two farthest UAVs in the UAV cluster, e 最近距离 is the distance between the two closest UAVs in the UAV cluster; R 前进 = α 朝向 × r l (e 朝向 ) + α 速度 × r l (e 速度 ) + α 位置 × r r (e 位置 ) + α 高度 × r l (e 高度 ), where e 朝向 is the difference between the orientation of each UAV and the target orientation, e 速度 is the difference between the speed of each UAV and the target speed, e 位置 is the difference between the position of each UAV and the target position, e 高度 is the difference between the altitude of each UAV and the target altitude; R 真机约束 = α 网络 × rl (e 网络 )+α 推力差 × l (e 推力差 )+α 推力和 × l (e 推力和 )+α 偏航角 × l (e 偏航角 ), where e 网络 is the difference between the neural network outputs in two consecutive time steps, e 推力差 is the difference in thrust of the drone motor in two consecutive time steps, e 推力和 is the sum of the thrusts of the four motors of the drone in the current time step, e 偏航角 is the difference between the yaw angle of the UAV and the yaw angle of the target; r l (e) r r (e) and r i (e) are the three reward functions involved in the multi-objective reward function, which are the linear function r l (e) = 1-|e|; reciprocal function r r (e) = 1 / (1+e); indicating function r i (e) = 1 e>阈值 , α is the scaling constant; R 避障 Only the distance d from the drone to the nearest obstacle 障碍物 About, when d 障碍物 Greater than the safety distance d 安全 When R 避障 =α 安全距离 ×(d 障碍物 -d 安全 );When d 障碍物 Less than the danger distance d 危险 When R 避障 =α 撞击系数 When d 危险 <d 障碍物 <d 安全 When R 避障 =α 危险距离 ×(d 障碍物 -d 危险 ).
[0011] Optionally, searching for a target weight vector of a scalar reward function in a simplified task environment includes: selecting a random weight vector; performing a task in the simplified task environment based on a scalar reward function corresponding to the random weight vector; using a multi-agent proximal policy optimization algorithm (MAPPO) to iteratively solve a target policy corresponding to the scalar reward function, and determining the target weight vector based on performance selection and expected behavior of the target policy.
[0012] Optionally, the simulation training environment includes, in sequence: an environment without obstacles, an environment with static obstacles, and an environment combining static obstacles with dynamic obstacles.
[0013] Optionally, the drones of the drone formation are provided with an observation encoder based on an attention mechanism, and the observation encoder uses the observation information as input of a reinforcement learning strategy.
[0014] Optionally, the observation encoder includes a self-perception part, a perception part for other drones, a perception part for static obstacles, and a perception part for dynamic obstacles, wherein each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension, and a multi-head self-attention mechanism is applied to the latent space vector to obtain other perception features, and the self-perception features and other perception features are spliced to obtain the observation information.
[0015] Optionally, during the simulation training process, it also includes: randomizing the initial position of each drone and clipping the output of the reinforcement learning strategy.
[0016] The second aspect of the present application provides a reinforcement learning-based obstacle avoidance device for a UAV formation, including: an acquisition module, used to acquire a multi-objective reward function for a UAV formation, convert the multi-objective reward function into a scalar reward function, and search for a target weight vector of the scalar reward function in a simplified task environment; a construction module, used to construct a simulation training environment with gradually increasing complexity, simulate training the reinforcement learning strategy of the UAV formation in the simulation training environment, and the reinforcement learning strategy uses the scalar reward function corresponding to the target weight vector to solve the obstacle avoidance task; a migration module, used to use the real machine data of the UAV formation as training constraints during the simulation training process, and use the collective thrust and body angular velocity as action representations of the reinforcement learning strategy, and migrate the reinforcement learning strategy to the real machine of the UAV formation after the simulation training is completed.
[0017] A third aspect of the present application provides a drone, which avoids obstacles based on the drone formation obstacle avoidance method based on reinforcement learning of the first aspect.
[0018] The fourth aspect of the present application provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the reinforcement learning-based obstacle avoidance method for drone formations of the first aspect is implemented.
[0019] Therefore, this application includes the following beneficial effects:
[0020] The embodiment of the present application obtains the multi-objective reward function of the drone formation, converts it into a scalar reward function, searches for the target weight vector of the scalar reward function in a simplified task environment, constructs a simulation training environment with gradually increasing complexity, performs simulation training on the reinforcement learning strategy of the drone formation in the simulation training environment, uses the real machine data of the drone formation as training constraints during the simulation training, and uses the collective thrust and body angular velocity as action representations of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is migrated to the real machine of the drone formation, and a random search method is used to uniformly sample in the combination space of multiple optimization objectives. At the same time, the training is divided into two stages, thereby reducing the training overhead, and introducing real machine constraints during the training process, thereby reducing the migration error of the strategy from simulation to real machine. In this way, the problems of multi-objective optimization difficulty, excessive search volume, and large differences between simulation and real machine in related technologies are solved.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 A flowchart of a UAV formation obstacle avoidance method based on reinforcement learning according to an embodiment of the present application;
[0024] Figure 2 A two-stage reinforcement learning training process provided according to an embodiment of the present application;
[0025] Figure 3 A structural diagram of an observation encoder provided according to an embodiment of the present application;
[0026] Figure 4 This is an example diagram of a UAV formation obstacle avoidance device based on reinforcement learning provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0028] The following describes a method, apparatus, and unmanned aerial vehicle (UAV) for UAV formation obstacle avoidance based on reinforcement learning according to embodiments of the present application. In view of the problems in the related art such as difficult multi-objective optimization, excessive search volume, and large differences between simulation and real aircraft mentioned in the above background art, the present application provides a method for UAV formation obstacle avoidance based on reinforcement learning. In this method, in embodiments of the present application, a multi-objective reward function of the UAV formation is obtained, converted into a scalar reward function, and the target weight vector of the scalar reward function is searched in a simplified task environment. A simulation training environment with gradually increasing complexity is constructed, and the reinforcement learning strategy of the UAV formation is simulated and trained in the simulation training environment. During the simulation training process, the real aircraft data of the UAV formation is used as a training constraint, and the collective thrust and body angular velocity are used as the action representation of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is migrated to the real aircraft of the UAV formation. A random search method is adopted to uniformly sample in the combined space of multiple optimization objectives. At the same time, the training is divided into two stages to reduce the training cost, and real aircraft constraints are introduced during the training process to reduce the migration error of the strategy from simulation to real aircraft. Thus, the problems in the related art such as difficult multi-objective optimization, excessive search volume, and large differences between simulation and real aircraft are solved.
[0029] Specifically, Figure 1 FIG. is a schematic flowchart of a method for UAV formation obstacle avoidance based on reinforcement learning provided by an embodiment of the present application.
[0030] As Figure 1 shown, the method for UAV formation obstacle avoidance based on reinforcement learning includes the following steps:
[0031] In step S101, a multi-objective reward function of the UAV formation is obtained, the multi-objective reward function is converted into a scalar reward function, and the target weight vector of the scalar reward function is searched in a simplified task environment.
[0032] Among them, the method of converting the multi-objective reward function into a scalar reward function uses the MORL (Multi-Objective Reinforcement Learning) method; the simplified task environment is an environment that only includes dynamic obstacles and does not add the constraint of real aircraft migration.
[0033] It can be understood that in embodiments of the present application, the MORL method is used to convert the multi-objective reward function into a scalar reward function, and the target weight vector of the scalar reward function is searched in the simplified environmental task. The search method will be described in detail below and will not be elaborated here.
[0034] In embodiments of the present application, the scalar reward function is a weighted combination of multi-objective reward vector functions:
[0035] R 总 = wT R
[0036] Among them, w is the weight vector, w=(w 编队 , w 前进 , w 避障 , w 真机约束 ), the weight vector will be searched by the algorithm, where each item w i Represents the corresponding sub-reward function R i R 总 The ratio of ∑w i =1,w T is the transpose of w; R is the reward vector, which consists of four sub-reward functions, R = (R 编队 , R 前进 , R 避障 , R 真机约束 ), R 编队 =α 形状 × r (e 形状 )+α 最远距离 × r (e 最远距离 )+α 最近距离 × i (e 最近距离 ), where e 形状 To use the Laplace matrix to measure the difference between the current UAV formation and the target formation, e 最远距离 is the distance between the two farthest drones in the drone cluster, e 最近距离 is the distance between the two nearest drones in the drone cluster; R 前进 =α 朝向 × l (e 朝向 )+α 速度 × l (e 速度 )+α 位置 × r (e 位置 )+a 高度 × l (e 高度 ), where e 朝向 is the difference between the orientation of each UAV and the orientation of the target, e 速度 is the difference between the speed of each UAV and the target speed, e 位置 is the difference between the position of each UAV and the target position, e 高度 is the difference between the height of each UAV and the target height; R 真机约束 =α 网络 × l (e 网络 )+α 推力差 × l (e 推力差 )+α推力和 ×r l (e 推力和 ) + α 偏航角 ×r l (e 偏航角 ),where e 网络 is the difference between the neural network outputs in two consecutive time steps, and e 推力差 is the difference in the thrusts of the drone motors in two consecutive time steps, and e 推力和 is the sum of the thrusts of the four motors of the drone in the current time step, and e 偏航角 is the difference between the yaw angle of the drone and the target yaw angle; r l (e), r r (e) and r i (e) are three reward functions involved in the multi - objective reward function, which are the linear function r l (e) = 1 - |e|; the reciprocal function r r (e) = 1 / (1 + e); the indicator function r i (e) = 1 e>阈值 , and α is a scaling constant; R 避障 is only related to the distance d 障碍物 from the drone to the nearest obstacle. When d 障碍物 is greater than the safety distance d 安全 , R 避障 = α 安全距离 ×(d 障碍物 - d 安全 ); when d 障碍物 is less than the dangerous distance d 危险 , R 避障 = α 撞击系数 ; when d 危险 <d 障碍物 <d 安全 , R 避障 = α 危险距离 ×(d 障碍物 - d 危险 ).
[0037] In the embodiment of the present application, searching for the target weight vector of the scalar reward function in a simplified task environment includes: selecting a random weight vector; performing tasks in the simplified task environment based on the scalar reward function corresponding to the random weight vector; using the multi - agent proximal policy optimization algorithm to iteratively solve the target policy corresponding to the scalar reward function, and determining the target weight vector according to the performance of the target policy and the expected behavior.
[0038] Among them, the multi - agent proximal policy optimization algorithm is a reinforcement learning algorithm that allows the agent to explore new behaviors without significantly deviating from the current policy, thereby stably updating the policy and avoiding the learning instability that may be caused by a large - scale change in the policy.
[0039] It can be understood that the embodiment of the present application searches for the target weight vector of the scalar reward function in a simplified task environment, and needs to repeat the following actions: select a random weight vector, obtain the scalar reward function corresponding to the vector, perform the task in the simplified task environment, use the multi-agent proximal policy optimization algorithm to iteratively solve the target strategy corresponding to the scalar reward function, and after repeated operations, determine the target weight vector based on the performance selection and expected behavior of the target strategy.
[0040] In step S102, a simulation training environment with gradually increasing complexity is constructed, and a reinforcement learning strategy of the UAV formation is simulated and trained in the simulation training environment. The reinforcement learning strategy uses a scalar reward function corresponding to a target weight vector to solve an obstacle avoidance task.
[0041] The simulation training environment includes static and dynamic obstacles, which will be described in detail below and will not be repeated here.
[0042] It can be understood that after obtaining the target weight vector, the embodiment of the present application reconstructs a simulation training environment with gradually increasing complexity, places the drone formation into the simulation training environment with gradually increasing complexity to perform tasks, and simulates training for the reinforcement learning strategy. The reinforcement learning strategy uses the scalar reward function corresponding to the obtained target weight vector to solve the obstacle avoidance task.
[0043] In the embodiment of the present application, the simulation training environment includes: an environment without obstacles, an environment with static obstacles, and an environment with a combination of static obstacles and dynamic obstacles.
[0044] It can be understood that the simulation training environment of the embodiment of the present application includes: an environment without obstacles, an environment with static obstacles, and an environment combining static obstacles with dynamic obstacles. The complexity can be gradually increased by first providing an environment without obstacles, then an environment containing only static obstacles, and finally an environment containing both static and dynamic obstacles.
[0045] In an embodiment of the present application, the drones in the drone formation are provided with an observation encoder based on an attention mechanism, and the observation encoder uses the observation information as the input of the reinforcement learning strategy.
[0046] Among them, the observation encoder is based on the attention mechanism, which is responsible for processing all observation information received by the drone and converting this information into a format suitable for reinforcement learning strategies.
[0047] It can be understood that the embodiment of the present application sets up an observation encoder based on the attention mechanism on the drone. The observation encoder is based on the attention mechanism and is responsible for processing all observation information received by the drone and converting this information into a format suitable for use by the reinforcement learning strategy. The observation information can be used as the input of the reinforcement learning strategy.
[0048] In an embodiment of the present application, the observation encoder includes a self-perception part, a perception part for other drones, a perception part for static obstacles, and a perception part for dynamic obstacles, wherein each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension, and a multi-head self-attention mechanism is applied to the latent space vector to obtain other perception features, and the self-perception features and other perception features are spliced to obtain observation information.
[0049] Among them, the latent space vector refers to an abstract representation form converted from the original observation information, which usually has a fixed dimension and can be used in the same way in different contexts.
[0050] It can be understood that the observation encoder of the embodiment of the present application includes a self-perception part, a perception part for other drones, a perception part for static obstacles, and a perception part for dynamic obstacles. These three parts can realize self-perception, perception of other drones, perception of static obstacles, and perception of dynamic obstacles, and the perception information obtained by each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension, and a multi-head self-attention mechanism is applied to obtain other perception features, and finally the self-perception features and other perception features are spliced to obtain the observation information.
[0051] In an embodiment of the present application, during the simulation training process, it also includes: randomizing the initial position of each drone and clipping the output of the reinforcement learning strategy.
[0052] Among them, the clipping is CTBR (Collective Thrust and Body Rates), which can limit the control commands generated by the reinforcement learning algorithm to ensure that these commands are within a safe and practical range.
[0053] It is understandable that in the simulation training process, the embodiment of the present application needs to randomize the initial position of each drone and use CTBR to clip the output of the reinforcement learning strategy to avoid extreme actions.
[0054] In step S103, during the simulation training process, the real aircraft data of the UAV formation is used as training constraints, and the collective thrust and body angular velocity are used as action representations of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is transferred to the real aircraft of the UAV formation.
[0055] It can be understood that the embodiment of the present application introduces data from real drones in the simulation training, and uses the real machine data of the drone formation as training constraints to ensure that the simulation environment is as close to reality as possible, and migrates the reinforcement learning strategy to the real machines of the drone formation after the simulation training is completed.
[0056] According to the obstacle avoidance method for drone formation based on reinforcement learning proposed in the embodiment of the present application, a multi-objective reward function of the drone formation is obtained, converted into a scalar reward function, and the target weight vector of the scalar reward function is searched in a simplified task environment, a simulation training environment with gradually increasing complexity is constructed, and the reinforcement learning strategy of the drone formation is simulated and trained in the simulation training environment. During the simulation training, the real machine data of the drone formation is used as training constraints, and the collective thrust and body angular velocity are used as action representations of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is migrated to the real machine of the drone formation, and a random search method is used to uniformly sample in the combination space of multiple optimization objectives. At the same time, the training is divided into two stages, thereby reducing the training overhead, and real machine constraints are introduced during the training process, thereby reducing the migration error of the strategy from simulation to real machine.
[0057] The following is a further description of the obstacle avoidance method for UAV formation based on reinforcement learning through a specific embodiment.
[0058] like Figure 2 As shown in the figure, in the first stage of reward scalarization, a random search is performed in a relatively simple scenario to find the multi-objective combination weight that best meets the expectations. After the weight vector is determined, the second stage is entered, the reward function is generalized, and the reward function corresponding to the weight is used to solve the more complex dynamic obstacle avoidance task of multi-drone formations, and to realize real-machine deployment. Here, course learning is used to accelerate training. The two stages are described separately as follows:
[0059] The first stage is reward scalarization. The multi-objective reinforcement learning method is used to convert the multi-objective rewards into scalar values. Assume that the scalar reward function is a linear combination of the rewards of each sub-objective, that is, R 总 =w T R, where the weight vector is w = (w 编队 , w 前进 , w 避障 , w 真机约束 ), the reward vector is R = (R 编队 , R 前进 , R 避障 , R 真机约束). The weight vector w represents the relationship between each sub-goal. Through the weight vector, the multi-objective problem can be transformed into a single-objective problem. Existing MORL methods are not suitable for solving multi-objective optimization problems of drones because they often assume that the reward function of each goal is highly monotonic, while the multi-objective tasks of drones do not have such characteristics. Therefore, when the sub-goal rewards are non-monotonic as the weight vector changes, the existing MORL methods will not be able to fully explore the weight space, making it difficult to obtain a suitable solution. Therefore, in order to scalarize the reward vector, the following actions are repeated: select a random weight vector, obtain the corresponding scalar reward function, use the multi-agent proximal policy optimization algorithm to solve the optimal policy corresponding to the reward function, and evaluate the performance of the policy. Since this process requires multiple iterative solutions, the computational cost is high, and it is very time-consuming, a suitable target weight vector is searched in a slightly simplified environment. This task environment only includes dynamic obstacles and no constraints on real machine migration are added. Finally, the weight vector that best matches the expected behavior is selected for the second stage.
[0060] The second stage is the generalization of the reward function. In the first stage, the combined weight w of multiple objectives was determined in a simplified environment. This weight and its corresponding reward function are applied to the obstacle avoidance task of a UAV formation, where multiple UAVs navigate through static and dynamic obstacles, and real machine constraints are added. Since directly training RL strategies for complex tasks from scratch often requires high computing power and often converges to suboptimal solutions, curriculum learning is used to gradually increase the difficulty of the tasks, thereby accelerating the training process and improving the final performance of the strategy. Specifically, a three-stage curriculum was designed, in which the UAV first learns to fly forward in an environment without obstacles, then flies in an environment containing only static obstacles, and finally flies in an environment containing both static and dynamic obstacles.
[0061] The network architecture of this embodiment is described below:
[0062] In order to enable the UAV strategy to cope with scenarios with different numbers of obstacles, an observation encoder based on the attention mechanism is designed to eliminate the limitation of the policy network on the input dimension and enhance the information representation related to the UAV task. The observation consists of four parts, such as Figure 3 As shown in the figure, it includes self-perception, perception of other drones, perception of static obstacles, and perception of dynamic obstacles. Each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension. A multi-head self-attention mechanism is applied to these latent space vectors to obtain the corresponding features. In order to better capture the relationship between drones and other environmental entities, these features are further processed by a multi-head cross-attention module, which uses self-perception features as queries and other features as keys and values. Finally, the self-perception features are spliced together with the output of the multi-head cross-attention module to obtain a complete state representation of the scene.
[0063] The following is the simulation real machine migration method of this embodiment:
[0064] Collective Thrust and Body Rates (CTBR) commands are used as the action of the strategy to balance the transfer effect from simulation to real aircraft and dexterous control. Specifically, the action of the i-th drone in the drone cluster is expressed as a = (c, ω roll ,ω pitch ,ω yaw ), where c represents the collective thrust and ω represents the body angular velocity corresponding to the three rotation axes. In order to better achieve the migration from simulation to real machine, a large amount of real machine data was collected, and the real machine reward function target was set according to its characteristics, encouraging the strategy to minimize the sum of the throttle outputs of the four rotors, the rotation angle on the yaw axis, and the throttle output difference and network output difference of each rotor between adjacent time steps. In addition, domain randomization and CTBR clipping were introduced during the training process. Specifically, the initial position of each drone is randomized within 0.05 meters of the original formation, and the initial rotation angle is randomized within the range of [-0.1π, 0.1π]. In addition, to avoid extreme actions, the policy output is clipped, limiting ω to the range of [-45, 45] degrees / second and c to the range of [0.4, 0.9].
[0065] Next, the obstacle avoidance device for UAV formation based on reinforcement learning proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0066] Figure 4 It is a block diagram of a UAV formation obstacle avoidance device based on reinforcement learning according to an embodiment of the present application.
[0067] like Figure 4 As shown, the UAV formation obstacle avoidance device 10 based on reinforcement learning includes: an acquisition module 201, a construction module 202 and a migration module 203.
[0068] Among them, the acquisition module 201 is used to obtain the multi-objective reward function of the UAV formation, convert the multi-objective reward function into a scalar reward function, and search for the target weight vector of the scalar reward function in a simplified task environment; the construction module 202 is used to construct a simulation training environment with gradually increasing complexity, and simulate the reinforcement learning strategy of the UAV formation in the simulation training environment. The reinforcement learning strategy uses the scalar reward function corresponding to the target weight vector to solve the obstacle avoidance task; the migration module 203 is used to use the real machine data of the UAV formation as training constraints during the simulation training process, and use the collective thrust and body angular velocity as the action representation of the reinforcement learning strategy, and migrate the reinforcement learning strategy to the real machine of the UAV formation after the simulation training is completed.
[0069] In this embodiment of the present application, the scalar reward function is:
[0070] R=u(R)=w T R
[0071] Among them, w is the weight vector, w=(w 编队 , w 前进 , w 避障 , w 真机约束 ); R is the reward vector, R = (R 编队 , R 前进 , R 避障 , R 真机约束 ).
[0072] In an embodiment of the present application, the acquisition module 201 is further used to: search for a target weight vector of a scalar reward function in a simplified task environment, including: selecting a random weight vector; performing a task in the simplified task environment based on a scalar reward function corresponding to the random weight vector; using a multi-agent proximal strategy optimization algorithm to iteratively solve a target strategy corresponding to the scalar reward function, and determining the target weight vector based on the performance selection and expected behavior of the target strategy.
[0073] In the embodiment of the present application, the simulation training environment includes: an environment without obstacles, an environment with static obstacles, and an environment with a combination of static obstacles and dynamic obstacles.
[0074] In an embodiment of the present application, the drones in the drone formation are provided with an observation encoder based on an attention mechanism, and the observation encoder uses the observation information as the input of the reinforcement learning strategy.
[0075] In an embodiment of the present application, the observation encoder includes a self-perception part, a perception part for other drones, a perception part for static obstacles, and a perception part for dynamic obstacles, wherein each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension, and a multi-head self-attention mechanism is applied to the latent space vector to obtain other perception features, and the self-perception features and other perception features are spliced to obtain observation information.
[0076] In an embodiment of the present application, during the simulation training process, it also includes: randomizing the initial position of each drone and clipping the output of the reinforcement learning strategy.
[0077] It should be noted that the aforementioned explanation of the embodiment of the UAV formation obstacle avoidance device based on reinforcement learning is also applicable to the UAV formation obstacle avoidance device based on reinforcement learning of this embodiment, and will not be repeated here.
[0078] According to the reinforcement learning-based obstacle avoidance device for a drone formation proposed in an embodiment of the present application, a multi-objective reward function of the drone formation is obtained and converted into a scalar reward function, and the target weight vector of the scalar reward function is searched in a simplified task environment to construct a simulation training environment with gradually increasing complexity. The reinforcement learning strategy of the drone formation is simulated and trained in the simulation training environment. During the simulation training, the real machine data of the drone formation is used as training constraints, and the collective thrust and body angular velocity are used as action representations of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is migrated to the real machine of the drone formation. A random search method is used to uniformly sample in the combination space of multiple optimization objectives. At the same time, the training is divided into two stages, thereby reducing the training overhead. Real machine constraints are introduced during the training process, thereby reducing the migration error of the strategy from simulation to real machine.
[0079] An embodiment of the present application also provides a drone, which avoids obstacles based on the above-mentioned drone formation obstacle avoidance method based on reinforcement learning.
[0080] An embodiment of the present application also provides a computer-readable storage medium on which a computer program or instruction is stored. When the computer program or instruction is executed, the above-mentioned reinforcement learning-based drone formation obstacle avoidance method is implemented.
[0081] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0082] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0083] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0084] It should be understood that the various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, the steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.
[0085] A person of ordinary skill in the art may understand that all or part of the steps carried by the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the above-mentioned program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0086] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A UAV formation obstacle avoidance method based on reinforcement learning, characterized in that: The following steps are involved: Obtaining a multi-objective reward function for a UAV formation, converting the multi-objective reward function into a scalar reward function, and searching for a target weight vector of the scalar reward function in a simplified task environment; Constructing a simulation training environment with gradually increasing complexity, and performing simulation training on the reinforcement learning strategy of the UAV formation in the simulation training environment, wherein the reinforcement learning strategy uses a scalar reward function corresponding to the target weight vector to solve an obstacle avoidance task; During the simulation training process, the real aircraft data of the UAV formation is used as training constraints, and the collective thrust and the body angular velocity are used as action representations of the reinforcement learning strategy. After the simulation training is completed, the reinforcement learning strategy is migrated to the real aircraft of the UAV formation.
2. The obstacle avoidance method for UAV formation based on reinforcement learning according to claim 1, characterized in that: The scalar reward function is a weighted combination of multiple objective reward vector functions: R 总 =w T R Among them, w is the weight vector, w=(w 编队 , w 前进 , w 避障 , w 真机约束 ), the weight vector will be searched by the algorithm, where each item w i Represents the corresponding sub-reward function R i R 总 The ratio of ∑w i =1,w T is the transpose of w; R is the reward vector, which consists of four sub-reward functions, R = (R 编队 , R 前进 , R 避障 , R 真机约束 ), R 编队 =α 形状 × r (e 形状 )+α 最远距离 × r (e 最远距离 )+α 最近距离 × i (e 最近距离 ), where e 形状 To use the Laplace matrix to measure the difference between the current UAV formation and the target formation, e 最远距距离 is the distance between the two farthest drones in the drone cluster, e 最近距离 is the distance between the two nearest drones in the drone cluster; R 前进 =α 朝向 × l (e 朝向 )+α 速度 × l (e 速度 )+α 位置 × r (e 位置 )+α 高度 × l (e 高度 ), where e 朝向 is the difference between the orientation of each UAV and the orientation of the target, e 速度 is the difference between the speed of each UAV and the target speed, e 位置 is the difference between the position of each UAV and the target position, e 高度 is the difference between the height of each UAV and the target height; R 真机约束 =α 网络 × l (e 网络 )+α 推力差 × l (e 推力差 )+α 推力和 × l (e 推力和 )+α 偏航角 × l (e 偏航角 ), where e 网络 is the difference between the neural network outputs in two consecutive time steps, e 推力差 is the difference in thrust of the drone motor in two consecutive time steps, e 推力和 is the sum of the thrusts of the four motors of the drone in the current time step, e 偏航角 is the difference between the yaw angle of the UAV and the yaw angle of the target; r l (e) r r (e) and r i (e) are the three reward functions involved in the multi-objective reward function, which are the linear function r l (e) = 1-|e|; reciprocal function r r (e) = 1 / (1+e); indicating function r i (e) = 1 e>阈值 , α is the scaling constant; R 避障 Only the distance d from the drone to the nearest obstacle 障碍物 About, when d 障碍物 Greater than the safety distance d 安全 When R 避障 =α 安全距离 ×(d 障碍物 -d 安全 );When d 障碍物 Less than the danger distance d 危险 When R 避障 =α 撞击系数 When d 危险 <d 障碍物 <d 安全 When R 避障 =α 危险距离 ×(d 障碍物 -d 危险 ).
3. The UAV formation obstacle avoidance method based on reinforcement learning according to claim 1, characterized in that: The step of searching for the target weight vector of the scalar reward function in the simplified task environment comprises: Choose a random weight vector; Execute the task in the simplified task environment based on a scalar reward function corresponding to the random weight vector; A multi-agent proximal policy optimization algorithm is used to iteratively solve the target strategy corresponding to the scalar reward function, and a target weight vector is determined according to the performance selection and expected behavior of the target strategy.
4. The UAV formation obstacle avoidance method based on reinforcement learning according to claim 1, characterized in that: The simulation training environment includes: an environment without obstacles, an environment with static obstacles, and an environment with both static obstacles and dynamic obstacles.
5. The UAV formation obstacle avoidance method based on reinforcement learning according to claim 1, characterized in that: The drones of the drone formation are provided with an observation encoder based on an attention mechanism, and the observation encoder uses observation information as input of the reinforcement learning strategy.
6. The UAV formation obstacle avoidance method based on reinforcement learning according to claim 5 is characterized in that: The observation encoder includes a self-perception part, a perception part for other drones, a perception part for static obstacles, and a perception part for dynamic obstacles, wherein each part is encoded by an independent multi-layer perceptron to generate a latent space vector of the same dimension, and a multi-head self-attention mechanism is applied to the latent space vector to obtain other perception features, and the self-perception features and the other perception features are spliced to obtain observation information.
7. The UAV formation obstacle avoidance method based on reinforcement learning according to claim 1, characterized in that: The simulation training process also includes: Randomize the initial position of each drone and clip the output of the reinforcement learning policy.
8. A UAV formation obstacle avoidance device based on reinforcement learning, characterized in that: include: An acquisition module is used to acquire a multi-objective reward function of a UAV formation, convert the multi-objective reward function into a scalar reward function, and search for a target weight vector of the scalar reward function in a simplified task environment; A construction module is used to construct a simulation training environment of gradually increasing complexity, and simulate the reinforcement learning strategy of the UAV formation in the simulation training environment, wherein the reinforcement learning strategy uses a scalar reward function corresponding to the target weight vector to solve the obstacle avoidance task; The migration module is used to use the real machine data of the UAV formation as training constraints during the simulation training process, and use the collective thrust and the body angular velocity as the action representation of the reinforcement learning strategy, and migrate the reinforcement learning strategy to the real machines of the UAV formation after the simulation training is completed.
9. A drone, characterized in that: The drone avoids obstacles based on the drone formation obstacle avoidance method based on reinforcement learning as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed, the reinforcement learning-based obstacle avoidance method for UAV formation is implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Unmanned vehicle dynamic obstacle avoidance method and system based on near-end strategy optimization
CN120491653A
Multi-agent virtual-real migration method and device
CN122472083A