Multi-unmanned aerial vehicle target hunting method based on internal and external reward mechanism reinforcement learning
By strengthening the learning method through internal and external reward mechanisms, the external reward mechanism enhances UAV collaboration, while the internal reward mechanism guides autonomous exploration. This solves the problems of insufficient exploration and poor convergence performance in multi-UAV collaborative target capture, and achieves efficient and safe target capture.
Patent Information
- Application Number
- CN202511450464.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies struggle to efficiently achieve multi-UAV collaborative target capture in dynamic environments, and traditional methods suffer from insufficient exploration and poor convergence performance.
A reinforcement learning approach based on internal and external reward mechanisms is adopted. An external reward mechanism is designed to enhance UAV collaboration, while an internal reward mechanism guides the UAV to autonomously explore the environment. A loss function is constructed and a policy model is trained. The loss function of the policy model is constructed through the external and internal reward mechanisms to guide the UAV to efficiently capture targets in complex environments.
It enables UAVs to efficiently and safely complete target capture missions in complex environments, avoiding the problems of insufficient exploration and poor convergence performance in traditional methods. It can quickly adapt to dynamic changes and plan efficient and safe capture paths.
Smart Images

Figure CN121008591A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned aerial vehicle control, and in particular to a multi-unmanned aerial vehicle target encirclement method based on an internal and external reward mechanism reinforcement learning. BACKGROUND
[0002] With the development of science and technology, unmanned aerial vehicles have been widely used in civilian and military fields due to their high flexibility and strong maneuverability. A multi-unmanned aerial vehicle system is a new and interesting field that not only inherits the advantages of single unmanned aerial vehicles but also performs advanced tasks through information exchange and behavior collaboration, significantly improving the efficiency of task execution and environmental adaptability to complex environments. Cooperative target encirclement as a typical application scenario of multi-unmanned aerial vehicle systems not only effectively reduces potential threats but also significantly reduces personnel risks. This approach helps to achieve continuous monitoring of enemy targets, prevent enemy targets from escaping, and protect the safety of friendly targets.
[0003] Multi-unmanned aerial vehicle cooperative target encirclement is a relatively complex problem, especially in dynamic environments, where there are a large number of unpredictable and unpredictable uncertainty factors. These uncertainties may arise from sudden target movements, environmental changes, and other unforeseen circumstances. In the face of such complex situations, unmanned aerial vehicles need to flexibly adjust their flight paths and attitudes to ensure that they can work closely together to form an effective encirclement posture.
[0004] However, the existing target encirclement method is difficult to achieve efficient encirclement of the target. SUMMARY
[0005] Therefore, it is necessary to provide a multi-unmanned aerial vehicle target encirclement method based on an internal and external reward mechanism reinforcement learning, which can achieve efficient encirclement of the target.
[0006] The present application adopts the following technical solutions: The present application provides a multi-unmanned aerial vehicle target encirclement method based on an internal and external reward mechanism reinforcement learning, comprising: designing an internal and external reward mechanism; the internal and external reward mechanism includes an external reward mechanism and an internal reward mechanism, the external reward mechanism is constructed according to feedback signals given by the environment in response to the individual behavior of the unmanned aerial vehicle, and is used to enhance the cooperation between unmanned aerial vehicles; the internal reward mechanism is constructed according to feedback signals obtained from the behavior of the unmanned aerial vehicle in response to changes in the state of the unmanned aerial vehicle, and is used to guide the unmanned aerial vehicle to explore the environment and acquire new knowledge autonomously; constructing a loss function according to the external reward mechanism and the internal reward mechanism, and training a policy model according to the loss function; The environmental state information of the target UAV at the previous moment is input into the strategy model to obtain the control input vector of the target UAV at the current moment, and the environmental state information of the target UAV at the current moment is updated with the control input vector at the current moment; the target UAV is any UAV in the target capture mission; the environmental state information includes its own speed and position, as well as the speed and position of the target to be captured; If the Euclidean distance between the positions of all drones and the position of the target to be captured is a preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured, then the target capture mission is considered complete.
[0007] Optionally, an external reward mechanism may be designed, including: Based on the constraints of the multi-UAV cooperative target capture mission, an external reward mechanism is designed; the constraints of the multi-UAV cooperative target capture mission include collision avoidance constraints, motion continuity constraints, and energy consumption constraints. The collision avoidance constraints are: ; in, This indicates the time consumed in the target capture mission. This indicates the number of drones used in the target capture mission. Indicates the first A drone in Location at any given moment Indicates the first A drone in Location at any given moment Indicates in time and The Euclidean distance between them Indicates a safe distance. Indicates the number of obstacles. Indicates the first The coordinates of the center of each obstacle Indicates in time and The Euclidean distance between them Indicates the first The radius of the obstacle Indicates the target to be apprehended is in Location at any given moment Indicates in time and The Euclidean distance between them; The motion continuity constraint is: ; in, Indicates the first The starting position of the drone. Indicates the first The initial speed of the drone Indicates a preset fixed distance. Indicates the speed of the target to be captured. Indicates the first A drone in The speed of time, This indicates the preset time interval. Indicates the first A drone in The speed of time, Indicates that drones are in The control input vector at time t, Indicates the first A drone in The position at that moment; Energy consumption constraints are: ; in, Indicates the first A drone in The propulsion power at any moment Indicates the first The maximum available energy consumption of a drone.
[0008] Alternatively, the external reward mechanism may be: ; in, Indicates the first A drone in External rewards at all times This represents the reward generated by the collision avoidance constraint. This represents the reward resulting from the constraint of motion continuity. This represents the reward generated by energy consumption constraints. Indicates the first The target reward obtained by the drone after successfully capturing it. , , and These represent the adjustable weighting parameters for each reward; ; ; ; ; in, and The first The rewards earned by a drone for avoiding collisions with other drones and with other static obstacles. and These are variable weight parameters that take non-negative values; , and They represent the first The drone and the first A drone in The relative position, relative speed, and relative yaw angle at any given moment; and They represent the first The drone and the first One obstacle The relative position and relative yaw angle at any given time. and This represents a predefined weighting factor. It is a fixed repulsion coefficient for the energy consumption of drones. It is the first The drone was shut down Energy consumed constantly It is a positive real number.
[0009] Optionally, an intrinsic reward mechanism may be designed, including: An intrinsic reward mechanism is constructed based on the difference between the state error reward and the extrinsic reward; the state error reward represents the difference between the predicted state of the UAV and the target state; the intrinsic reward mechanism is as follows: ; in, Indicates the first A drone in Intrinsic rewards of moments Indicates the reward for state error. It is a single-layer fully connected network used to compare state error rewards. External rewards The differences between them.
[0010] Optionally, the loss function for: ; in, , express The L2 norm.
[0011] Optionally, the policy model includes a policy network and an evaluation network; the training process of the policy model includes: input the historical environment state information of the sample unmanned plane at the first time into the initial strategy network, predict the control input vector of the sample unmanned plane at the second time, and update the predicted historical environment state information of the sample unmanned plane at the second time with the control input vector at the second time; determine the loss function value according to the predicted historical environment state information at the second time and the target historical environment state information at the second time; update the initial strategy network and the initial evaluation network according to the loss function value, continue to iteratively train the updated initial strategy network and the initial evaluation network, and determine the initial strategy network and the initial evaluation network that meet the iteration condition as the strategy network and the evaluation network of the policy network model until a preset iteration condition is met.
[0012] Optionally, the environment state information of the target unmanned plane at the previous time is input into the policy model to obtain the control input vector of the target unmanned plane at the current time, comprising: input the environment state information of the target unmanned plane at the previous time into the strategy network to obtain the control input vector of the target unmanned plane at the current time.
[0013] The application provides a multi-unmanned plane target encircling device based on an internal and external reward mechanism reinforcement learning, comprising: The acquisition module is used for designing an internal and external reward mechanism; the internal and external reward mechanism comprises an external reward mechanism and an internal reward mechanism; the external reward mechanism is constructed according to feedback signals given by an environment in dependence on individual behaviors of unmanned planes, and is used for enhancing cooperation between the unmanned planes; the internal reward mechanism is constructed according to feedback signals obtained from behaviors of the unmanned planes in dependence on changes in states of the unmanned planes, and is used for guiding the unmanned planes to autonomously explore the environment and acquire new knowledge; The training module is used for constructing a loss function according to the external reward mechanism and the internal reward mechanism, and training a policy model according to the loss function; The running module is used for inputting historical environment state information of a target unmanned plane at a previous time into the policy model to obtain a control input vector of the target unmanned plane at a current time, and updating environment state information of the target unmanned plane at the current time with the control input vector at the current time; the target unmanned plane is any unmanned plane in a target encircling task; the environment state information comprises a speed and a position of the target unmanned plane and a speed and a position of a target to be encircled; The encircling module is used for determining that the target encircling task is completed if the Euclidean distances between the positions of all the unmanned planes and the position of the target to be encircled are all preset fixed distances, and the speeds of all the unmanned planes are consistent with the speed of the target to be encircled.
[0014] The application provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to realize the multi-unmanned aerial vehicle target encircling method based on the internal and external reward mechanism reinforcement learning.
[0015] The application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, and the processor realizes the multi-unmanned aerial vehicle target encircling method based on the internal and external reward mechanism reinforcement learning when executing the program.
[0016] The application adopts the above at least one technical scheme to achieve the following beneficial effects: In the application, the internal and external reward mechanism is innovatively designed to improve the explorativeness and convergence of the strategy model in the training stage; the loss function of the strategy model is constructed through the external reward mechanism and the internal reward mechanism, the external reward mechanism is derived from the individual behaviors of the unmanned aerial vehicles, these behaviors are helpful to achieve the group target and represent the interests of the unmanned aerial vehicle group, and are used to enhance the cooperation between the unmanned aerial vehicles to achieve the overall target of the unmanned aerial vehicle group; the internal reward mechanism is the feedback signal obtained by the unmanned aerial vehicle from its own behavior according to the change of its state, and represents the advantages of each individual unmanned aerial vehicle and is used to guide the unmanned aerial vehicle to explore the environment and acquire new knowledge; in this way, for the multi-unmanned aerial vehicle cooperative target encircling problem, the control input vector of the target unmanned aerial vehicle is determined through the strategy model, which not only helps the unmanned aerial vehicle to fully explore the environment and avoid the dilemma of insufficient convergence, but also plans an efficient and safe encircling path for each unmanned aerial vehicle, and solves the pain points of insufficient exploration and poor convergence performance of the traditional reinforcement learning method in a complex environment with sparse or inaccurate rewards, thereby efficiently helping the unmanned aerial vehicle to complete the cooperative target encircling task. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are included to provide a further understanding of the application and constitute a part of this application, illustrate certain illustrative embodiments of the application and together with the description serve to explain the application. In the drawings:
[0018] Figure 1 A flowchart of a multi-unmanned aerial vehicle target encircling method based on the internal and external reward mechanism reinforcement learning is provided for the application; Figure 2 A training process diagram of a strategy model is provided for the application; Figure 3 A framework diagram of an internal and external reward mechanism is provided for the application; Figure 4 A computer device diagram for realizing the multi-unmanned aerial vehicle target encircling method based on the internal and external reward mechanism reinforcement learning is provided for the application. DETAILED DESCRIPTION
[0019] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in connection with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work belong to the protection scope of the present application.
[0020] In the process of developing the trapping strategy, a large number of constraints must also be fully considered. For example, the anti-collision constraint is crucial, and the unmanned aerial vehicles need to maintain a safe distance to avoid collision accidents, which not only relates to the normal operation of the equipment, but also relates to the success or failure of the entire trapping task. At the same time, the energy consumption constraint cannot be ignored, and the energy of the unmanned aerial vehicles is limited, which must be reasonably planned in the premise of ensuring the completion of the task to maximize the saving of energy. Due to these multiple constraints and the complexity of the dynamic environment, it is extremely difficult to find the optimal trapping strategy. From the perspective of computational complexity, the multi-unmanned aerial vehicle cooperative target trapping problem has the NP-hard characteristic, which means that the difficulty of solving will increase exponentially with the increase of the problem size. To find the optimal solution under such complex conditions puts high requirements on related technologies and algorithms.
[0021] In recent years, a large number of research work has been carried out on the target surrounding problem, and a series of feasible solutions with different constraints and objectives have been proposed. In general, the existing solutions to the cooperative target surrounding problem can be roughly divided into three categories: deterministic methods, heuristic methods, and reinforcement learning methods.
[0022] Deterministic methods rely on precise mathematical models and pre-defined control rules to coordinate the actions of unmanned aerial vehicles. By using detailed environment modeling and strategic planning, deterministic methods can optimize the behavior sequence of unmanned aerial vehicles and effectively surround static or dynamic targets. In deterministic methods, model predictive control (MPC) is undoubtedly the most typical representative algorithm, which has been applied to solve the cooperative target trapping problem of multiple unmanned aerial vehicles by many scholars. However, deterministic methods are usually designed and adopted in specific application scenarios, and their effectiveness mainly depends on the accuracy of mathematical models and control rules. In addition, deterministic methods lack learning ability and cannot quickly adapt to dynamic changes, resulting in unsatisfactory solutions in open and complex environments.
[0023] Heuristic methods simplify complex problems by leveraging existing knowledge and inspiration to quickly find approximate or satisfactory solutions. They are often used for problems that cannot be solved using traditional exact algorithms, such as combinatorial optimization problems and the traveling salesman problem. In the multi-UAV cooperative target encirclement problem, heuristic methods not only combine heuristic information and rules to guide the encirclement process but also can quickly adapt to dynamically changing environments and effectively adjust strategies to achieve the task. However, heuristic methods rely heavily on initial parameter settings and cannot effectively explore the entire solution space, easily getting trapped in local optima. Furthermore, heuristic methods are typically highly dependent on assumptions and specific problem structures. If the problem characteristics and assumptions are not met, the effectiveness of heuristic methods may be severely compromised.
[0024] Reinforcement learning allows agents to continuously learn and optimize their decision-making strategies through real-time interaction with the environment. In these methods, agents do not need to make any assumptions or build precise environmental models; they can obtain optimal behavioral strategies through trial and error, while quickly adapting to different environments and tasks. Due to its unique trial-and-error learning mechanism, reinforcement learning has been widely applied in recent years to solve problems related to multi-UAV cooperative target encirclement. However, traditional reinforcement learning methods place high demands on the state space of the environment and the action space of the agent. Especially in complex optimization problems with continuous and high-dimensional state spaces observed by a large number of agents, these methods may encounter combinatorial explosion, making it difficult to find the optimal solution. Furthermore, when solving multi-UAV cooperative target encirclement problems, especially in complex environments with sparse or imprecise rewards, traditional reinforcement learning methods often suffer from insufficient exploration and poor convergence performance. The design of efficient learning frameworks and salient reward functions remains a key focus of current reinforcement learning research.
[0025] Although the multi-UAV cooperative target encirclement problem has been extensively studied by experts and scholars both domestically and internationally in numerous publications, the existing three types of solutions still have many problems that urgently need to be addressed. Deterministic methods, while relatively simple to implement, have significant limitations. They lack the ability to effectively learn and adapt to dynamic environments, making it difficult to make flexible adjustments and optimizations based on environmental changes in a short period. When placed in open and complex environments, this method often fails to provide satisfactory solutions and struggles to cope with the challenges posed by various uncertainties and variability. Heuristic methods, while having clear advantages in efficiency, simplicity, practicality, and adaptability, are prone to getting stuck in local optima due to their over-reliance on initial parameter settings, making it difficult to effectively explore the entire solution space. Traditional reinforcement learning methods, while helping UAVs intelligently chase and capture enemy targets through their unique trial-and-error learning mechanism and continuously optimize their strategies based on environmental feedback, thus improving the efficiency of encirclement operations and adaptability to dynamic environments, face challenges in defining reward functions for good or bad behavior and providing explicit reward signals due to the extremely large number of state-action combinations in the multi-UAV cooperative target encirclement problem. Especially in complex environments where rewards are sparse or imprecise, traditional reinforcement learning methods generate reward signals by evaluating the effects of behavior in the external environment, which often leads to problems such as insufficient exploration and poor convergence performance.
[0026] To address the problems of traditional reinforcement learning methods, this invention proposes a multi-UAV target encirclement method based on an intrinsic and extrinsic reward mechanism. This mechanism overcomes the inaccuracy of reward signals during the learning process, effectively enabling cooperative target encirclement of multiple UAVs. The intrinsic and extrinsic reward mechanism emphasizes that the reward signals guiding the agent's learning consist of two parts: extrinsic and intrinsic rewards. Extrinsic rewards originate from the individual actions of the agents, which contribute to achieving the group's goals and represent the interests of the UAV swarm. These rewards enhance cooperation among UAVs to achieve the overall swarm objective. On the other hand, intrinsic rewards are feedback signals obtained by the agents from their own actions based on changes in their state. These rewards represent the strengths of each individual agent and guide them to autonomously explore the environment and acquire new knowledge. For the multi-UAV cooperative target encirclement problem, this method not only helps agents fully explore the environment and avoids insufficient convergence, but also plans efficient and safe encirclement paths for each UAV, enabling the enemy target to be quickly and evenly surrounded.
[0027] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0028] Figure 1 This is a flowchart illustrating a multi-UAV target capture method based on reinforcement learning with internal and external reward mechanisms, as described in this invention. The method specifically includes the following steps: S101, design internal and external reward mechanisms; the internal and external reward mechanisms include external reward mechanisms and internal reward mechanisms. The external reward mechanism is constructed based on the feedback signals given by the individual behavior of the drone in the environment, and is used to enhance the cooperation between drones. The internal reward mechanism is constructed based on the feedback signals obtained by the drone from its own behavior based on the changes in the drone's state, and is used to guide the drone to autonomously explore the environment and acquire new knowledge.
[0029] A corresponding system model is constructed based on the characteristics of the multi-UAV cooperative target encirclement problem. In the task scenario of multi-UAV cooperative target encirclement, it is assumed that... drone Need to include A static obstacle In an environment for moving targets A coordinated encirclement operation is conducted. Each autonomous drone has a different flight speed and energy consumption. To simplify the drone model, it is assumed that all drones fly at the same altitude, and the physical properties of the drones, such as sharpness, size, and color, are ignored. Therefore, the... drone The dynamic model can be defined as:
[0030] (1); in, express The first derivative, express The first derivative, , and They represent the first drone exist Position, velocity, and control input vector at any given time. and It is the damping factor. Used to represent drones Safe distance from obstacles or other drones.
[0031] During the encirclement and capture operation, the drone requires propulsion energy to maintain high altitude and fly to the designated location. The drone's propulsion power consists of three components: blade profile power to overcome the drag generated by blade rotation, induced power to overcome induced drag and generate lift, and parasitic power to overcome parasitic friction drag generated during airborne motion. Therefore, the drone's propulsion power can be modeled as follows:
[0032] (2); in, It is the flight speed of the drone, coefficient to The modeling parameters depend on the inherent characteristics of the drone and environmental conditions. Used to represent Maximum available energy consumption in target capture missions.
[0033] Static obstacles mainly consist of enemy anti-aircraft guns, radar threat areas, and natural terrain. These obstacles are geographically fixed and will not move during the encirclement mission. To simplify the obstacle model, static obstacles are abstracted as circles with a fixed radius. Use tuples It means that among them and They represent obstacles. The center coordinates and radius of the circle. Simultaneously, the moving target... Considered to have a starting position and fixed moving speed The point mass. Therefore, the moving target at time point... The location can be accessed via Perform the calculation.
[0034] Among them, the design of external reward mechanisms includes: designing external reward mechanisms based on the constraints of multi-drone collaborative target capture missions.
[0035] Based on the constraints and optimization objectives of a multi-UAV cooperative target acquisition task, a problem formula for the task is established. The constraints of the multi-UAV cooperative target acquisition task include collision avoidance constraints, motion continuity constraints, and energy consumption constraints.
[0036] (1) Collision Avoidance Constraint: The UAV should not collide with obstacles, other UAVs, or moving targets. That is, at any given time, the distance between any two UAVs or between a UAV and an obstacle should be greater than a predefined safe distance. The time consumed to complete the target encirclement task is represented by the following collision avoidance constraint:
[0037] (3); in, This indicates the time consumed in the target capture mission. This indicates the number of drones used in the target capture mission. Indicates the first A drone in Location at any given moment Indicates the first A drone in Location at any given moment Indicates in time and The Euclidean distance between them Indicates a safe distance. Indicates the number of obstacles. Indicates the first The coordinates of the center of each obstacle Indicates in time and The Euclidean distance between them Indicates the first The radius of the obstacle Indicates the target to be apprehended is in Location at any given moment Indicates in time and The Euclidean distance between them.
[0038] (2) Motion Continuity Constraint: Each UAV should take off from its initial position and continuously update its motion state until the target encirclement mission is completed. When the moving target is encircled, all UAVs should maintain a fixed distance from the target and have the same moving speed as the target. The motion continuity constraint is as follows:
[0039] (4); in, Indicates the first The initial position of the drone, i.e., the drone at the start of the target capture mission. The starting position, Indicates the first The initial speed of the drone Indicates a preset fixed distance. Indicates the speed of the target to be captured. Indicates the first A drone in The speed of time, This indicates the preset time interval. Indicates the first A drone in The speed of time, Indicates that drones are in The control input vector at time t, Indicates the first A drone in The location at any given moment.
[0040] (3) Energy Constraint: Each UAV should complete its target encirclement mission before running out of energy. Therefore, it is necessary to ensure that the energy consumption of the UAV at any given time does not exceed the available value. The energy constraint is:
[0041] (5); in, Indicates the first A drone in The propulsion power at any moment Indicates the first The maximum available energy consumption of a drone.
[0042] The optimization objective of a multi-UAV cooperative encirclement mission is to find a safe flight path for each UAV to quickly encircle and capture the enemy target while satisfying collision avoidance, motion continuity, and energy consumption constraints. Its mathematical form can be expressed as: .
[0043] Before using reinforcement learning algorithms based on internal and external reward mechanisms to solve the multi-UAV cooperative target capture problem, it is necessary to convert the constraints of the multi-UAV cooperative target capture task established above into a Markov decision process model.
[0044] (1) State Design: The environmental information observed by each UAV includes four aspects: its own position and velocity, the position and velocity of neighboring UAVs, the position of detected obstacles, and the position and velocity of moving targets. Using Indicates drone exist The environmental state information at any given time is represented as follows:
[0045] (6).
[0046] (2) Motion design: The UAV can control the input vector The parameters in the data are used to update its motion state (i.e., position and velocity). Therefore, any drone exist Moment of action It can be represented as:
[0047] (7).
[0048] (3) External reward mechanism design: The rewards obtained by the agent from the external environment should fully consider the constraints and objectives of the multi-UAV cooperative target capture problem. Therefore, the following can be used: , , and They represent the first A drone in The rewards are derived from collision avoidance constraints, motion continuity constraints, energy consumption constraints, and the objective at each step. Furthermore, to design a more reasonable and effective reward scheme, a relative motion model is used to analyze the calculation process of these four types of rewards:
[0049] (8); in, , and They represent the first drone With the drone exist The relative position, relative speed, and relative yaw angle at any given moment.
[0050] During the drone's pursuit, the reward generated by the collision avoidance constraint... It can be expressed by the following formula: (9); in, and The first The rewards earned by a drone for avoiding collisions with other drones and with other static obstacles. and These are variable weight parameters that take non-negative values. and They represent the first The drone and the first One obstacle The relative position and relative yaw angle at any given moment.
[0051] Based on the above analysis of motion continuity constraints, a reward system for the UAV during the encirclement process can be designed, which is generated by the motion continuity constraints. as follows: (10); in, and These are predefined weighting factors; it is important to note that... and The value may change as the drone rapidly approaches or follows the target's movement.
[0052] Based on the above analysis of energy consumption constraints, all drones should complete the target encirclement mission before running out of energy. Therefore, the reward generated by the energy consumption constraint... The following formula can be used for calculation: (11); in, It is a fixed repulsion coefficient for the energy consumption of drones. It is the first The drone was shut down. Energy consumed constantly.
[0053] To incentivize drones to successfully complete target encirclement missions, a target reward is designed for drones upon successful encirclement. as follows: (12); in, It is a positive real number, such as 10000. Combining the four types of rewards mentioned above, the [number]th [reward] can be calculated. drone exist The external rewards received at any given time, i.e., the external reward mechanism, are as follows:
[0054] (13); in, Indicates the first A drone in External rewards at all times , , and This represents the adjustable weighting parameter for each reward.
[0055] Optionally, an intrinsic reward mechanism is designed, including: constructing an intrinsic reward mechanism based on the difference between the state error reward and the extrinsic reward; the state error reward represents the difference between the predicted state of the UAV and the target state; the intrinsic reward mechanism is as follows: (14); in, Indicates the first A drone in Intrinsic rewards of moments Indicates the reward for state error. It is a single-layer fully connected network used to compare state error rewards. External rewards The differences between them.
[0056] S102. Based on the external and internal reward mechanisms, a loss function is constructed, and a policy model is trained based on the loss function.
[0057] loss function for: (15); in, , express The L2 norm.
[0058] Optionally, the policy model includes a policy network and an evaluation network; the training process of the policy model includes: inputting the historical environmental state information of the sample UAV at the first moment into the initial policy network, predicting the control input vector of the sample UAV at the second moment, and updating the predicted historical environmental state information of the sample UAV at the second moment with the control input vector at the second moment; determining the loss function value based on the predicted historical environmental state information at the second moment and the target historical environmental state information at the second moment; updating the initial policy network and the initial evaluation network based on the loss function value, and continuing to iteratively train the updated initial policy network and the initial evaluation network until the preset iteration conditions are met, and determining the initial policy network and the initial evaluation network that meet the iteration conditions as the policy network and evaluation network of the policy network model.
[0059] Specifically, such as Figure 2 As shown, Figure 2 The flowchart illustrates the construction process of the policy model. The multi-agent reinforcement learning framework based on internal and external reward mechanisms is described below. It consists of several intelligent agents, each controlling a corresponding single drone. Therefore, for any given drone... In this context, the corresponding intelligent agent consists of a policy network (Actor) and an evaluation network (Critic), which are respectively used as... and To represent, where These represent the parameters of the policy network. These represent the parameters used to evaluate the network.
[0060] For the parameters of the policy network The policy is updated using a gradient ascent approach to maximize the expected reward. (16); in, It contains a large number of joint experience sequences from drone swarms. The experience replay pool. It is the global joint state information of the drone swarm. This indicates the actions of all drones. Indicates that the drone swarm is in a joint state The following joint actions were taken. External rewards received. This refers to the joint state information at the next time step. It's important to note that when solving for the policy gradient, besides the agent... Selected action Apart from that, all other data comes from the experience replay pool. .
[0061] For evaluating the parameters of the network Iterative updates are performed using a time-difference approach, and the loss function is calculated as follows: (17); in, and Representing intelligent agents respectively The target policy network and target evaluation network have the same network structure as the policy network and evaluation network, respectively, and are mainly used for overestimation problems in environment reinforcement learning. It is a drone The total reward value obtained is determined by the internal and external reward mechanisms based on external rewards from environmental feedback. Intrinsic rewards derived from changes in internal state The weighted calculation yielded the result.
[0062] The intrinsic and extrinsic reward mechanism mainly consists of an encoder that extracts and compresses features from the original state information for further processing and analysis, a feedforward module that predicts the feature vector of the next state, and a reward prediction error module that generates intrinsic rewards by analyzing prediction errors. First, the encoder is used to process the agent's state... and the next state Mapped to the corresponding feature vector and In the middle. Then, and actions The input is fed into the feedforward module to predict the agent's behavior. The feature vector of the next state :
[0063] (18); in, It is an intelligent agent A single-layer fully connected network (forward module) whose parameters are minimized by the feature vector and Differences between Update: (19).
[0064] After passing through the forward module, Later, through comparison and The similarity between them is used to calculate the current state error reward using the Euclidean norm (L2 norm) calculation method. : (20).
[0065] Subsequently, by rewarding the state error and external rewards Simultaneously, the error is input into the reward prediction module to obtain the agent's information. Intrinsic rewards : (twenty one); in, It is also a single-layer fully connected network (reward prediction error module) used for intelligent agents. Comparison of state error reward External rewards The difference between them is used to generate intrinsic rewards. Since the optimization objective of the reward prediction error module is to maximize extrinsic rewards... and intrinsic rewards Minimize the error between them, therefore loss function It can be represented as:
[0066] (twenty two).
[0067] Combining the above formula, the loss function of the internal and external reward mechanisms can be expressed as: (twenty three); in, This is used to balance the importance of the feedforward module and the reward prediction error module. This is achieved by minimizing the loss function. To update the network and The parameters.
[0068] like Figure 3 As shown, Figure 3 This is a framework diagram of internal and external reward mechanisms.
[0069] In one embodiment, an extrinsic-and-intrinsic reward-based multi-agent reinforcement learning (EIR-MARL) algorithm is used to solve the Markov decision process model established for the multi-UAV cooperative target encirclement problem. This is a multi-UAV target encirclement method based on extrinsic-and-intrinsic reward-based reinforcement learning. The training process of its policy model specifically includes: S201, for each intelligent agent The network parameters involved are initialized, specifically including the policy network. Evaluation Network Target-Policy Network Target evaluation network And the network involved in internal and external reward mechanisms and the Internet .
[0070] S202, Initialize the experience replay pool .
[0071] S203, initialize the environment and observe each agent. initial state .
[0072] S204, for each agent Perform the same operation: from the joint state Read its own state Then it is input into its own policy network. In the middle, the action is obtained by sampling. .
[0073] S205, Environment based on joint action Feedback is given to the agent to execute the joint action. External rewards obtained later At the same time, return to the next union state. .
[0074] S206 combines the states, actions, external rewards, and next states of all agents to construct a joint experience sequence. And store it in the experience replay pool. middle.
[0075] S207, if experience replay pool If the number of joint empirical sequences in the sample does not meet the sampling criterion, the current joint state is updated. Then, proceed to step (4) to continue generating experience sequences. If the number of experience sequences reaches the target, then retrieve them from the experience replay pool. A batch of random samples of size The joint experience sequence is then used to perform subsequent training steps.
[0076] For each agent, the joint experience sequence obtained from sampling is used. Perform the same training operations: S208 uses internal and external reward mechanisms to calculate drones Total reward value obtained : (twenty four); in, and This represents an adjustable weighting factor used to weigh the importance of extrinsic and intrinsic rewards.
[0077] S209, based on total reward value The evaluation network is calculated based on the formula of temporal difference. loss value : (25).
[0078] Then based on the calculations... The evaluation network is updated using gradient descent. parameters .
[0079] S210, Computational Policy Network policy gradient : (26).
[0080] Then based on the policy gradient Update the policy network using gradient ascent. parameters .
[0081] S211, Update the target policy network using a soft update method. and target evaluation network parameters and : (27); (28); in, This represents the soft update rate.
[0082] S212, Calculate the loss value of internal and external reward mechanisms. The forward module is updated using gradient descent. and reward prediction error module The parameters are then used to update the current joint state. Then, proceed to step S204 to continue execution until the environment's termination condition is met, at which point this round of training ends.
[0083] The iteration steps S203 to S212 are repeated until the sum of the reward values returned by the algorithm in each round of training tends to stabilize.
[0084] S103, input the historical environmental state information of the target UAV at the previous moment into the strategy model to obtain the control input vector of the target UAV at the current moment, and update the environmental state information of the target UAV at the current moment with the control input vector at the current moment; the target UAV is any UAV in the target capture mission; the environmental state information includes its own speed and position, as well as the speed and position of the target to be captured.
[0085] Optionally, the environmental state information of the target UAV at the previous moment is input into the policy model to obtain the control input vector of the target UAV at the current moment, including: inputting the environmental state information of the target UAV at the previous moment into the policy network to obtain the control input vector of the target UAV at the current moment.
[0086] S104. If the Euclidean distance between the positions of all drones and the position of the target to be captured is a preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured, then the target capture mission is determined to be completed.
[0087] Based on the environmental status information of all drones at the current moment, determine the relationship between the Euclidean distance between the position of all drones and the position of the target to be captured and the preset fixed distance, as well as the relationship between the speed of all drones and the speed of the target to be captured. If the Euclidean distance between the position of all drones and the position of the target to be captured is not the preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured, then continue to execute step S103 to update the environmental status information of the drones until the Euclidean distance between the position of all drones and the position of the target to be captured is the preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured.
[0088] If the Euclidean distance between the positions of all drones and the position of the target to be captured is a preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured, then the target capture mission is considered complete.
[0089] This invention studies a multi-agent reinforcement learning algorithm based on internal and external reward mechanisms and applies it to the multi-UAV cooperative encirclement problem. The main innovations of this invention are as follows:
[0090] (1) A system model of the multi-UAV cooperative target capture problem was established, and its constraints and optimization objectives were analyzed to abstract it into a multi-constraint combinatorial optimization problem. At the same time, a precise problem formula was given, which more clearly described the essential characteristics of the multi-UAV cooperative target capture problem.
[0091] (2) A Markov decision process model for multi-UAV cooperative target capture task was established, which enables reinforcement learning algorithms to be applied more effectively to multi-UAV cooperative target capture problem.
[0092] (3) Based on the traditional multi-agent reinforcement learning algorithm, a multi-agent reinforcement learning algorithm based on internal and external reward mechanism is proposed. The internal and external reward mechanism solves the pain points of insufficient exploration and poor convergence of the traditional reinforcement learning algorithm in the sparse reward environment, so that the proposed algorithm can solve the multi-UAV cooperative target capture problem more efficiently.
[0093] Compared with the prior art, the present invention has two main advantages: (1) At the algorithm design level, a multi-agent reinforcement learning algorithm based on internal and external reward mechanism is proposed, which can use internal and external reward mechanism to solve the shortcomings of traditional reinforcement learning algorithm in sparse reward environment, such as insufficient exploration and poor convergence.
[0094] (2) In practical applications, the proposed EIR-MARL algorithm can be combined with the multi-UAV cooperative target capture problem, which can minimize the time for UAV swarms to complete the target capture task and promote the application of UAV swarm technology in other fields.
[0095] When applying the multi-UAV target capture method based on internal and external reward mechanisms provided by this invention, it is not necessary to consider... Figure 1 The steps shown are executed in sequence. The specific execution order of each step can be determined as needed, and this invention does not impose any restrictions on it.
[0096] The above describes a multi-UAV target acquisition method based on reinforcement learning with internal and external reward mechanisms, provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding multi-UAV target acquisition device based on reinforcement learning with internal and external reward mechanisms, the device comprising: The acquisition module is used to design internal and external reward mechanisms. These mechanisms include external and internal reward mechanisms. External reward mechanisms are constructed based on feedback signals from the individual behavior of the drone in relation to the environment, and are used to enhance cooperation between drones. Internal reward mechanisms are constructed based on feedback signals obtained from the drone's own behavior based on changes in its state, and are used to guide the drone to autonomously explore the environment and acquire new knowledge. The training module is used to construct a loss function based on the external and internal reward mechanisms, and to train the policy model based on the loss function. The execution module is used to input the historical environmental state information of the target UAV at the previous moment into the strategy model, obtain the control input vector of the target UAV at the current moment, and update the environmental state information of the target UAV at the current moment with the control input vector at the current moment; the target UAV is any UAV in the target capture mission; the environmental state information includes its own speed and position, as well as the speed and position of the target to be captured; The encirclement module is used to determine that the target encirclement mission is completed if the Euclidean distance between the positions of all drones and the position of the target to be encircled is a preset fixed distance, and the speed of all drones is the same as the speed of the target to be encircled.
[0097] Specific limitations regarding the multi-UAV target capture device based on reinforcement learning with internal and external reward mechanisms can be found in the limitations of the multi-UAV target capture method based on reinforcement learning with internal and external reward mechanisms mentioned above, and will not be repeated here. Each module in the aforementioned multi-UAV target capture device based on reinforcement learning with internal and external reward mechanisms can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0098] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The proposed method is a multi-UAV target capture method based on reinforcement learning with internal and external reward mechanisms.
[0099] The present invention also provides Figure 4 The schematic diagram of the computer device shown is as follows: Figure 4 As shown, at the hardware level, this computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above. Figure 1 The proposed method is a multi-UAV target capture method based on reinforcement learning with internal and external reward mechanisms.
[0100] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this invention.
Claims
1. A method for multi-UAV target acquisition based on reinforcement learning with internal and external reward mechanisms, characterized in that, include: Design internal and external reward mechanisms; The internal and external reward mechanisms include external reward mechanisms and internal reward mechanisms. External reward mechanisms are constructed based on feedback signals given by the individual behavior of the drone in the environment, and are used to enhance cooperation between drones. Internal reward mechanisms are constructed based on feedback signals obtained from the drone's own behavior based on changes in the drone's state, and are used to guide the drone to autonomously explore the environment and acquire new knowledge. Based on the external and internal reward mechanisms, a loss function is constructed, and a policy model is trained based on the loss function. The environmental state information of the target UAV at the previous moment is input into the strategy model to obtain the control input vector of the target UAV at the current moment, and the environmental state information of the target UAV at the current moment is updated with the control input vector at the current moment; the target UAV is any UAV in the target capture mission. Environmental status information includes its own speed and position, as well as the speed and position of the target to be captured; If the Euclidean distance between the positions of all drones and the position of the target to be captured is a preset fixed distance, and the speed of all drones is the same as the speed of the target to be captured, then the target capture mission is considered complete.
2. The method according to claim 1, characterized in that, Design an external reward mechanism, including: Based on the constraints of the multi-UAV cooperative target capture mission, an external reward mechanism is designed; the constraints of the multi-UAV cooperative target capture mission include collision avoidance constraints, motion continuity constraints, and energy consumption constraints. The collision avoidance constraints are: ; in, This indicates the time consumed in the target capture mission. This indicates the number of drones used in the target capture mission. Indicates the first A drone in Location at any given moment Indicates the first A drone in Location at any given moment Indicates in time and The Euclidean distance between them Indicates a safe distance. Indicates the number of obstacles. Indicates the first The coordinates of the center of each obstacle Indicates in time and The Euclidean distance between them Indicates the first The radius of the obstacle Indicates the target to be apprehended is in Location at any given moment Indicates in time and The Euclidean distance between them; The motion continuity constraint is: ; in, Indicates the first The starting position of the drone. Indicates the first The initial speed of the drone Indicates a preset fixed distance. Indicates the speed of the target to be captured. Indicates the first A drone in The speed of time, This indicates the preset time interval. Indicates the first A drone in The speed of time, Indicates that drones are in The control input vector at time t, Indicates the first A drone in The position at that moment; Energy consumption constraints are: ; in, Indicates the first A drone in The propulsion power at any moment Indicates the first The maximum available energy consumption of a drone.
3. The method according to claim 2, characterized in that, The external reward mechanism is as follows: ; in, Indicates the first A drone in External rewards at all times This represents the reward generated by the collision avoidance constraint. This represents the reward resulting from the constraint of motion continuity. This represents the reward generated by energy consumption constraints. Indicates the first The target reward obtained by the drone after successfully capturing it. , , and These represent the adjustable weighting parameters for each reward; ; ; ; ; in, and The first The rewards are earned by avoiding collisions with other drones and other static obstacles. and These are variable weight parameters that take non-negative values; , and They represent the first The drone and the first A drone in The relative position, relative speed, and relative yaw angle at any given moment; and They represent the first The drone and the first One obstacle The relative position and relative yaw angle at any given time. and This represents a predefined weighting factor. It is a fixed repulsion coefficient for the energy consumption of drones. It is the first The drone was shut down Energy consumed constantly It is a positive real number.
4. The method according to claim 3, characterized in that, Design an intrinsic reward mechanism, including: An intrinsic reward mechanism is constructed based on the difference between the state error reward and the extrinsic reward; the state error reward represents the difference between the predicted state of the UAV and the target state; the intrinsic reward mechanism is as follows: ; in, Indicates the first A drone in Intrinsic rewards of moments Indicates the reward for state error. It is a single-layer fully connected network used to compare state error rewards. External rewards The differences between them.
5. The method according to claim 4, characterized in that, loss function for: ; in, , express The L2 norm.
6. The method according to claim 1, characterized in that, The policy model includes a policy network and an evaluation network; the training process of the policy model includes: The historical environmental state information of the sample UAV at the first moment is input into the initial policy network to predict the control input vector of the sample UAV at the second moment, and the predicted historical environmental state information of the sample UAV at the second moment is updated with the control input vector at the second moment. The loss function value is determined based on the predicted historical environmental state information at the second time step and the target historical environmental state information at the second time step. The initial policy network and initial evaluation network are updated based on the loss function value. The updated initial policy network and initial evaluation network are then iteratively trained until the preset iteration conditions are met. The initial policy network and initial evaluation network that meet the iteration conditions are then determined as the policy network and evaluation network of the policy network model.
7. The method according to claim 6, characterized in that, The environmental state information of the target UAV at the previous moment is input into the policy model to obtain the control input vector of the target UAV at the current moment, including: The environmental state information of the target UAV at the previous moment is input into the policy network to obtain the control input vector of the target UAV at the current moment.