A method for unmanned aerial vehicle cluster cooperative target search based on reinforcement learning
By constructing a 3D motion model and optimizing resource allocation in collaborative target search of UAV swarms, and combining the heuristically embedded MAPPO algorithm, the problems of search efficiency and resource allocation of UAV swarms in dynamic environments are solved, achieving efficient target discovery and algorithm generalization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2025-08-29
- Publication Date
- 2026-08-04
AI Technical Summary
Existing UAV swarm collaborative target search methods suffer from low search efficiency, suboptimal allocation of computing resources, and poor algorithm generalization when facing dynamically changing environments, especially in three-dimensional space and large-scale UAV swarm scenarios.
Based on the heuristically embedded MAPPO algorithm, combined with the motion model of UAVs in three-dimensional continuous space, a collaborative target search architecture for UAV swarms is constructed. The three-dimensional motion trajectory, charging decision, and unloading position and computational resource allocation of the UAVs are optimized. A Beta distribution update strategy network is adopted, and a safety action controller and multi-head attention mechanism are introduced to design a heuristic energy-saving task unloading method.
It improves the search efficiency and target discovery rate of UAV swarms in dynamic environments, reduces the training difficulty in large-scale UAV swarm cooperation scenarios, and enhances the convergence performance and generalization of the algorithm.
Smart Images

Figure CN120973011B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) target search technology, specifically relating to a UAV swarm cooperative target search method based on reinforcement learning. Background Technology
[0002] For the cooperative target search problem in UAV swarms, current research mainly focuses on heuristic-based and reinforcement learning-based search schemes. Heuristic schemes rely on prior knowledge of a specific environment and typically exhibit limited flexibility and adaptability, especially in dynamically changing search environments. In contrast, reinforcement learning-based methods find near-optimal solutions through continuous exploration and trial and error in unknown environments, effectively adapting to dynamic changes in the search environment and demonstrating significant advantages in solving complex decision-making problems. Therefore, reinforcement learning-based cooperative target search methods for UAV swarms have become a current research hotspot.
[0003] In recent years, a series of UAV swarm cooperative target search methods based on reinforcement learning have been proposed, and these methods have all achieved good results. For example, a multi-UAV high-low altitude cooperative search architecture based on the MAPPO algorithm is proposed to solve the path planning problem of dynamic targets in a 3D scene for multi-UAV cooperative search; a multi-UAV distributed cooperative target search method based on MADDPG can operate efficiently in complex and large-scale scenes; and the DQN algorithm is used to jointly make optimal task offloading decisions and flight direction selections for multi-UAV cooperative target search, etc.
[0004] However, existing research approaches still face the following technical challenges. First, current studies typically assume that UAVs perform discrete motion on a fixed two-dimensional plane, or simply deploy different UAVs at different heights on different horizontal planes, which does not match real-world scenarios. While some studies use pitch angle, heading angle, and velocity to describe the three-dimensional motion of UAVs, these studies do not establish the relationship between the UAV's trajectory and the uncertainty of the search area information. This low-precision, discrete motion method may reduce the search efficiency of UAVs. Second, existing studies do not consider the joint optimization of computational offloading decisions and computational resource allocation during the UAV search process. Instead, they use a fixed CPU computing frequency to perform search image processing locally on the UAV, or use a fixed transmission power to offload image processing tasks to a remote base station. This single optimization approach reduces the task completion rate and search efficiency of UAVs for latency-sensitive search tasks. Finally, in UAV swarm collaborative search scenarios, since each UAV can only observe a local area, UAVs need to communicate with a variable number of neighboring UAVs, and the total number of UAVs participating in the search may also change. Existing research addresses this variation by maintaining sufficient dimensions in the neural network input layer. However, this approach suffers from scalability issues when dealing with a large number of drones. Furthermore, when the number of drones in the search scenario changes, the method requires retraining the reinforcement learning algorithm, resulting in poor generalization. Summary of the Invention
[0005] In view of this, the present invention provides a cooperative target search method for UAV swarms based on reinforcement learning, which realizes action planning for each UAV in the UAV swarm based on the heuristic embedding MAPPO algorithm for the target search task.
[0006] This invention provides a cooperative target search method for UAV swarms based on reinforcement learning, comprising the following steps:
[0007] Based on the motion model of UAVs in three-dimensional continuous space, a collaborative target search architecture for UAV swarms is constructed.
[0008] Under the collaborative target search architecture of UAV swarms, a joint optimization problem is established for UAV 3D motion trajectory, charging decision, search task unloading location and computing resource allocation.
[0009] The MAPPO algorithm based on heuristic embedding solves the joint optimization problem of UAV 3D motion trajectory, charging decision, search task unloading location and computing resource allocation, and obtains the planned actions of each UAV.
[0010] The drones execute predetermined planned actions to complete the drone swarm collaborative target search task.
[0011] Furthermore, the UAV swarm collaborative target search architecture is as follows: the search area is discretized into multiple grids, and the search task of each grid is represented by the amount of search data and the required CPU processing density; a base station with an edge server is set at the center of the search area to provide image recognition services for the target search task, and laser charging piles are set at the boundaries; the target search task is divided into multiple sub-time slots of equal length, including flight sub-time slots and unloading sub-time slots; the UAV has pre-loaded energy and a top-down camera is set at its bottom. In the flight sub-time slot, it moves at a fixed speed in any direction in three-dimensional space, and in the unloading sub-time slot, it charges using laser charging piles or performs target search tasks. When performing target search tasks, the top-down camera captures the search area directly below to form a search observation image, which is divided into multiple observation sub-images; in the unloading sub-time slot, after all observation sub-images have completed image recognition, the UAV updates the target existence probability of the corresponding grid according to the recognition results.
[0012] Furthermore, the image recognition of the observed sub-image is performed by either the drone or by the base station after being unloaded from the drone.
[0013] Furthermore, the joint optimization problem of the UAV's three-dimensional motion trajectory, charging decision, search task unloading location, and computational resource allocation is expressed as:
[0014]
[0015] st
[0016]
[0017]
[0018] in, For UAV f in flight sub-time slot The magnitude of the fixed flight speed within; For UAV f in flight sub-time slot The angle between the direction of the flight velocity within the space and the z-axis in the three-dimensional Cartesian coordinate system; For UAV f in flight sub-time slot The angle between the direction of the flight velocity within the space and the x-axis in the three-dimensional Cartesian coordinate system; Charging strategies for drones; The unloading strategy for the observed sub-image is OK; The CPU frequency for observing sub-images is OK; The transmission power of the observed sub-image ok; T is the set of all sub-slots; K is the set of all grids within the search area; P t (x k y k ) represents the grid (x) at the start of sub-slot t.k y k The probability that the target exists in (u); t (x k y k ) = 1 indicates that a target exists in the grid, u t (x k y k ) = 0 indicates that there is no target in the grid; ue t (x k y k ) = 1 indicates that the target has been detected in the grid at the start of sub-slot t, ue t (x k y k E = 0 indicates that no target was detected in the grid at the start of sub-slot t; t (x k y k w1 represents the information uncertainty of the grid; w2 and w1 are both weighting factors. Let f be the maximum flight speed of the drone; The maximum CPU computing frequency of the drone f; This represents the maximum transmit power of the UAV f. The maximum battery capacity limit for drone f; L s The safe distance between the two drones; L x With L y These are the length and width of the search area, respectively; and These represent the minimum and maximum safe flight altitudes for the drones, respectively; F represents the set of all drones. P represents the actual search area processed by the UAV. t+1 (x k ,y k (ue) t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k )) represents the increase in the probability that the UAV swarm detects the presence of a target within the grid within sub-time slot t; ∑ k∈K (P t+1 (x k ,y k (ue)t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k E represents the increase in the probability that the UAV swarm detects the presence of a target within the search area during sub-time slot t; t (x k y k )-E (t+1) (x k y k ) represents the reduction in information uncertainty of the grid within sub-time slot t; ∑ k∈K (E t (x k y k )-E (t+1) (x k y k )) represents the reduction in information uncertainty of the search region within sub-time slot t;
[0019] C1, C2, C3, C4, C5, C6, C7, and C8 are all constraints. Among them, C1 is the constraint on the UAV's flight speed and direction in the three-dimensional coordinate system; C2 is the UAV's charging scheduling constraint; C3 is the UAV's data unloading position constraint; C4 and C5 are the constraints on the CPU computing frequency and transmission power allocated when the UAV unloads the observed sub-image OK, respectively; C6 is the UAV's energy constraint during the execution of the search task; C7 is the constraint on the safe distance between any two UAVs; and C8 is the UAV's flight range constraint.
[0020] Furthermore, the heuristic embedding MAPPO algorithm includes:
[0021] A distributed partially observable Markov decision process is used to construct a heuristically embedded MAPPO algorithm model for the collaborative target search task of UAV swarms. The model includes: an environment state space, an observation space of the UAVs, an action space of the UAVs, and a reward function. Each agent corresponds to one UAV, and the agent includes a policy network and an evaluation network. The policy network includes an embedding layer, an attention layer, and an output layer, and the evaluation network includes an embedding layer, two levels of attention layers, and an output layer.
[0022] The training process of the agent is as follows: At the beginning of sub-time slot t, the observation state of the UAV is input into the agent's policy network. Intermediate results are calculated through the embedding layer and single-level attention layer of the policy network based on a single-level multi-head attention mechanism. The intermediate results are then input into the output layer to obtain the parameters β1 and β2 of the probability density function of the Beta distribution. The movement actions and charging strategies of the UAV within sub-time slot t are obtained by sampling the Beta distribution. Based on the current position of the UAV and the movement actions output by the policy network, the safety action controller obtains the movement strategy of the UAV. Based on the position of the UAV and the communication channel conditions between the UAV and the base station, the unloading strategy of the observed image and the allocation of computing resources are determined by a heuristic energy-saving task offloading method.
[0023] The global environment state is input into the evaluation network of the agent based on a two-level multi-head attention mechanism to obtain the score of the policy network; the advantage function is calculated based on the score of the policy network and the reward function; the policy network is updated using the objective function based on the advantage function and the ratio of the current policy to the original policy; the policy network is then trained until the preset termination condition is reached, and the planned actions of the UAV are output.
[0024] Furthermore, the environmental state space is:
[0025]
[0026] in, The remaining energy of the drone f at the beginning of sub-time slot t; L represents the three-dimensional position coordinates of the UAV f at the beginning of sub-time slot t; BS and L laser These are the coordinates of the base station and the laser charging station, respectively; P t (x k y k ), k∈K represents the grid (x) at the beginning of sub-slot t. k y k The probability of the target existing; RT represents the number of remaining time steps in one round of reinforcement learning;
[0027] The observation space of the drone is:
[0028]
[0029] in, OBS represents the probability distribution of targets within the grid that can be observed by the UAV's field of view. f The communication range of the UAV f; SNR thr The minimum signal-to-noise ratio (SNR) threshold for communication between drones; f,f′ Let SNR be the signal-to-noise ratio (SNR) between UAV f and UAV f′; when SNR f,f′ ≥SNR thrWhen f′∈OBS f ;otherwise,
[0030] The drone's action space is:
[0031]
[0032] in, and All are continuous variables, determining the flight sub-time slot of the UAV f. Flight trajectory within; As a discrete variable, it determines the unloading sub-slot of the UAV. Internal charging strategy; when season This indicates that the drone f is unloading the sub-time slot. Internally charged from a laser charging station; when season This indicates that the drone f is unloading the sub-time slot. Perform search area probing and probe data unloading;
[0033] The reward function is:
[0034]
[0035] Where CR is the penalty value for UAV f colliding with other UAVs; w1 and w2 are weighting factors balancing target detection and environment search; K is the set of all grids in the search area; ue t (x k y k )∈{0,1} indicates whether a target has been detected in the grid at the start of sub-slot t; P t+1 (x k ,y k (ue) t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k )) represents the increase in the probability that the UAV swarm detects the presence of a target within the grid within sub-time slot t; ∑ k∈K (P t+1 (x k ,y k (ue)t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k E represents the increase in the probability that the UAV swarm detects the existence of a target within the entire search area within sub-time slot t; t (x k ,y k )-E (t+1) (x k ,y k ) represents the reduction in information uncertainty of the grid within sub-time slot t; ∑ k∈K (E t (x k ,y k )-E (t+1) (x k ,y k )) represents the reduction in information uncertainty of the entire search region within sub-time slot t.
[0036] Furthermore, the heuristic-based energy-saving task offloading method is as follows:
[0037] S1. Variables related to the heuristic energy-saving task unloading method for UAV initialization;
[0038] S2. In each sub-time slot t, the UAV f traverses the search area. For each observed sub-image, calculate the minimum CPU computation frequency required for local processing. and the minimum transmission power required to transmit it to the base station. in, The calculation formula is:
[0039]
[0040] Where B is the channel bandwidth. For noise power, CG f,bs (t) represents the hybrid channel gain between the UAV f and the base station bs, N ok To measure the amount of detection data for the observed sub-images, The length of the sub-time slot unloaded within the sub-time slot;
[0041] like and The maximum computational frequency of the UAV f is given. If the maximum transmission power of UAV f is given, then execute S2.1; if and Then execute S2.2; if and Then execute S2.3; if and Then execute S2.4;
[0042] Step S2.1: The UAV adds the currently observed sub-image to the set of observed sub-images that cannot be unloaded, and then executes S2.5;
[0043] Step S2.2: The UAV unloads the current observation sub-image to the base station and adds it to the set of observation sub-images that can only be unloaded to the base station, and then executes S2.5;
[0044] S2.3, The UAV can only update the variable tFre, which is the sum of the CPU frequencies allocated to all the observation sub-images that are unloaded to the local machine and the observation sub-images that are unloaded to the local machine. Then execute S2.5.
[0045] The update method for FO is as follows:
[0046]
[0047] Among them, κ f A coefficient related to the power f of the UAV;
[0048] S2.4 The UAV unloads the current observation sub-image to its local machine and updates the set of observation sub-images FBO and tFre that can be unloaded to the local machine and the base station, then executes S2.5; the FBO is updated in the following way:
[0049]
[0050] like Then the drone f will unload ok to the base station, the drone f will update the observation sub-image set BO that can only be unloaded to the base station, and then execute S2.5;
[0051] S2.5, UAV f traverses the search area If all observed sub-images are OK, execute S3;
[0052] S3, if Execute S4 if necessary; otherwise, execute S5.
[0053] S4. The drone will sort the observation sub-images that can be unloaded to the local area and the base station;
[0054] S5. The final unloading strategy for all observation sub-images generated by the UAV includes: BO, FO, FBO, NO, and computing resource allocation strategy. and
[0055] Furthermore, the policy network includes:
[0056] The local observations of the UAV f are divided into its own characteristics. And its communication range OBS f Features of other UAVs j
[0057] Using an embedded layer fully connected network, from o respectively f and o j Extracting feature e f and e j And send it to the multi-head attention head to obtain the attention value x. f x f for:
[0058]
[0059] Among them, e f Let e be the eigenvalue of the drone f. j (j∈OBS f ) represents the communication range of the UAV f (OBS) f The eigenvalues of other UAVs j; W k W q and W v These are the weight matrices for keys, queries, and values, respectively. For W k The transpose of d; k As the key dimension; α f,j Score for attention;
[0060] Finally, x f With e f The data are concatenated and sent to the fully connected network of the output layer to obtain the probability distribution of the drone f's actions.
[0061] Furthermore, the evaluation network includes:
[0062] Global state information is divided into environmental information S. m ={L BS L laser ,RT,P t (x k y k Information on )} and drone f Embedded layer fully connected network from S m With Sf Extract two types of features e respectively m and e f (f∈F), e m For environmental information S m eigenvalues, e f (f∈F) represents the information S of the drone f. f eigenvalues;
[0063] e f (f∈F) is sent to |F| first-level multi-head attention heads, which extract the attention value x respectively. f (f∈F), where |F| is the total number of drones;
[0064] Connect e f and x f It is then sent to a fully connected network to obtain the attention output vector g. f (f∈F); e m and g f (f∈F) is sent to the two-level multi-head attention head to obtain the attention value x. m ;
[0065] series e m and x m It is then sent to the fully connected network at the output layer to obtain the final state value evaluation score.
[0066] Furthermore, the energy of the drone f at the start of sub-time slot t+1 is:
[0067]
[0068] in, and These represent the energy of the drone f at the beginning of sub-time slot t and sub-time slot t+1, respectively. Given the battery capacity limit of the drone f, the min(·) function ensures that the drone's energy consumption during charging will not exceed the battery capacity limit. The energy consumed by the UAV f to transmit the observed sub-images to the remote base station. For UAV f in flight sub-time slot The energy consumed by flight within the spacecraft; This indicates that the drone f is unloading the sub-time slot. Perform the search task. This indicates that the drone f is charging from the laser charging station; To unload sub-slots The total computing energy consumed by the local computing of the drone. For drones f in Energy collected during the period; if The drone f then enters an energy depletion state and terminates its current mission.
[0069] Beneficial effects:
[0070] This invention constructs a collaborative target search architecture for UAV swarms based on the motion model of UAVs in a continuous three-dimensional space. Under this architecture, a joint optimization problem is established, encompassing the UAV's three-dimensional motion trajectory, charging decisions, search task unloading location, and computational resource allocation, to minimize search area uncertainty and increase the target detection rate. Then, the MAPPO algorithm based on heuristic embedding is used to solve this joint optimization problem, yielding the planned actions for each UAV. In the algorithm design, a heuristic energy-saving task offloading method is used to efficiently determine the offloading location of search tasks and allocate computational resources for search tasks, thereby improving the efficiency of search task offloading. A Beta distribution is used to update the MAPPO policy network, avoiding the policy gradient bias caused by the forced truncation of actions exceeding boundary values in Gaussian distributions, thus improving the algorithm's convergence performance. A safety action controller is introduced to assist in UAV trajectory design, ensuring that the UAV always flies within a safe altitude above the search area, improving the algorithm's convergence speed. A policy network based on a single-level multi-head attention mechanism and an evaluation network based on a two-level multi-head attention mechanism are designed to help handle the UAV network with partial observability and dynamic size, improving the algorithm's convergence performance and generalization. A learning mechanism is used to reuse the strategies and experiences of small-scale UAV swarm collaborative target search in large-scale UAV swarm collaborative target search, reducing the difficulty of policy learning in large-scale UAV swarm cooperation scenarios, accelerating algorithm convergence, and showing significant advantages in reducing search area uncertainty and increasing the proportion of target discovery. Furthermore, it can greatly reduce the training difficulty of the algorithm in large-scale UAV swarm collaborative target search scenarios. Attached Figure Description
[0071] Figure 1 This is a flowchart illustrating a cooperative target search method for UAV swarms based on reinforcement learning, provided by the present invention.
[0072] Figure 2 This example illustrates a scenario of a drone swarm collaboratively searching for a target.
[0073] Figure 3 This is a motion model of the UAV within a flight sub-time slot in the embodiment.
[0074] Figure 4 This is the search model of the UAV during the unloading sub-slot in the embodiment.
[0075] Figure 5(a) shows the relationship between the observation area of the UAV and the actual search area processed in the embodiment.
[0076] Figure 5(b) shows an image segmentation method for the search area actually processed by the UAV in one embodiment.
[0077] Figure 6(a) shows a policy network based on the heuristic embedding MAPPO algorithm in one embodiment.
[0078] Figure 6(b) shows the evaluation network of the MAPPO algorithm based on heuristic embedding in one embodiment.
[0079] Figure 7 This is a flowchart illustrating a heuristic energy-saving task offloading method in one embodiment. Detailed Implementation
[0080] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0081] This invention provides a cooperative target search method for UAV swarms based on reinforcement learning, which specifically includes the following steps:
[0082] Step 1: Based on the motion model of UAVs in three-dimensional continuous space, construct a collaborative target search architecture for UAV swarms.
[0083] The UAV swarm collaborative target search architecture first discretizes the search area into several grids, with the search task for each grid represented by the amount of search data and the required CPU processing density. A base station with an edge server is located at the center of the search area to provide image recognition services for the target search task. A laser charging station is located at the boundary of the search area to charge the UAVs, extending their flight and search time. The entire time period of the UAV swarm collaborative target search task is divided into multiple equal-length sub-time slots, each containing a flight sub-time slot and an unloading sub-time slot. Each UAV in the swarm carries a certain amount of pre-loaded energy, which is consumed during task execution. Each UAV is equipped with an onboard top-down camera on its underside, capable of capturing images of the grids within the search area. During each flight sub-time slot, the UAV can move at a fixed speed in any direction within three-dimensional space. During each unloading sub-time slot, the UAV can use a remote laser charging station to recharge its battery. Alternatively, it can perform target search tasks. During the charging process using a remote laser charging station, the drone's own energy cannot exceed its maximum battery capacity. When performing target search tasks, the onboard camera can overlook and capture a specific search area directly below it to form a search observation image. The drone can use image segmentation technology to divide the captured search observation image into multiple observation sub-images of specific sizes. The image recognition task for each observation sub-image can be performed by the drone itself, or it can be unloaded to a base station at the center of the search area for image recognition. The drone consumes energy when performing image recognition or unloading images to the base station; the amount of energy consumed depends on the locally allocated CPU computing resources and the communication channel conditions between the drone and the base station. After completing the recognition task for all observation sub-images in the captured image during the unloading sub-slot, the drone dynamically updates the target existence probability of the search area grid based on the recognition results of each observation sub-image.
[0084] Step 2: Under the collaborative target search architecture of UAV swarm, establish a joint optimization problem of UAV 3D motion trajectory, charging decision, search task unloading location and computing resource allocation, so as to minimize the uncertainty of the search area and increase the target discovery rate.
[0085] Specifically, the goal of UAV swarm collaborative target search is to minimize the uncertainty of the search area and increase the target discovery rate. The UAV swarm collaborative target search task needs to meet the following constraints: maximum flight speed and azimuth limits for UAVs; collision avoidance between UAVs; UAVs not flying out of the search area boundary or beyond the safe flight altitude; UAV energy always being greater than 0 and less than the maximum battery capacity limit during the mission; computational tasks for each observation sub-image captured by the UAV need to be offloaded to the UAV's local or remote base station; the computational resources allocated to the UAV for offloading observation sub-images must not exceed the UAV's maximum CPU frequency or the UAV's maximum transmission power to the base station; UAVs can only perform charging or target search tasks during the offloading sub-slots; UAVs or base stations can use image recognition technology to process each offloaded observation sub-image in parallel; the decision variables for the UAV swarm collaborative target search task are the UAV's trajectory planning strategy, UAV charging strategy, offloading strategy for each observation sub-image, and the allocation of computational resources for each observation sub-image. This example comprehensively considers the above constraints, with the optimization objective of minimizing the uncertainty of the search area and increasing the target discovery rate, and establishes a joint optimization problem for UAV 3D motion trajectory, charging decision, search task offloading location, and computational resource allocation.
[0086] Step 3: Based on the heuristic embedding MAPPO algorithm, solve the joint optimization problem of UAV three-dimensional motion trajectory, charging decision, search task unloading location and computing resource allocation to obtain the planned actions of each UAV.
[0087] In summary, the heuristically embedded MAPPO algorithm improves search task unloading efficiency by efficiently deciding on the unloading location and allocating computational resources for search tasks through a heuristic energy-saving task offloading method. The heuristically embedded MAPPO algorithm uses a Beta distribution to update the MAPPO policy network, avoiding the policy gradient bias caused by the forced truncation of actions exceeding boundary values in Gaussian distributions, thus improving the algorithm's convergence performance. A safety action controller is introduced to assist in UAV trajectory design, ensuring that the UAV always flies within a safe altitude above the search area, improving the algorithm's convergence speed. A policy network based on a single-level multi-head attention mechanism and an evaluation network based on a two-level multi-head attention mechanism are designed to help handle the UAV network's partial observability and dynamic size, improving the algorithm's convergence performance and generalization. Through a learning mechanism, strategies and experiences from small-scale UAV swarm collaborative target search are reused for large-scale UAV swarm collaborative target search, reducing the difficulty of policy learning in large-scale UAV swarm cooperative scenarios and accelerating algorithm convergence.
[0088] Specifically, MAPPO is a deep reinforcement learning algorithm for multi-agent proximal policy optimization, which uses the classic Actor-Critic architecture to find the optimal policy for generating the best actions of the agents.
[0089] First, to significantly improve the efficiency of search task unloading, a heuristic-based energy-saving unloading method is designed to efficiently determine the unloading location and computational resource allocation for search tasks. Second, to prevent UAVs from flying out of the search area boundary or safe flight altitude, and to accelerate algorithm convergence, a safety action controller is designed to limit the UAV's flight trajectory. Third, to handle the local observability of UAVs and the dynamic UAV network with variable UAV numbers, a policy network based on a single-level multi-head attention mechanism and an evaluation network based on a two-level multi-head attention mechanism are proposed. Fourth, to improve algorithm convergence performance, a Beta distribution is introduced to update the MAPPO policy network. Finally, to accelerate the efficiency of large-scale UAV swarm collaborative target search, a learning mechanism is introduced to reuse strategies and experiences from small-scale UAV swarm collaborative target search scenarios.
[0090] Step 4: Each drone performs the corresponding planned action to complete the drone swarm collaborative target search task.
[0091] This invention constructs a collaborative target search architecture for UAV swarms based on the motion model of UAVs in a continuous three-dimensional space. Under this architecture, a joint optimization problem is established, encompassing the UAV's three-dimensional motion trajectory, charging decisions, search task unloading location, and computational resource allocation, to minimize search area uncertainty and increase the target detection rate. Then, the MAPPO algorithm based on heuristic embedding is used to solve this joint optimization problem, yielding the planned actions for each UAV. In the algorithm design, a heuristic energy-saving task offloading method is used to efficiently determine the offloading location of search tasks and allocate computational resources for search tasks, thereby improving the efficiency of search task offloading. A Beta distribution is used to update the MAPPO policy network, avoiding the policy gradient bias caused by the forced truncation of actions exceeding boundary values in Gaussian distributions, thus improving the algorithm's convergence performance. A safety action controller is introduced to assist in UAV trajectory design, ensuring that the UAV always flies within a safe altitude above the search area, improving the algorithm's convergence speed. A policy network based on a single-level multi-head attention mechanism and an evaluation network based on a two-level multi-head attention mechanism are designed to help handle the UAV network's partial observability and dynamic size, improving the algorithm's convergence performance and generalization. A learning mechanism is used to reuse the strategies and experiences of small-scale UAV swarm collaborative target search in large-scale UAV swarm collaborative target search, reducing the difficulty of policy learning in large-scale UAV swarm cooperative scenarios and accelerating algorithm convergence. This invention has significant advantages in reducing search area uncertainty and increasing the proportion of target discovery, and can greatly reduce the training difficulty of the algorithm in large-scale UAV swarm collaborative target search scenarios.
[0092] In one embodiment, step 1 includes: discretizing the search area into several grids, where the search task for each grid is represented by the amount of search data and the required CPU processing density; equipping the center of the search area with a base station equipped with an edge server to provide image recognition services for the target search task; equipping the boundary of the search area with a laser charging station to charge the drones, thereby extending the drones' flight and search time; and dividing the entire time period of the drone swarm collaborative target search task into multiple equal-length sub-time slots, each sub-time slot sequentially containing a flight sub-time slot. and unloading sub-slots Each drone in the drone swarm carries a certain amount of pre-loaded energy, which is consumed during mission execution. Each drone is equipped with an onboard top-view camera on its underside, capable of capturing images of the search area grid. F drones take off from their respective starting points. When the signal-to-noise ratio (SNR) of the communication channel between drones exceeds the drone communication SNR threshold, they can communicate via a predefined communication protocol. The communication information exchanged between drones includes their position and remaining energy. During each flight sub-slot, a drone can fly at a fixed speed in any direction within three-dimensional space. During each unloading sub-slot, the drone is stationary and can use a remote laser charging station to recharge its power or perform target search missions. When a drone is performing a target search mission, the onboard camera can overlook and capture a specific search area directly below it. This search area contains multiple grids to form a search observation image. The drone can use image segmentation technology to divide the captured search observation image into multiple observation sub-images of specific sizes. Image recognition for each observation sub-image can be performed by the drone itself. The image recognition can be performed by the base station at the center of the search area, which is then handled by the base station. The drone consumes energy when performing image recognition or unloading images to the base station, and the amount of energy consumed depends on the amount of CPU computing resources allocated to the drone and the communication channel conditions between the drone and the base station. After completing the image recognition task for all observation sub-images in the search observation image during the unloading sub-slot, the drone dynamically updates the target existence probability and target existence status of the search area grid based on the recognition results of each observation sub-image. Before the drone begins its target search task, the target existence probability in each grid is 0.5, and the uncertainty index is 1. During the target search task, the drone collects observation data using its onboard sensors and dynamically updates the target existence probability in the search area grid based on the collected observation data. The target existence probability detected by drone f at time t+1 is updated based on the detection probability and false alarm probability of drone f at different altitudes at time t, according to the Bayesian criterion. The environmental uncertainty of the grid is defined as the information entropy of the target existence probability within the corresponding grid. Figure 2 This embodiment demonstrates a scenario of drone swarm collaborative target search.
[0093] Specifically, the search region is discretized into several grids as follows:
[0094] Define the search region as having a size of L. x *L y In a two-dimensional plane, this region is divided into |K| grids, each with a side length of L, denoted as (x... k y k ), k∈K={1,2,…,|K|}; the search task for the k-th grid is Tk ={N k C k}, k∈K represents, where N k C represents the search data size for grid k. k This represents the data processing density required for the search grid k (in CPU cycles / bit); assuming N k Obligation expectation value The normal distribution, i.e. |N| targets are randomly distributed within a grid in the search area, denoted as n∈N={1,2,…,|N|}.
[0095] In the initial stage, the probability of the target existing in each grid is P(x). k y k The uncertainty index U(x) is 0.5 for both. k y k All values are 1. Furthermore, |F| UAVs take off from their respective starting points and can fly within a safe flight altitude space above the search area, denoted as f∈F={1,2,…,|F|}. The UAVs' mission is to explore the unknown search area, reduce uncertainty within the search area, and discover as many targets as possible. In particular, considering the specific detection probabilities and false alarm probabilities of the UAVs' onboard sensors, the UAVs need to visit the same search area multiple times to reduce information uncertainty and confirm the exact existence of targets.
[0096] The UAV motion model used in this invention, such as Figure 3 As shown in the figure, the motion model of the UAV within each flight sub-time slot t is illustrated. Within this area, each drone can move continuously in three-dimensional space above the search area. Assume the drone's flight speed is... If the position of the drone f remains constant, then the position update can be expressed as:
[0097]
[0098] in, and These represent the position coordinates of UAV f at the beginning of sub-time slot t and sub-time slot t+1, respectively. This indicates that the UAV f is in the flight sub-time slot. Fixed flight speed within, Indicates the duration of the flight sub-slot. This indicates that the UAV f is in the flight sub-time slot. The angle between the direction of the flight velocity and the z-axis in the three-dimensional Cartesian coordinate system. This indicates that the UAV f is in the flight sub-time slot. The angle between the direction of the flight velocity and the x-axis in the three-dimensional Cartesian coordinate system.
[0099] The UAV search model used in this invention, such as Figure 4 As shown in the figure, the search model of the UAV within the unloading sub-slot is illustrated. The unloading sub-slot is located within each sub-slot t. Within the UAV f, the relationship between its observation area and its location is as follows:
[0100]
[0101] Among them, FOV f W represents the field of view of the onboard camera of the drone f. f H represents the field of view width of an airborne camera. f Indicates the height of the airborne camera's field of view. This indicates that the drone f is unloading the sub-time slot. The altitude of the drone f. The observation area of the drone f. Represented as:
[0102]
[0103] in, The aspect ratio of the field of view of the onboard camera of the drone.
[0104] The observation area of UAV f Search area with actual processing The relationship is shown in Figure 5(a). In each unloading sub-slot... In the middle, the observation area of UAV f This may partially cover certain grid cells. In this invention, it is assumed that the UAV can only process its observation area. Full coverage of the grid cells means that some grid cells located at the boundary of the observation area cannot be unloaded from the sub-slot. Since the area within the search area is not detected, this invention reduces the impact of this phenomenon by improving the discretization accuracy of the grid cells. Simultaneously, the continuous motion space design of the UAV and the reinforcement learning-based search scheme can further inspire the UAV to learn a reasonable flight path, maximizing coverage of the entire grid cell to improve search efficiency. Furthermore, when the UAV's observation area exceeds the boundary of the search area, the UAV only searches the area overlapping with the search area. Let... This indicates the observation area of the UAV f. Let f represent the actual search area processed by the drone.
[0105] This invention constructs a search task offloading model, specifically, it uses image segmentation to divide the search area actually processed by the UAV. The image segmentation method of the search area actually processed by the UAV is shown in Figure 5(b). Different observation sub-images are formed after image segmentation. The data can be offloaded to the UAV or a base station in parallel, thereby improving the efficiency of image recognition computation. For ease of task processing, this invention sets the side length of the observed sub-image to an integer multiple of the grid length, for example, twice in Figure 5(b). The relationship between the observed sub-image and the grid can be represented as follows:
[0106]
[0107] Where ok represents a sub-observation sub-image, k represents a grid, and k∈ok represents the set of all grids in the observation sub-image ok. N ok This represents the amount of detection data for the observed sub-image "ok", expressed in bits. (C) ok This indicates the processing density required to observe the sub-image ok, in CPU cycles per bit.
[0108] In this invention, This indicates that the observed sub-image ok is unloaded in sub-time slot t. During this period, local image processing was performed on the drone f. This indicates that the observed sub-image 'ok' undergoes image processing at a remote base station. When the observed sub-image 'ok' performs image recognition processing locally, it satisfies the following:
[0109]
[0110] in, This indicates that the drone f is unloading the sub-time slot. The local CPU computing frequency allocated to the observed sub-image ok during the period; This represents the computation time consumed by the UAV f in locally recognizing the observed sub-image ok. The sum of the CPU computing frequencies allocated by the UAV f for all observed sub-images ok should not exceed the maximum CPU computing frequency of the UAV f. Right now:
[0111]
[0112] To represent data transmission between the UAV and the base station, this invention employs a ground-to-air path loss communication model. Specifically, this invention considers both line-of-sight (LAS) and non-LAS transmission channels. Let the probabilities of using LAS and non-LAS transmission between the UAV f and the base station be respectively... and Modeling can be done using the following formula:
[0113]
[0114] Here, ω and η are both constants, reflecting the environmental indicators of the search area. This represents the horizontal distance between the drone f and the base station. This represents the flight altitude of drone f. The channel gain between drone f and the base station. for:
[0115]
[0116] Where, d f,bs Let represent the absolute distance between the drone f and the base station. c∈{LoS, NLoS} indicates whether the communication channel is a line-of-sight (LOS) or non-line-of-sight (NLOS) channel. β c α represents the channel gain at a distance of 1 meter. c η is the path loss constant. c Using Gaussian distribution Modeling the hybrid channel gain, it can be expressed as:
[0117]
[0118] The signal-to-noise ratio of communication between the drone f and the base station is:
[0119]
[0120] in, This indicates the data transmission power of the drone f transmitting the observed sub-image ok to the base station. Let be the noise power. Therefore, the channel transmission rate from the UAV to the base station can be expressed as:
[0121]
[0122] Where B represents the channel bandwidth, and the transmission time of the observed sub-image ok from the drone f to the base station is:
[0123]
[0124] also, It should not exceed the maximum transmit power of the drone.
[0125]
[0126] Considering the strong processing power of the base station and the small size of the image recognition result, this invention ignores the time required for base station image recognition and the return time of the image recognition result from the base station to the drone. The total processing time for the observed sub-image OK is... It can be represented as:
[0127]
[0128] In this invention, image recognition for all observed sub-images (OK) needs to be performed in the unloading sub-time slot. Completed within [time period]. (Exceeding [time limit]) The image recognition task will be judged as a failure. Indicates that the observed sub-image is OK. Complete within the time limit, otherwise Then we have:
[0129]
[0130] set up This indicates that drones or base stations can... Successful analysis and computation of the internal grid (x k y k The detection data of ) . The observation sub-image ok contains multiple grids k. The completion of the calculation of the observation sub-image ok means that the calculation of all k within it is completed, therefore:
[0131]
[0132] In this invention, the drones exchange only basic control information, including their three-dimensional coordinates. and remaining energy This invention uses a line-of-sight channel to model the air-to-air communication channel gain between UAVs:
[0133]
[0134] Where, d f,f′ This represents the distance between drone f and another drone f′. The signal-to-noise ratio (SNR) of communication between drone f and drone f′ is:
[0135]
[0136] Among them, PT f,f′ This represents the transmission power between the two drones. Let the communication neighborhood of drone f be...
[0137] OBS f ={f′∈F|f′≠f,SNR f,f′ (t)≥SNR thr}, SNR thr OBS represents the minimum signal-to-noise ratio threshold required for drones to transmit communication information. f This indicates that the drone f can only communicate with its surroundings, and the signal-to-noise ratio is greater than the SNR. thr The drone f′ communicates with the drone (i.e., transmits three-dimensional coordinates and remaining energy).
[0138] The confidence probability graphical model in this invention uses the target probability distribution graph P. t (x k y k) to describe each grid (x) in the entire search region at the start of sub-slot t. k y k The probability that the target exists in P. Specifically, P t (x k y k )∈[0,1] indicates that the target exists in the grid (x) k x k The probability in the equation. At the start of the search task, the probability of the target existing is set to P. 1 (x k y k =0.5 indicates that the drone swarm has no prior information about the search area.
[0139] When a drone swarm performs a search mission, the drones use onboard cameras to collect detection data for each grid. The collected grid detection data needs to be analyzed and calculated on the drones or a remote base station, and then the target probability distribution P is updated based on the calculation results. t (x k y k This invention uses... This indicates whether the drone's onboard sensors detected the target within time slot t. This indicates that the drone has detected a target. This indicates that the drone did not detect the target. Due to the limited detection accuracy of airborne sensors, this invention uses a Bayesian model to update the target presence probability in the grid:
[0140]
[0141] Among them, P t (x k y k ) represents the grid (x) at the start of sub-slot t. k y k The probability that the target exists, P s This represents the confidence level when the sensor detects a target in the grid. This is related to the drone's flight altitude. If the sensors do not detect the target, then... Replace P s If more than one drone simultaneously detects the grid (x) k y k If P, then t (x k y k The same number of updates will be performed.
[0142] This invention uses a linear model to describe the confidence level of an airborne camera. and drone flight altitude Relationship between: Confidence level of airborne cameras Flight altitude of drones Inversely proportional. The lower the drone's flight altitude (the closer to the grid points), the higher the detection accuracy of the onboard camera, which can be specifically expressed as:
[0143]
[0144] in, This indicates that the drone f is at an altitude of [height missing]. The confidence level of the onboard camera. and These represent the maximum and minimum safe flight altitudes of the drone f, respectively. and These represent the confidence levels of the onboard camera at the maximum and minimum safe flight altitudes of the drone, respectively.
[0145] The goal of drone swarm search missions is to reduce the information uncertainty in the search area, which can be achieved by considering the target probability distribution P. t (x k y k This is reflected in the information entropy. The definition of information entropy is as follows:
[0146] E t (x k y k )=-P t (x k y k log2P t (x k y k )-(1-P t (y k y k log2(1-P) t (x k y k ))
[0147] Grid (x) k y k A high information entropy value indicates high uncertainty in the grid information. When P t (x k y k When ) = 0.5, E t (x k y k ) = 1; when P t (x k y k When )∈{0,1}, E t (x k y k ) = 0.
[0148] The energy model used in this invention, namely the UAV f in the flight sub-time slot The energy consumed by flight within the spacecraft can be expressed as: in, This indicates the propulsion power of the drone f. It can be represented as:
[0149]
[0150] Among them, P a and P b Representing the blade profile and induced power, u tip v represents the tip speed of the rotor blades when the UAV is hovering. a Indicates the induced rotor speed during hovering, κ, τ a f a A and A represent air density, rotor solidity, fuselage drag ratio, and rotor disk area, respectively.
[0151] The CPU computing power consumed by the UAV f in processing the observed sub-image ok locally can be expressed as:
[0152]
[0153] Among them, κ f This is a coefficient related to the drone's power f. During the unloading sub-slot... During this period, the total computing energy consumed by the UAV f in local computing is:
[0154]
[0155] The energy consumed by the drone f to transmit the observed sub-image ok to the remote base station can be expressed as:
[0156]
[0157] In target search missions involving drone swarms, considering the high energy consumption of drone flight, this invention employs laser charging to recharge the drones, thereby improving mission efficiency. During the unloading sub-slot... In this context, the drone either performs search missions or recharges its power from laser charging stations. Let... Indicates that the drone f is in Perform the search task. This indicates that drone f is charging from a laser charging station. Drone f is... The energy collected during the period is:
[0158]
[0159] Where ∈{0,1} is the energy conversion efficiency; This represents the laser charging channel; PL1 is the laser transmission power; G is the area of the laser collector; φ represents the optical efficiency of the combined transmitter and receiver; φ represents the channel medium attenuation coefficient. ξ represents the distance between the drone f and the laser charging station; F represents the initial laser beam size; ξ represents angular diffraction.
[0160] In summary, the energy update model for UAV f within sub-time slot t can be expressed as:
[0161]
[0162] in, and These represent the energy of the drone f at the beginning of sub-time slot t and sub-time slot t+1, respectively. This represents the battery capacity limit of the drone f. The min(·) function ensures that the drone's energy will not exceed the battery capacity limit while charging. Then the drone f will enter a state of energy depletion and terminate its current mission.
[0163] In one embodiment, step 2 includes: establishing a joint optimization problem for the UAV's 3D motion trajectory, charging decision, search task unloading location, and computational resource allocation under the UAV swarm collaborative target search architecture, in order to minimize the uncertainty of the search area and increase the target detection rate. The joint optimization problem is:
[0164]
[0165] st
[0166]
[0167]
[0168] The objective function of this invention is to jointly optimize the flight trajectory of the UAV. Drone charging strategy Data offloading strategy for observed sub-image OK Drone CPU computing frequency and transmission power The allocation is designed to minimize the uncertainty of the search area and increase the proportion of targets found within the search area. This indicates that the UAV f is in the flight sub-time slot. The magnitude of the fixed flight speed within; This indicates that the UAV f is in the flight sub-time slot. The angle between the direction of the flight velocity within the space and the z-axis in the three-dimensional Cartesian coordinate system; This indicates that the UAV f is in the flight sub-time slot. The angle between the direction of the flight velocity within the space and the x-axis in the three-dimensional Cartesian coordinate system; This indicates the drone's charging strategy; The unloading strategy indicates whether the observed sub-image is OK; This indicates the CPU frequency allocation for observing sub-images. The transmission power allocation for the observed sub-image 'ok' is represented by: T; T represents the set of all sub-slots in the entire search task; K represents the set of all grids within the search area; P t (x k y k ) indicates that at the beginning of sub-slot t, the grid (x) k y k The probability that the target exists in (u); t (x k x k )∈{0,1} represents the grid (x) k y k Whether or not a target actually exists in u t (x k y k ) = 1 indicates that the grid (x) k y k There is a target in ) u t (x k y k ) = 0 indicates that the grid (x) k y k There is no target in ); ue t (x k y k )∈{0,1} represents the grid (x) at the start of sub-slot t. k y k Has the target been found in )? t (x k y k ) = 1 indicates that the grid (x) starts at the beginning of sub-slot t. k y k The target has been found in ) ue t (x k y k ) = 0 indicates that the grid (x) starts at the beginning of sub-slot t. k y k No target has been found in E; t (x k y k ) represents the grid (x) k y k Information uncertainty; w1 and w2 represent weighting factors used to balance target detection and environment search; This represents the maximum flight speed of the drone f; This represents the maximum CPU computing frequency of the drone f; This indicates the maximum transmit power of the UAV f; Indicates the maximum battery capacity limit of the drone f; L s Indicates the safe distance between two drones; L x With L y Indicates the length and width of the search area; and This represents the minimum and maximum safe flight altitudes for drones; F represents the set of all drones. P represents the actual search area processed by the UAV f. t+1 (x k ,y k (ue) t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k )) indicates that the drone swarm discovered the grid (x) within sub-time slot t. k y k The increase in the probability of the target's existence within the range; ∑ k∈K (P t+1 (x k ,y k (ue) t+1 (x k ,y k )-u t+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k E represents the increase in the probability that the UAV swarm detects the presence of a target within the search area during sub-time slot t; t (x k y k )-E (t +1)(x k y k ) represents the grid (x) k y k The reduction in information uncertainty within sub-time slot t; ∑ k∈K (E t (x k y k )-E (t+1) (x k y k )) represents the reduction in information uncertainty of the search region within sub-time slot t; C1, C2, C3, C4, C5, C6, C7, and C8 represent eight constraints.
[0169] Specifically, C1 restricts the UAV's flight speed and direction in the three-dimensional coordinate system, ensuring they do not exceed the maximum flight speed and direction. C2 relates to UAV charging scheduling, meaning the UAV either performs a charging operation or a search task in each time slot. C3 restricts the location where UAV detection data is unloaded, ensuring that each observation sub-image OK is either calculated locally by the UAV or unloaded to a remote base station. C4 and C5 restrict the CPU computing frequency and transmission power allocated to the UAV for unloading observation sub-image OK, ensuring they do not exceed the UAV's maximum resource limits. C6 ensures that the UAV's energy level remains positive and does not exceed the maximum battery capacity limit during the search task. C7 ensures that the distance between any two UAVs is greater than the safe distance L. s To avoid drone collisions. C8 indicates that drones cannot fly outside the search area boundaries or outside the safe flight altitude range.
[0170] In one embodiment, step 3 includes: modeling the UAV swarm collaborative target search task as a distributed partially observable Markov decision process; which includes: an environmental state space, an UAV observation space, a UAV action space, and a reward function; constructing a MAPPO algorithm model based on heuristic embedding; wherein each agent corresponds to one UAV, and each agent includes a policy network and an evaluation network; the policy network includes an embedding layer, a single-level multi-head attention layer, and an output layer; the evaluation network includes an embedding layer, a two-level multi-head attention layer, and an output layer; wherein the training process of a single agent includes: at the beginning of each sub-time slot t, inputting the UAV's observation state into the policy network of the corresponding agent, and through the embedding layer and single-level multi-head attention mechanism of the policy network, the UAV's observation state is input into the policy network of the corresponding agent. The attention layer calculates intermediate results and then inputs them into the output layer to obtain the parameters β1 and β2 of the probability density function of the Beta distribution. By sampling the Beta distribution, the movement actions and charging strategy of the UAV within sub-time slot t are obtained. Based on the UAV's current position and the movement actions output by the policy network, the final movement strategy of the UAV is obtained based on the safety action controller. According to the UAV's location and the communication channel conditions between the UAV and the base station, a heuristic energy-saving task offloading method is used to determine the offloading strategy and computational resource allocation for the observed images captured by the UAV. The global environment state is input into the evaluation network based on a two-level multi-head attention mechanism for the corresponding agent to obtain a score for the policy network. Based on the policy network's score and the reward function, the advantage function is calculated as follows:
[0171]
[0172] in, R represents the dominance function. f (t) represents the reward function, V f (s(t)) represents the score of the policy network, s(t) represents the global environment state, γ represents the discount factor, and λ represents the generalized advantage estimation parameter.
[0173] Based on the advantage function and the ratio of the current policy to the original policy, the policy network is updated according to the objective function; policy network training continues until a preset termination condition is met, and the planned UAV actions are output. The objective function is:
[0174]
[0175] in, θ is the average reward at time t; f These are the trainable parameters of the policy network; ∈ is the cutoff coefficient; and φ represents the current policy and the original policy, respectively. f (θf ,t) is the ratio between the current policy and the original policy.
[0176] In one embodiment, the environment state space is:
[0177]
[0178] in, This represents the remaining energy of the drone f at the beginning of sub-time slot t; L represents the three-dimensional position coordinates of the UAV f at the beginning of sub-time slot t; BS and L laser P represents the coordinates of the base station and the laser charging station, respectively; t (x k y k ), k∈K represents the grid (x) at the beginning of sub-slot t. k y k The probability of the target existing; RT represents the remaining time steps in a set (round) of reinforcement learning.
[0179] The observation space of the drone is:
[0180]
[0181] in, Let OBS represent the probability distribution of targets that UAV f can observe. Since UAV f has a limited field of view, it can only obtain the target probability distribution within the grid of its field of view; the target probability of grids outside the field of view is represented by 0. f Let SNR represent the communication range of the drone f. thr The minimum signal-to-noise ratio (SNR) threshold for communication between drones; f,f′ This represents the signal-to-noise ratio (SNR) of communication between drone f and drone f′. When SNR... f,f′ ≥SNR thr When f′∈OBS f ;otherwise,
[0182] The drone's action space is:
[0183]
[0184] in, and It is a continuous variable that determines the flight sub-time slot of the UAV f. The flight path within. It is a discrete variable that determines the unloading sub-slot of the drone. The charging strategy within. To facilitate the design of the MAPPO strategy network, this invention makes the discrete charging strategy continuous: that is, it has when season This indicates that the drone f is charging from the laser charging station within sub-time slot t; otherwise... This indicates that the UAV f performs search area exploration and data offloading. It is particularly important to note that the UAV f's action space does not include decision variables related to data offloading and computational resource allocation. This invention uses a heuristic, energy-efficient task offloading method to determine the UAV's data offloading strategy and computational resource allocation.
[0185] The reward function is:
[0186]
[0187] Where CR is the penalty value for drone f colliding with other drones, and it is a large positive number. w1 and w2 are weighting factors that balance target detection and environment search. K represents the set of all grids in the search area; u t (x k y k )∈{0,1} represents the grid (x) k y k Whether or not a target actually exists in ) t (x k y k )∈{0,1} represents the grid (x) at the start of sub-slot t. k x k Has the target been found in (P)? t+1 (x k ,y k (ue) t+1 (x k ,y k )-u t + 1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k )) indicates that the drone swarm discovered the grid (x) within sub-time slot t. k y k The increase in the probability of the target's existence within the range; ∑ k∈K (P t+1 (x k ,y k (ue) t+1 (x k ,y k )-ut+1 (x k ,y k ))-P t (x k ,y k (ue) t (x k ,y k )-u t (x k ,y k E represents the increase in the probability that the UAV swarm detects the presence of a target within the search area during sub-time slot t; t (x k ,y k )-E (t+1) (x k ,y k ) represents the grid (x) k y k The reduction in information uncertainty within sub-time slot t; ∑ k∈K (E t (x k ,y k )-E (t+1) (x k ,y k The )) represents the reduction in information uncertainty of the search area within sub-time slot t. It is particularly important to note that the reward function does not consider the case where the UAV f flies out of the search area boundary or beyond the safe flight altitude. To improve the convergence speed of the MAPPO algorithm, this application uses a safe action controller to handle the case where the UAV flies out of the search area boundary or beyond the safe flight altitude.
[0188] In a drone swarm collaborative target search scenario, each drone can only perform local observations (i.e., communicate with a varying number of neighboring drones), and the total number of drones in the search scenario may also change. Traditional multi-agent reinforcement learning algorithms handle the dynamic changes in the number of drones by pre-reserving input dimensions, resulting in poor scalability; furthermore, as the number of drones increases or decreases, both the policy network and value network of MAPPO need to be retrained. To address these issues, this invention employs a two-level multi-head attention mechanism to handle changes in the number of drones and improves the algorithm's ability to understand and capture complex dependencies between agents.
[0189] Specifically, the policy network based on the heuristic embedding MAPPO algorithm proposed in this invention has the structure shown in Figure 6(a). Here, the local observations of the UAV f are divided into its own features. (including environmental characteristics) and its communication range OBS f Features of other UAVs j MLP stands for Fully Connected Network; this invention uses an MLP embedding layer to respectively start from o f and o j Extracting two types of features e f and e j And send it to the multi-head attention head to obtain the attention value x. f ;x f Includes relevant information about the drone itself, while ignoring irrelevant information. f It can be represented as:
[0190]
[0191] Among them, e f E represents the eigenvalues of the drone f; j (j∈OBS f ) indicates the communication range of the drone f (OBS). f The eigenvalues of other UAVs j; W k W q and W v These are the weight matrices for keys, queries, and values, respectively. It is W k The transpose of d; k It is the key dimension; α f,j It is the attention score; x f It is the attention head output; finally, x f With e f The data are concatenated and sent to the output layer MLP to obtain the motion probability distribution of the drone f.
[0192] The evaluation network based on the heuristic embedding MAPPO algorithm proposed in this invention has the structure shown in Figure 6(b). The evaluation process of the evaluation network is as follows:
[0193] (1) Global state information is divided into environmental information S m ={L BS L laser ,RT,P t (y k y k Information on )} and drone f MLP embedding layer from S m With S f (f∈F) Extract two types of features e respectively m and e f (f∈F), e m Represents environmental information S m eigenvalues, e f (f∈F) represents the information S of the drone f. f eigenvalues.
[0194] (2)ef (f∈F) is sent to |F| first-level multi-head attention heads to extract attention values x respectively. f (f∈F), x f (f∈F) extracts relevant information between drone f and other drones, where |F| represents the total number of drones.
[0195] (3)e f and x f The vectors are concatenated and sent to the MLP to obtain the attention output vector g. f (f∈F).
[0196] (4)e m and g f (f∈F) is sent to a two-level multi-head attention head to obtain the attention value x. m x m Relevant information between environmental conditions and all drones was extracted.
[0197] (5)e m and x m The values are concatenated and sent to the output MLP layer to obtain the final state value evaluation score. Notably, this application uses a single-level multi-head attention head to extract the correlation between all drone agents (homogeneous agents) and a two-level multi-head attention head to extract the correlation between the environment and drone agents (also known as heterogeneous agents).
[0198] Through the above network design, both the policy network and the evaluation network are decoupled from the number of agents, so they can both cope with changes in the number of dynamic agents.
[0199] The output of each UAV's policy network is a continuous four-dimensional distribution. As the number of UAVs increases, the cooperative patterns of the UAV swarm in the search environment and the entire policy space grow exponentially, posing a significant challenge to policy optimization. To facilitate cooperative search learning in large-scale UAV swarms, this invention utilizes a curriculum learning mechanism. Curriculum learning advocates starting with the simplest scenarios and gradually increasing the training difficulty to improve final asymptotic performance or reduce training time. Therefore, this invention proposes that both the MAPPO policy network and evaluation network support a dynamically increasing number of agents. First, the policy and evaluation networks are trained for small-scale UAV swarm cooperative target search scenarios. Then, at the start of training for large-scale UAV swarm cooperative target search scenarios, a model reloading mechanism is used to load the trained policy and evaluation networks from the small-scale UAV swarm cooperative target search scenarios onto each UAV, and training of the policy and evaluation networks continues. In this way, the strategies and experiences of small-scale UAV swarm cooperative target search can be reused in large-scale UAV swarms, thereby reducing the difficulty of policy learning in large-scale UAV swarm cooperative scenarios and accelerating algorithm convergence.
[0200] In the output layer of the UAV policy network, this invention uses a Beta distribution to update the MAPPO policy network, avoiding the policy gradient bias caused by the forced truncation of actions exceeding boundary values by the Gaussian distribution, thus improving the algorithm's convergence performance. Specifically, existing methods for handling continuous action spaces typically select actions by sampling from an unbounded Gaussian distribution. Since the Gaussian distribution requires forced truncation of values exceeding action boundaries, this leads to biased policy gradients. This invention uses a bounded Beta distribution for action sampling, replacing the existing method of sampling continuous actions from a Gaussian distribution, thereby avoiding boundary effects. The Beta distribution can be represented as:
[0201]
[0202] Here, β1 and β2 are parameters of the Beta distribution, output by the policy network output layer MLP; Γ(x) is the gamma function. Since the Beta distribution samples actions within the range [0, 1], it is necessary to map the sampled actions back to the original action range, for example,
[0203] In collaborative target search by drone swarms, the final movement strategy of the drone is obtained based on the current position of the drone and the movement actions output by the policy network, using a safety action controller.
[0204] Specifically, according to constraint C8, each UAV cannot fly outside the search area boundary or within the safe altitude range above the search area. While violating this constraint can typically be penalized by adding a negative reward value to the reward function, using constraint penalties does not help UAVs quickly learn safe flight maneuvers; instead, it slows down the convergence speed of the training algorithm. Therefore, this invention designs a safe action controller to assist in UAV trajectory design. Specifically, this invention uses the flight maneuver of UAV f within sub-time slot t, output by the policy network, and the current position of UAV f. Predict the position of drone f in the next sub-time slot t+1. if If a drone is not within the safe altitude range above the search area, it will abandon the drone flight actions output by the policy network and remain stationary for the current sub-time slot t. This "early prediction" mechanism ensures that each drone always flies within the safe altitude range above the search area, accelerating the convergence speed of the MAPPO algorithm.
[0205] In one embodiment, the flow of the heuristic energy-saving task offloading method is as follows: Figure 7 As shown, it specifically includes:
[0206] S1. Variables related to the energy-saving task unloading process of the drone initialization heuristic.
[0207] Drone f initialization and tFre = 0, where, The empty set is represented by NO, which stores observation sub-images that cannot be unloaded, BO, which stores observation sub-images that can only be unloaded to the base station, and FO, which stores observation sub-images that can only be unloaded to the UAV. FBO stores observation sub-images that can be unloaded to either the UAV or the base station, but are preferentially unloaded to the UAV because unloading to the UAV consumes less UAV energy. tFre records the sum of CPU frequencies allocated to all observation sub-images that are unloaded to the UAV.
[0208] S2, Minimum computational resources required for unloading the UAV's computational observation sub-image.
[0209] In each sub-slot t, the UAV f traverses the search area it actually processes. The observed sub-images are divided into several parts; for each observed sub-image, the minimum CPU computing frequency required for local processing by the UAV is calculated. Minimum transmission power required to transmit to the base station
[0210] The calculation formula is: Where, N ok C represents the amount of detection data for the observed sub-image "ok", in bits. ok The processing density required to observe sub-image OK is expressed in CPU cycles per bit. This indicates the length of the unloading sub-slot within sub-slot t.
[0211] The calculation formula is:
[0212]
[0213] Where B represents the channel bandwidth; Indicates noise power; CG f,bs (t) represents the hybrid channel gain between the UAV f and the base station. The UAV f simultaneously compares... Maximum computational frequency of UAV f as well as With the maximum transmit power of the UAV f Size; if and Then execute S2.1; if and Then execute S2.2; if and Then execute S2.3; if and Then execute S2.4;
[0214] S2.1. Processing observation sub-images that cannot be unloaded by the UAV.
[0215] This indicates that unloading the observation sub-image ok from UAV f would exceed the maximum computing resource allocation of UAV f. Therefore, UAV f cannot unload the observation sub-image ok. UAV f updates NO = NO∪ok and executes S2.5.
[0216] S2.2 The UAV can only process the observation sub-images that are offloaded to the base station.
[0217] This indicates that unloading the observed sub-image "ok" to the drone's local storage would exceed the drone's maximum CPU computing frequency. Due to limitations, drone f can only offload ok to the base station, drone f updates BO = BO∪ok, and executes S2.5.
[0218] S2.3 The UAV can only process observation sub-images that are offloaded locally.
[0219] This indicates that if the drone f were to offload the observed sub-image OK to the base station, it would exceed the maximum transmission power of the drone f. Due to limitations, drone f can only unload ok locally, update FO and tFre, and execute S2.5. The update methods for FO and tFre are as follows:
[0220]
[0221] Among them, κ f This represents a coefficient related to the power f of the UAV; tFre represents the sum of CPU frequencies allocated to all observation sub-images ok that the drone f updates and unloads locally.
[0222] S2.4 The drone executes an energy-saving unloading strategy.
[0223] This indicates that the UAV f can either unload the observed sub-image ok to its local machine or to the base station; in this case, the UAV f is calculated according to the following formula.
[0224]
[0225] in, This indicates that the drone f uses the minimum CPU computing frequency. Power when processing observed sub-images is OK; UAV f comparison and like This indicates that unloading the observed sub-image ok to the local machine consumes less drone energy than unloading it to the base station. Therefore, drone f unloads ok to the local machine, updates FBO and tFre, and executes S2.5; the updates of FBO and tFre are as follows:
[0226]
[0227] like This means that unloading the observed sub-image ok to the base station consumes less drone energy than unloading it locally. Therefore, drone f unloads ok to the base station, updates BO = BO∪ok, and executes S2.5.
[0228] S2.5 The UAV completes the initial unloading of all observed sub-images.
[0229] The drone f traversed the search area it actually processed. All observed sub-images are OK, and S3 is executed.
[0230] S3. The sum of the CPU computing frequencies required for the UAV to unload all observed sub-images is greater than the maximum CPU computing frequency.
[0231] The sum of the CPU frequencies allocated to all observation sub-images (ok) unloaded from the drone and tFre is compared with... like This indicates that the sum of the CPU computation frequencies required for all observed sub-images (ok) unloaded locally by the drone f, tFre, is greater than the maximum CPU computation frequency of the drone f. Then execute S4; if Then execute S5.
[0232] S4. The drone sorting can be offloaded to the drone's local location or to the observation sub-image of the base station.
[0233] The drone will carry out various tasks in the FBO. Key-value pairs based on value Sort in descending order, let the sorted set be . set up The number of key-value pairs is M, and S4.1 is executed.
[0234] S4.1, UAV sorting can only be unloaded to the observation sub-images local to the UAV.
[0235] The drone will carry out various tasks in FO. Key-value pairs based on value Sort in descending order, let the sorted set be . And execute S4.2.
[0236] S4.2. The merged sorting result of the UAVs is
[0237] The UAV f sets an ordered set where the elements of are arranged before and execute S4.3.
[0238] S4.3. The UAV initializes and traverses the relevant variables.
[0239] Define times to record the number of key-value pairs that have been traversed in, let times = 0, and execute S4.4.
[0240] S4.4. The UAV traverses each observed sub-image in and adjusts its offloading strategy.
[0241] The UAV f traverses in order each key-value pair {ok, priority} in, where priority represents the value of the key-value pair; for each the UAV f updates and compares the sizes of times and M. If times < M, it means that the observed sub-image ok needs to adjust its offloading position to the base station, then the UAV f executes BO = BO ∪ ok and deletes the key-value pairs corresponding to the observed sub-image ok in FBO and FO; if times ≥ M, it means that the offloading strategy of the observed sub-image ok needs to be adjusted to not offload, then the UAV f executes NO = NO ∪ ok and deletes the key-value pairs corresponding to the observed sub-image ok in FBO and FO; afterwards, the UAV f updates times = times + 1 and executes S4.5.
[0242] S4.5. The UAV jumps out of the traversal.
[0243] The UAV f compares tFre with the size of. If then it jumps out of the traversal of and executes S5; if then it continues to traverse until or the traversal is completed, and then execute S5.
[0244] S5. The UAV generates the final offloading strategy for all observed sub-images.
[0245] The UAV f generates the search area it actually processes Unloading strategies (BO, FO, FBO, NO) and computational resource allocation strategies for all observed sub-images within the region.
[0246] It can be seen that in S1, the heuristic energy-saving task offloading method first initializes the relevant variables used by the algorithm. In S2, S2.1, S2.2, S2.3, S2.4, and S2.5, the UAV f is calculated based on the maximum CPU computing frequency. and maximum transmission power Calculate the actual search area it processes. The optimal unloading location for each observed sub-image (ok) is determined. Here, S2.1 indicates that the drone (f) cannot unload ok; S2.2 indicates that the drone (f) can only unload ok to the base station; S2.3 indicates that the drone (f) can only unload ok to its local location; and S2.4 indicates that the drone (f) uses an energy-saving method to unload ok. In S3, S4, S4.1, S4.2, S4.3, S4.4, and S4.5, the sum of the CPU computation frequencies required by the heuristic energy-saving task unloading method to process the local image unloading of the drone (f) exceeds the maximum CPU computation frequency of the drone. In this situation, the adjustment mechanism for the OK unloading position of the observed sub-image follows these rules: First, for A heuristic approach to offloading energy-saving tasks redirects them to the base station for offloading. This represents the set of observed sub-images that can be unloaded from either the drone itself or the base station; secondly, for A heuristic approach to energy-saving task unloading adjusts it to not be unloaded, whereby... This represents the set of observed sub-images that can only be unloaded from the drone itself; third, for Heuristic energy-saving task offloading methods prioritize adjusting the required CPU computing frequency. Large, and the energy increase is small when offloaded to the base station. Fourth, for the unloading location of the observed sub-image OK; Heuristic energy-saving task offloading methods prioritize offloading tasks to the drone's local storage, which consumes more energy. The unloading location of the observed sub-image is OK; Fifth, when the sum of the CPU computing frequencies required for unloading the local image of the UAV f exceeds the maximum CPU computing frequency of the UAV. At that time, the drone f priority adjustment The observed sub-image is unloaded at the OK location, and then the position is adjusted. The sum of CPU computation frequencies required for unloading the observed sub-image at the ok unloading position until the adjusted local image is unloaded does not exceed the maximum CPU computation frequency of the UAV.
[0247] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for cooperative target searching of a UAV swarm based on reinforcement learning, characterized in that, Includes the following steps: Based on the motion model of UAVs in three-dimensional continuous space, a collaborative target search architecture for UAV swarms is constructed. Under the collaborative target search architecture of UAV swarms, a joint optimization problem is established for UAV 3D motion trajectory, charging decision, search task unloading location and computing resource allocation. The MAPPO algorithm based on heuristic embedding solves the joint optimization problem of UAV 3D motion trajectory, charging decision, search task unloading location and computing resource allocation, and obtains the planned actions of each UAV. The drones execute predetermined planned actions to complete the drone swarm collaborative target search task; The MAPPO algorithm for heuristic embedding includes: A distributed partially observable Markov decision process is used to construct a heuristically embedded MAPPO algorithm model for the collaborative target search task of UAV swarms. The model includes: an environment state space, an observation space of the UAVs, an action space of the UAVs, and a reward function. Each agent corresponds to one UAV, and the agent includes a policy network and an evaluation network. The policy network includes an embedding layer, an attention layer, and an output layer, and the evaluation network includes an embedding layer, two levels of attention layers, and an output layer. The training process of the agent is as follows: at the beginning of sub-time slot t, the observation state of the UAV is input into the agent's policy network. Intermediate results are calculated through the embedding layer and single-level attention layer of the policy network based on a single-level multi-head attention mechanism. The intermediate results are then input into the output layer to obtain the parameters of the Beta distribution probability density function. and The movement actions and charging strategies of the UAV within sub-time slot t are obtained by sampling the Beta distribution; the safety action controller obtains the movement strategy of the UAV based on the current position of the UAV and the movement actions output by the policy network; based on the position of the UAV and the communication channel conditions between the UAV and the base station, the unloading strategy of the observed image and the allocation of computing resources are determined by adopting a heuristic energy-saving task unloading method. The global environment state is input into the evaluation network of the agent based on a two-level multi-head attention mechanism to obtain the score of the policy network; the advantage function is calculated based on the score of the policy network and the reward function; the policy network is updated using the objective function based on the advantage function and the ratio of the current policy to the original policy; the policy network is then trained until the preset termination condition is reached, and the planned actions of the UAV are output. The heuristic-based energy-saving task offloading method is as follows: S1. Variables related to the heuristic energy-saving task unloading method for UAV initialization; S2, in each sub-slot t, the drone f traverses the search area calculates the minimum CPU computation frequency required for its local processing and the minimum transmission power required for its transmission to the base station wherein, the calculation formula is: , Wherein, B is a channel bandwidth, is a noise power, is a mixed channel gain between the unmanned aerial vehicle f and the base station bs, is a detection data amount of the observation sub-image, is a length of the offloading sub-slot in the sub-slot; If and , is the maximum computation frequency of the drone f, is the maximum transmission power of the drone f, then S2.1 is performed; if and , then S2.2 is performed; if and , then S2.3 is performed; if and , then S2.4 is performed; Step S2.1: The UAV adds the currently observed sub-image to the set of observed sub-images that cannot be unloaded, and then executes S2.5; Step S2.2: The UAV unloads the current observation sub-image to the base station and adds it to the set of observation sub-images that can only be unloaded to the base station, and then executes S2.5; S2.3, the UAV updates can only be unloaded to the local observation sub-image set and the variable of the sum of the CPU frequencies allocated to all observation sub-images unloaded to the local S2.5 is performed again. The update method is as follows: , wherein, is a coefficient related to the power of the drone f; S2.4, the unmanned aerial vehicle unloads the current observation sub-image locally and updates the set of observation sub-images that can be unloaded to the local and base station With S2.5 is performed again; The update manner is that: , like Then, drone f will be unloaded to the base station, and drone f can only update the set of observed sub-images unloaded to the base station. ,in, For the drone f, calculate at the minimum CPU frequency Process the power when the observed sub-image is OK, and then execute S2.5; S2.5, UAV f traverses the search area If all observed sub-images are OK, execute S3; S3, if If yes, then execute S4; otherwise, execute S5. S4. The drone will sort the observation sub-images that can be unloaded to the local area and the base station; S5. The final unloading strategy for all observation sub-images generated by the UAV includes: , , , Computing resource allocation strategy and , Used to store observation sub-images that could not be unloaded.
2. The UAV swarm cooperative target search method according to claim 1, characterized in that, The drone swarm collaborative target search architecture is as follows: the search area is discretized into multiple grids, and the search task of each grid is represented by the amount of search data and the required CPU processing density; a base station with an edge server is set at the center of the search area to provide image recognition services for the target search task, and laser charging piles are set at the boundaries; the target search task is divided into multiple sub-time slots of equal length, including flight sub-time slots and unloading sub-time slots; the drone has pre-loaded energy and a top-down camera at its bottom, and moves at a fixed speed in any direction in three-dimensional space during the flight sub-time slot, and charges using laser charging piles or performs target search tasks during the unloading sub-time slot. When performing target search tasks, the top-down camera captures the search area directly below to form a search observation image, which is divided into multiple observation sub-images; after all observation sub-images have completed image recognition in the unloading sub-time slot, the drone updates the target existence probability of the corresponding grid according to the recognition results.
3. The UAV swarm cooperative target search method according to claim 2, characterized in that, The image recognition of the observed sub-image is performed either by the UAV or by the UAV being unloaded to the base station and then performed by the base station.
4. The UAV swarm cooperative target search method according to claim 1, characterized in that, The joint optimization problem of UAV three-dimensional motion trajectory, charging decision, search task unloading location, and computing resource allocation is expressed as: , , , , , , , , , , , in, For UAV f in flight sub-time slot The magnitude of the fixed flight speed within; For UAV f in flight sub-time slot The angle between the direction of the flight velocity within the space and the z-axis in the three-dimensional Cartesian coordinate system; For UAV f in flight sub-time slot The angle between the direction of the flight velocity within the space and the x-axis in the three-dimensional Cartesian coordinate system; Charging strategies for drones; The unloading strategy for the observed sub-image is OK; The CPU frequency for observing sub-images is OK; The transmission power of the observed sub-image ok; T is the set of all sub-slots; K is the set of all grids within the search area; The grid at the start of sub-slot t The probability that a target exists within it; This indicates that a target exists in the grid. This indicates that there is no target in the grid; This indicates that a target has been detected in the grid at the start of sub-slot t. This indicates that no target was detected in the grid at the start of sub-slot t; This indicates whether a target has been detected in the grid at the start of time slot t+1; The information uncertainty of the grid; and All are weighting factors; Let f be the maximum flight speed of the drone; The maximum CPU computing frequency of the drone f; This represents the maximum transmit power of the UAV f. This is the maximum battery capacity limit for the drone f; The safe distance between the two drones; and These are the length and width of the search area, respectively; and These are the minimum and maximum safe flight altitudes for drones, respectively. A collection of all drones; This refers to the actual search area processed by the UAV. This represents the increase in the probability that the drone swarm detects the presence of a target within the grid within sub-time slot t; This represents the increase in the probability that the drone swarm detects the existence of a target within the search area during sub-time slot t; This represents the reduction in information uncertainty of the grid within sub-time slot t; This represents the reduction in information uncertainty within the search region during sub-time slot t; for and The distance between them; , , , , , , and All are constraints, among which, Constraints on the flight speed and direction of the UAV in a three-dimensional coordinate system; Charging scheduling constraints for drones; Unload location constraints for drone detection data; and These are constraints on the CPU computing frequency and transmission power allocated when the UAV unloads the observed sub-image and it is OK; Energy constraints for UAVs during search missions; Constraints on the safe distance between any two drones; To constrain the flight range of drones.
5. The UAV swarm cooperative target search method according to claim 1, characterized in that, The environmental state space is: , in, The remaining energy of the drone f at the beginning of sub-time slot t; Let f be the three-dimensional position coordinates of UAV f at the beginning of sub-time slot t; and These are the coordinates of the base station and the laser charging station, respectively. This indicates the grid at the start of sub-slot t. The probability of the target existing; This represents the number of remaining time steps in one round of reinforcement learning. The observation space of the drone is: , in, Let f be the probability distribution of targets in the grid that can be observed within the field of view of the UAV. The communication range of the drone f; This represents the minimum signal-to-noise ratio threshold for communication between drones. For drones f and drones The signal-to-noise ratio of communication between them; when hour, ;otherwise, ; The drone's action space is: , in, , and All are continuous variables, determining the flight sub-time slot of the UAV f. Flight trajectory within; As a discrete variable, it determines the unloading sub-slot of the UAV. Internal charging strategy; when season This indicates that the drone f is unloading the sub-time slot. Internally charged from a laser charging station; when season This indicates that the drone f is unloading the sub-time slot. Perform search area probing and probe data unloading; Let f be the maximum flight speed of the drone; The reward function is: , in, The penalty value for drone f colliding with other drones; and All are weighting factors; K is the set of all grids in the search region; Indicates whether a target has been detected in the grid at the start of sub-slot t; This represents the increase in the probability that the drone swarm detects the presence of a target within the grid within sub-time slot t; This represents the increase in the probability that the drone swarm detects the existence of a target within the entire search area within sub-time slot t; This represents the reduction in information uncertainty of the grid within sub-time slot t; This represents the reduction in information uncertainty across the entire search region within sub-time slot t.
6. The UAV swarm cooperative target search method according to claim 1, characterized in that, The policy network includes: The local observations of the UAV f are divided into its own characteristics. and its communication range Features of other UAVs j , ; Using embedded layer fully connected networks, respectively from and Extracting features and And send it to the multi-head attention head to obtain attention value. , for: , , in, Let f be the eigenvalue of the UAV. For the communication range of UAV f The eigenvalues of other UAVs j; , and These are the weight matrices for keys, queries, and values, respectively. for The transpose of the matrix; As a key dimension; Score for attention; at last, and The data are concatenated and sent to the fully connected network of the output layer to obtain the probability distribution of the drone f's actions.
7. The UAV swarm cooperative target search method according to claim 5, characterized in that, The evaluation network includes: Global state information is divided into environmental information. Information about drones Embedded layer fully connected network from and Extract two types of features respectively and , For environmental information eigenvalues, Information for drone f eigenvalues; Will and Send to Each primary multi-head attention head extracts attention values. and , Total number of drones; and It is then sent to a fully connected network to obtain the attention output vector. and ;Will and Send to two levels of multi-head attention heads to obtain attention values. ; and It is then sent to the fully connected network at the output layer to obtain the final state value evaluation score.
8. The UAV swarm cooperative target search method according to claim 1, characterized in that, The energy of UAV f at the start of sub-time slot t+1 is: ,in, and These represent the energy of the drone f at the beginning of sub-time slot t and sub-time slot t+1, respectively. Due to the battery power limitation of the drone f, The function ensures that the drone's energy consumption will not exceed the battery's capacity limit while it is charging. The energy consumed by the UAV f to transmit the observed sub-images to the remote base station. For UAV f in flight sub-time slot The energy consumed by flight within the spacecraft; Indicates that the drone f is in Perform the search task. This indicates that the drone f is charging from the laser charging station; To unload sub-slots The total computing energy consumed by the local computing of the drone. For drones f in Energy collected during the period; if If the drone f is depleted of energy, it will terminate its current mission.