Extensible multi-agent deep reinforcement learning unmanned aerial vehicle cluster collaborative search method based on digital twinning

Through digital twin technology and deep reinforcement learning of multiple agents, a high-fidelity training environment is built, reward functions and graphical representations are designed, and the collaborative search of drone clusters is optimized, which solves the problem of collaborative search of large-scale drone clusters in dynamic environments, and achieves efficient and scalable target search.

CN120409986APending Publication Date: 2025-08-01NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510147160.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing UAV cluster collaborative target search method is difficult to achieve efficient and scalable collaborative search in dynamic environments, especially in large-scale scenarios. The existing methods are difficult to achieve ideal results in reality after training in simulated environments, and the computing resources are large.

Method used

Using a scalable multi-agent deep reinforcement learning method based on digital twins, we use a high-fidelity training environment to design reward functions and combine graphic representations and environmental cognition maps to optimize target search rate, regional coverage, search time and collaborative security, and use local information to make distributed decisions.

Benefits of technology

It improves the collaborative search performance and scalability of drone clusters in dynamic environments, can discover more targets and reduce collisions in different scale tasks, and improves search efficiency and coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409986A_ABST
    Figure CN120409986A_ABST
Patent Text Reader

Abstract

The invention discloses an extensible deep reinforcement learning unmanned aerial vehicle cluster collaborative search method based on digital twinning. According to the method, an extensible multi-agent reinforcement learning algorithm SAMARL is adopted, a multi-agent near-end strategy optimization algorithm and a multi-head attention mechanism are fused, and the effectiveness and the adaptability of the algorithm are improved. In an SAMARL framework, a complex observation space containing graphic representation and an environment cognitive map is constructed, and the target search rate, the area coverage rate, the search time and the collaborative security are comprehensively optimized. An observation space based on a view field range is designed, and the model training speed is improved. Then, a search training framework of'twin training, distributed execution and persistent evolution 'is provided, and migration of the decision model from simulation to reality is achieved; the intelligent decision model is deployed to each unmanned aerial vehicle, and the cluster uses a field intensity meter and a communication radio station to cooperatively search unknown signal sources in a distributed manner. A large number of experimental results show that the method is superior to an existing advanced strategy in search efficiency and system expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of cooperative target search for UAV swarms, and particularly relates to a cooperative target search method for UAV swarms based on digital twin technology and scalable deep reinforcement learning, which helps to realize the migration of the intelligent decision-making model from the simulation environment to the real scenario. Background Art

[0002] With the rapid progress of computing technology, UAV (Unmanned Aerial Vehicle) target search has been widely applied in many fields, such as rescuing survivors, tracking suspects, and locating suspicious targets. Compared with a single UAV, multi-UAV systems show obvious advantages in flexibility, robustness, and fault tolerance, especially when performing complex tasks. Therefore, multi-UAV systems have received extensive attention in target search tasks. However, cooperative target search (CTS) is still a challenging problem. On the one hand, UAVs need to avoid collisions with surrounding UAVs or other threats in dynamic scenarios when performing target search tasks; on the other hand, due to the unknownness and dynamicity of large-scale scenarios, combined with the limitation of the detection range, it is difficult for UAVs to achieve real-time cooperation to cope with changes.

[0003] To solve the CTS problem, researchers have proposed various algorithms, including traditional algorithms and swarm intelligence optimization (SIO) algorithms. Traditional algorithms such as formation search and partition search are mainly applicable to static target search in simple and known environments. However, when facing dynamic threats or changes in the number of UAVs, these methods need to re-plan the flight path. Algorithms based on swarm intelligence optimization, such as genetic algorithm (GA) and ant colony optimization (ACO), still face challenges in making real-time decisions in highly dynamic changes, especially when the search scenario expands, it is difficult to obtain global information and meet the search accuracy requirements.

[0004] In recent years, Deep Reinforcement Learning (DRL) has gradually received attention in CTS problems. However, applying DRL to multi-agent systems in unknown environments still poses many challenges. Therefore, researchers have turned to Multi-Agent Deep Reinforcement Learning (MADRL) to enhance the autonomy and intelligence of UAV systems. In CTS problems, uneven target distribution and dynamic threats can significantly affect search performance. Compared with centralized methods, distributed methods have advantages in reducing computational burden and dependence on global information, and are more suitable for large-scale collaborative search tasks. In complex scenarios, it is crucial to develop distributed methods with high collaboration efficiency, scalability, and adaptability. The CTS problem is usually regarded as a Partially Observable Markov Decision Process (POMDP), which further highlights the potential of MADRL methods. Currently, most popular CTS methods are trained in simulation environments and then applied to real scenarios. However, due to the gap between simulation and reality and the complexity of the task environment, it is often difficult to achieve ideal results in practical applications. In addition, since large-scale UAV swarm collaboration requires a large amount of computing resources, CTS models based on reinforcement learning are rarely applied to actual UAVs, especially in large-scale scenarios. Existing research mainly focuses on small-scale tasks.

[0005] Learning-based UAV swarm algorithms often face challenges such as low sample efficiency, difficult model training, and insufficient scalability in real-world applications. In recent years, Digital Twin (DT) technology has emerged as a potential solution to bridge the physical and digital domains. By combining Machine Learning (ML) technology, DT can use real-time state data for model training, and then construct a high-fidelity digital twin in the virtual space, significantly improving training efficiency and model accuracy. This technology has been applied in the field of UAV cooperative control, such as trajectory planning and task allocation. However, there is currently little research on combining algorithms based on Multi-Agent Deep Reinforcement Learning (MADRL) with DT technology to achieve large-scale, distributed collaborative target search, especially in real-world environments full of uncertainties. Therefore, there are still many challenges that need to be addressed urgently to improve the performance and scalability of UAV swarms in distributed collaborative search. Summary of the Invention

[0006] The object of the present invention is to address the problem of collaborative target search for UAV swarms, and propose a scalable multi-agent deep reinforcement learning UAV swarm collaborative search method based on digital twin. This research constructs a high-fidelity training environment for the UAV swarm collaborative target search model based on MADRL, taking into account both the training speed and the effectiveness of the decision-making model. To achieve this object, the steps adopted by the present invention are as follows:

[0007] Step 1: According to the detection results of UAV sensors at each time step, construct a UAV swarm sensing and detection model to evaluate the confidence of the UAV swarm in detecting the presence of targets in the search area at each time step. By constructing a complex observation space including graphical representation and environmental cognitive map, comprehensive optimization of various performance indicators such as target search rate, area coverage rate, search time, and collaborative security is achieved.

[0008] Step 2: Design a scalable multi-agent deep reinforcement learning UAV swarm collaborative search reward function. The reward function includes exploration and cognition reward, target discovery reward, time consumption reward, area coverage reward, and collaborative security reward, and the final reward function is a linear coupling of the above five parts.

[0009] Step 3: Obtain the observations of the UAV swarm collaborative search problem in the simulation environment, input them into the policy network of SAMARL to generate UAV detection actions and obtain environmental feedback rewards. Subsequently, continuously train the SAMARL model until the reward function converges stably, and finally generate a UAV swarm collaborative search decision-making model.

[0010] In the actual environment, due to the limited detection radius and actual communication of UAVs, UAVs cannot obtain global environmental information. In the present invention, each UAV detects the unknown X-band target signal source by carrying a field strength meter and returns the signal strength from the signal sources in four directions. Assume the speed v at the current time step i , t and the angle between the speed v i , t+1 at the next time step cannot exceed the maximum turning angle ψ, then the speed of the UAV must satisfy:

[0011]

[0012] The signal strength detected in each direction can be described by the following formula:

[0013]

[0014] Among them, P r represents the received power, P t represents the transmitted power, G t and G rThey are the transmitting antenna gain and the receiving antenna gain respectively, λ is the wavelength of the signal, and d is the distance from the signal source to the receiver. The signal intensity finally detected by the drone is the signal intensity of the maximum value among the four directions.

[0015] P r,max =max(P r1 ,P r2 ,P r3 ,P r4 ) (3)

[0016] Among them, P r1 、P r2 、P r3 and P r4 are the received signal intensities from the four directions respectively.

[0017] First, considering factors such as the limited detection ability of the sensor due to noise, the conditional probabilities of the positive and negative characteristics of the sensor model can be expressed as:

[0018]

[0019] For the convenience of calculation, a linear update method is introduced. Formula (4) is transformed into

[0020]

[0021] Among them:

[0022]

[0023] Then, formula (5) is equivalent to:

[0024]

[0025] Furthermore, the observation space and action space of the UAV swarm cooperative search problem are specifically:

[0026] The observation space of the UAV is defined as a square area centered on itself. When designing the state space for each agent, the present invention sets a field of view (FOV) in the state space to reduce the dimension of the neural network input and improve the training speed. More importantly, a graph representation method is introduced to construct an expandable agent state space, thereby enhancing the generalization ability for task regions of any size. At each time step, the UAV extracts local information from its field of view as the input of the agent network, which helps the UAV determine its next action. The information extracted from the environmental cognitive map can be divided into four parts:

[0027] (1) Target presence probability map ρ1: It is extracted from the local probability map of the UAV and updated through information sharing. The target presence probability outside the boundary is assumed to be 0;

[0028] (2) Drone visit count information graph ρ2: If a drone visits a cell in a certain time step, the visit count of that cell is incremented by one, and the visit count outside the boundary is assumed to be 0;

[0029] (3) Neighbor and threat information graph ρ3: A threat in the field of view is represented as 0.5, other drones in the field of view are represented as 1, and empty cells or cells outside the boundary are represented as 0;

[0030] (4) Drone position ρ4: An observation vector is added to help the drone understand its own state.

[0031] The action space of the drone consists of the drone's action choices. At each time step, the drone selects from up to five possible actions based on its current position, and these actions correspond to the north, south, east, west, and stationary directions. Any actionable move that may cause the drone to go out of bounds in the next step will be discarded.

[0032] Furthermore, the specific reward function for scalable multi-agent deep reinforcement learning cluster cooperation target search in digital twin training is as follows:

[0033] (1) Exploration and cognition reward: This reward function is used to guide the drone to search for more targets. If the target existence probability is greater than the predefined threshold ξ, it is considered that the target has been found. It should be noted that the target reward can only be obtained when the first agent finds the target for the first time. The target discovery reward at time step t is defined as:

[0034]

[0035] where α is a positive constant, is the information entropy at time t.

[0036] (2) Target discovery reward: This reward function is used to guide the drone to discover more targets. Therefore, in order to motivate the drone to discover new targets at time t, the drone i will be given a positive reward:

[0037]

[0038] where β is a positive constant.

[0039] (3) Time consumption reward: This reward function aims to guide the drone to complete the search task in the shortest time. The time consumption reward for each time step is given by:

[0040] r time,t =-λ (10)

[0041] where λ is a positive constant.

[0042] (4) Area Coverage Reward: This reward function is used to guide the drone to cover more unexplored grid cells. Therefore, the unknown area coverage reward is defined as:

[0043]

[0044] where μ is a positive constant used to adjust the reward, and N u is the number of drones. N new,t,i is the number of grid cells newly covered by drone i at time step t. S tot is the total number of grid cells in the entire area.

[0045] (5) Cooperative Safety Reward: This reward function aims to guide the drones away from threats and other drones. Therefore, the collision penalty is defined as:

[0046]

[0047] where γ and η are positive constants.

[0048] In summary, the complete reward function at time step t is expressed as:

[0049] r t = r cog,t + r tar,t + r time,t + r cov,t + r safe,t (13)

[0050] Furthermore, the training method of the SAMARL network is specifically as follows:

[0051] SAMARL is an extension based on the multi-agent proximal policy optimization algorithm. Among them, the policy π outputs the probability distribution of actions according to the environmental input s, draws an action from this distribution, and executes the selected action a in the environment. Subsequently, the agent receives the reward r and transitions to the next state s'. Let π θ represent the policy parameterized by the parameter θ, and J(π θ ) represent the long-term cumulative reward under this policy. SAMARL follows the framework of "centralized training, decentralized execution" to train the policy network and the evaluation network. In the centralized training stage, the evaluation network evaluates the policy network of each drone based on the global state. In the decentralized execution stage, each drone makes real-time behavior decisions according to its policy network using local observation data. The evaluation network is defined by the parameter φ and is trained to minimize the loss function:

[0052]

[0053] where B represents the batch size, N represents the number of agents, represents the discounted reward;

[0054] SAMARL integrates the multi - head attention mechanism in the policy network, enabling the agent to focus more on important observation data from neighbors. Specifically, after weighted aggregation of the suggestions according to the confidence, the local cognition of each drone is enhanced to a collection intention:

[0055]

[0056] Among them, f u represents the features of other agents except the current agent i. l and v are regarded as cognition and intention respectively. The confidence is defined as the dot product of l of other agents and its own query q. q, l, and v are linear projections of the local observation o i . Specifically, l j represents the cognition of agent j, v j represents the suggestion of agent j, and d x is a scaling factor.

[0057] The final action value is determined by combining the collective intention and the personal information of the agent, which is given as follows:

[0058]

[0059] Among them, f(·) is a fully - connected neural network layer. Based on the collective intention z i and the historical information of the agent use the MLP network to carefully extract insightful information and further refine the details of the individual agent. The personal information of the agent is:

[0060]

[0061] Among them, GRU(·) represents the gated recurrent unit. After obtaining the loss function, all parameters are trained by scalable multi - agent deep reinforcement learning through backpropagation. Description of the Drawings

[0062] Figure 1 is a system model diagram of scalable drone swarm collaborative target search based on digital twin;

[0063] Figure 2 is a block diagram of the scalable multi - agent deep reinforcement learning method proposed by the present invention;

[0064] Figure 3 is a training environment diagram of drone swarm collaborative target search based on digital twin;

[0065] Figure 4It is a comparison chart of the search performance results of the present invention and existing methods for tasks of different scales;

[0066] Figure 5 It is a comparison chart of the anti-collision performance results of the present invention and existing methods for tasks of different scales;

[0067] Figure 6 It is the target search probability chart of the present invention in a large-scale scenario. Detailed implementation manners

[0068] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0069] In the following description, this specification abbreviates the scalable multi-agent deep reinforcement learning cluster collaborative target search method based on digital twin training proposed by the present invention as SAMARL.

[0070] SAMARL first sets the following operating conditions:

[0071] 1. As Figure 1 shown, the search area Ω is divided into L x ×L y equal-sized discrete grids. The central coordinate of each grid is defined as c x,y =(x, y), where x ∈ {1, 2,..., L x}}, y ∈ {1, 2,...,, L y}}. Each target can only occupy one grid, and the probability of the existence of a target in the grid is modeled by a Bernoulli distribution, that is, τ xy =1 represents the probability that a single target exists in the cell C x,y is P x,y , while τ xy =0 represents the probability that there is no target in the cell C x,y is 1-P x,y . Suppose there are M unknown target signal sources, denoted as M = {1, 2,..., M}, and each target emits signals in all directions. Considering the dynamic threat environment is randomly distributed in the search area Ω. The coordinates of threat k are denoted as z k =(x, y), where 1 ≤ k ≤ N o .

[0072] 2. The unmanned aerial vehicle cluster system U = {U1, U2,...U i …, U Nu}(1 ≤ i ≤ N u ) contains N u homogeneous unmanned aerial vehicles. For the convenience of analysis, the mission cycle is discretized into T time steps, and the unmanned aerial vehicle only makes a movement decision at the beginning of each time step. Therefore, the coordinate of unmanned aerial vehicle i at time step t (1 ≤ t ≤ T) is denoted as ui,t =(x, y). Each drone can choose one direction from five directions to move: up, down, left, right, and stay still. Let σ denote the moving distance, then the coordinates of the i-th drone at time step t + 1 can be expressed as u i,t+1 =(x, y). It should be noted that in order to adapt to the motion constraints of actual drones, the motion direction of the drone cannot switch violently between two consecutive steps. Assume that at time step t, the speed of the drone is v i,t , then in the next time step, the speed v i,t+1 must satisfy the following conditions:

[0073]

[0074] The global task environment is invisible to the drones, and there are threats and targets at unknown locations. Each drone detects the target signal source through a field strength meter and returns the signal strength from four direction sources. The signal strength detected in each direction can be described as:

[0075]

[0076] where P r represents the received power, P t represents the transmitted power, G t and G r are the transmitting antenna gain and the receiving antenna gain respectively, λ is the wavelength of the signal, and d is the distance from the signal source to the receiver. The signal strength finally detected by the drone is the signal strength of the maximum value among the four directions.

[0077] P r,max = max(P r1 , P r2 , P r3 , P r4 ) (20)

[0078] where P r1 , P r2 , P r3 and P r4 are the received signal strengths from the four directions respectively. Considering factors such as limited detection ability of the sensor due to noise, the conditional probabilities of the positive and negative characteristics of the sensor model can be expressed as:

[0079]

[0080] where p, 1 - p, q, 1 - q represent the detection, missed detection, false alarm, and non-detection probabilities respectively. To ensure effective detection in a noisy environment, the detection probability p and the false alarm probability q are set in the ranges of [0.5, 1] and [0, 0.5] respectively.

[0081] Based on the above conditions, SAMARL proposed in the present invention has been implemented in Ubuntu, and the experimental results have proved the effectiveness of this method. The specific implementation steps of SAMARL are as follows:

[0082] Step 1: According to the detection results of the UAV sensors at each time step, construct a UAV swarm sensing detection model to evaluate the confidence of the UAV swarm in detecting the presence of targets in the search area at each time step.

[0083] Each UAV i maintains an information probability map of an independent search area Ω at time step t First of all, due to the imperfect detection ability of the sensors, each UAV needs to update its own information probability map according to the observation results. The commonly used update method is based on the Bayesian rule, that is:

[0084]

[0085] For the convenience of calculation, a linear update method is introduced. First, transform formula (22) into:

[0086]

[0087] Set:

[0088]

[0089] Then, formula (22) is equivalent to:

[0090]

[0091] Step 2: Combine the probability of the presence of each grid target by the UAVs to define the observation space and action space of the UAV swarm collaborative search problem. The observation space of the UAV consists of the target presence probability map, the access times map, the neighbor and threat map, and the position of the UAV.

[0092] In order to optimize the state space design of each agent, the present invention adopts the FOV range, reduces the input dimension, and accelerates the training of the neural network. In addition, the present invention also designs an observation space based on graphical representation for the agent, enhancing the generalization ability of tasks for UAV swarms of different scales. At each time step, the UAV extracts the local information within the field of view as the input of the agent network to help the UAV determine the next action. The information extracted from the environmental cognitive map can be divided into four parts:

[0093] (1) Target presence probability map ρ1: It is extracted from the local probability map of the UAV and updated through information sharing. The target presence probability outside the boundary is assumed to be 0;

[0094] (2) Information graph of access times ρ2: If a drone accesses a cell in a certain time step, the access times of that cell increase by one, and the access times outside the boundary are assumed to be 0;

[0095] (3) Information graph of neighbors and threats ρ3: The threats in the field of view are represented as 0.5, other drones in the field of view are represented as 1, and empty cells or cells outside the boundary are represented as 0;

[0096] (4) Position of the drone ρ4: Observation vectors are added to help the drone understand its own state.

[0097] In each time step, the drone can choose from up to five possible actions according to its current position, and these actions correspond to the directions of north, south, east, west, and stationary. Any possible action that may cause the drone to cross the boundary in the next step is discarded from the set of actionable actions.

[0098] In summary, the observation space of the drone consists of the target presence probability information graph, the grid access times information graph, the neighbor and threat information graph, and the self-position information.

[0099] O i =[ρ 1,i ,ρ 2,i ,ρ 3,i ,p 4,i (26)

[0100] Step 3: Design a reward function for scalable multi-agent deep reinforcement learning for UAV swarm collaborative search. The reward function includes exploration and cognition reward, target discovery reward, time consumption reward, area coverage reward, and collaborative security reward, and the final reward function is a linear coupling of the above five parts.

[0101] The specific definition of the reward function is as follows:

[0102] (1) Exploration and cognition reward: This reward function is used to guide the drone to search for more targets. If the target presence probability is greater than the predefined threshold ξ, it is considered that the target has been found. It should be noted that only when the first agent finds the target for the first time can it obtain the target reward. The target discovery reward at time step t is defined as:

[0103]

[0104] where α is a positive constant, is the information entropy at time t.

[0105] (2) Target discovery reward: This reward function is used to guide the drone to discover more targets. Therefore, in order to motivate the drone to discover new targets at time t, the drone i will be given a positive reward:

[0106]

[0107] Among them, β is a positive constant.

[0108] (3) Time - consumption reward: This reward function aims to guide the UAV to complete the search task in the shortest time. The time - consumption reward for each time step is given by:

[0109] r time,t = -λ (29)

[0110] Among them, λ is a positive constant.

[0111] (4) Area - coverage reward: This reward function is used to guide the UAV to cover more unexplored grid cells. Therefore, the unknown - area coverage reward is defined as:

[0112]

[0113] Among them, μ is a positive constant used to adjust the reward, N u is the number of UAVs. N new,t,i is the number of grid cells newly covered by UAV i at time step t. S tot is the total number of grid cells in the entire area.

[0114] (5) Cooperative - safety reward: This reward function aims to guide the UAVs to stay away from threats and other UAVs. Therefore, the collision penalty is defined as:

[0115]

[0116] Among them, ω5 is a positive constant.

[0117] To sum up, the complete reward function at time step t is expressed as:

[0118] r t = r cog,t + r tar,t + r time,t + r cov,t + r safe,t (32)

[0119] Step 4: Obtain the observations of the UAV swarm cooperative - search problem in the simulation environment, input them into the policy network of SAMARL, generate UAV detection actions and obtain environmental feedback rewards. Subsequently, continuously train the SAMARL model until the reward function converges stably, and finally generate the UAV swarm cooperative - search decision - making model.

[0120] SAMARL adopts the policy gradient method, where the policy π receives the environmental input s, outputs a probability distribution of an action, samples an action from this distribution, and executes the selected action in the environment. Subsequently, the agent receives a reward r and transfers to the next state s′. Let π θ denote the policy parameterized by θ, and J(π θ ) represent the long-term cumulative reward under this policy. The gradient update rule of the policy gradient method is as follows:

[0121]

[0122] where τ represents the agent's trajectory, and A πθ denotes the generalized advantage estimation based on the γ discount factor and the λ weighting factor, which is given by the following formula:

[0123]

[0124] To improve the data utilization efficiency on-policy, the importance sampling technique is introduced, using the data sampled from the previous policy instead of being restricted to the data under the current policy. Suppose the previous policy is π ω , and the updated policy is π θ , the gradient update formula can be expressed as:

[0125]

[0126] where r(a|s) is the probability ratio, which is given by:

[0127]

[0128] Then the loss function is expressed as:

[0129]

[0130] Although the modified gradient formula improves the data relevance, the significant difference between the old and new policy distributions may lead to inaccurate estimation of the gradient expectation. Therefore, the difference between the old and new policies must be restricted. To ensure that the difference between the old and new policy distributions is not too large, SAMARL adopts a clipping operation. The Actor network is parameterized by θ and is trained by maximizing the following objective

[0131]

[0132] Among them, clip(·) represents the truncation function, and ε represents the hyperparameter used to measure the interval between the old and new policy distributions. This formula constrains the ratio of the old and new policy distributions within the range of (1 - ε, 1 + ε). SAMARL uses a framework of "twin training, distributed execution, and continuous evolution" to train the Actor network and the Critic network. In the centralized training stage, the Critic network evaluates the Actor network of each drone based on the global state. In the distributed execution stage, each drone makes real-time decisions using its Actor network according to the local observations. The Critic network is represented by the parameter φ, and its loss function is defined as follows:

[0133]

[0134] As the number of drones continues to increase, the complexity of the model increases significantly, resulting in a decline in the performance of traditional fully connected networks and a slow convergence speed. For this reason, SAMARL integrates the multi-head attention mechanism into the Actor network, enabling the agent to focus on more important observation data from neighbors. Specifically, after weighting and aggregating the suggestions according to the confidence, the local cognition of each drone is enhanced to collect intentions,

[0135]

[0136] where f u represents the features of other agents except the current agent i. Let l and v be regarded as cognition and intention respectively. The confidence is the dot product between l of other agents and the query q of itself. q, l, and v are linear projections of the local observation o i . Among them, l j represents the cognition of agent j, v j represents the suggestion of agent j, and d x is the scaling coefficient.

[0137] Then, the collective intention of the agent is combined with the individual information to determine the final action value, where is:

[0138]

[0139] where f(·) is a fully connected neural network layer. Based on the collective intention z i and the historical information of the agent , use the MLP network to carefully extract insightful information and further refine the details of the individual agent. The personal information of the agent is:

[0140]

[0141] Among them, GRU(·) represents the gated recurrent unit.

[0142] The performance of a scalable deep reinforcement learning cluster collaborative search method based on digital twin proposed in the present invention has been verified by experimental results. Different from the existing algorithms trained based on learning algorithms, this method also introduces a digital twin-driven training framework, uses a high-fidelity digital twin model to improve the effectiveness of the intelligent decision-making model, and realizes continuous evolution through parallel sampling in the virtual environment and periodic update of the integrated model weights. In the training stage, the decision-making model is trained using data sampled simultaneously from multiple digital twin training copies, providing a rich variety of training samples for the reinforcement learning algorithm. These copies achieve a wide range of behaviors to adapt to different task conditions. Each drone receives the output of the decision-making model as the search strategy. In the execution stage, the trained model is deployed on the drones to perform the target search task in a distributed manner. The distributed drone swarm digital twin collaborative target search verification system adopted in the present invention, including real flight control, communication simulation tools, and a three-dimensional physical engine, helps to solve the problem of the difficult practical application of the collaborative target search strategy based on multi-agent deep reinforcement learning.

[0143] The detection probability p of the drone sensor is 0.85, the false alarm probability q is 0.15, the initial target existence probability is 0.5, the safety distance is 2, and the detection threshold ξ is 0.99. The simulation experiment compares the scalability of SAMAR with MAPPO, QMIX, MADDPG, and ACO on four search tasks: "30a5u15t3o" (in a 30×30 area, 5 drones search for 15 targets and avoid 3 dynamic threats), "40a10u20t3o" (in a 40×40 area, 10 drones search for 20 targets and avoid 3 dynamic threats), "50a15u25t5o" (in a 50×50 area, 15 drones search for 25 targets and avoid 5 dynamic threats), "60a20u30t70" (in a 60×60 area, 20 drones search for 30 targets and avoid 7 dynamic threats). Attached Figure 4 and attached Figure 5 gives the performance comparison of the scalable cluster collaborative target search method proposed in the present invention and the existing methods in terms of the number of searched targets and the number of collisions under different scale task conditions, Figure 6 which is the target search probability map for large-scale scenarios. The results show that the method proposed in the present invention can discover more unknown targets and improve the regional coverage rate compared with the existing comparative methods, proving the effectiveness and scalability of the present invention in different scale tasks.

[0144] The content not described in detail in the present invention application belongs to the prior art well-known to those skilled in the art.

Claims

1. A scalable deep reinforcement learning UAV swarm collaborative search method based on digital twin, and the steps adopted are as follows: Step 1: According to the detection results of UAV sensors at each time step, construct a UAV swarm sensing detection model to evaluate the confidence of the UAV swarm in detecting the presence of targets in the search area at each time step; Each UAV detects the target signal source through a field strength meter and returns the signal strength from four direction sources. The signal strength detected in each direction can be described as: Where, P r represents the received power, P t represents the transmission power, G t and G r are the transmitting antenna gain and the receiving antenna gain respectively, λ is the wavelength of the signal, d is the distance from the signal source to the receiver, and the signal intensity finally detected by the drone is the signal intensity of the maximum value among the four directions: P r,max = max(P r1 , P r2 , P r3 , P r4 ) (2) Where Pr1, Pr2, Pr3, and Pr4 are the received signal strengths from four directions respectively. Considering factors such as noise and limited detection capabilities of sensors, the conditional probabilities of the positive and negative characteristics of the sensor model can be expressed as: Where p, 1 - p, q, and 1 - q represent the detection, missed detection, false alarm, and non - detection probabilities respectively; To ensure effective detection in a noisy environment, the detection probability p and the false alarm probability q are set in the ranges of [0.5, 1] and [0, 0.5] respectively; For the convenience of calculation, a linear update method is introduced. First, transform formula (3) into: For the convenience of calculation, a linear update method is introduced. First, transform formula (4) into: Set: Then, formula (5) is equivalent to: Step 2: Combine the confidence level of the UAV in the probability of the existence of each grid target, define the observation space and action space of the UAV swarm collaborative search problem, and realize the comprehensive optimization of various performance indicators such as target search rate, area coverage rate, search time, and collaborative security by constructing a complex observation space including graphical representation and environmental cognitive map; To optimize the state - space design of each agent, the present invention adopts the field - of - view (FOV) range, which reduces the input dimension and accelerates the training of the neural network. In addition, the present invention also designs an observation space based on graphical representation for the agent, enhancing the generalization ability of tasks for UAV swarms of different scales. At each time step, the UAV extracts local information within the FOV as the input of the agent network to help the UAV determine the next action. The information extracted from the environmental cognitive map can be divided into four parts: (1) Target existence probability map ρ1: It is extracted from the local probability map of the UAV and updated through information sharing. The target existence probability outside the boundary is assumed to be 0; (2) Visit count information map ρ2: If a UAV visits a cell in a certain time step, the visit count of this cell increases by one. The visit count outside the boundary is assumed to be 0; (3) Neighbor and threat information map ρ3: The threat in the field of view is represented as 0.5, other UAVs in the field of view are represented as 1, and empty cells or cells outside the boundary are represented as 0; (4) Position of the UAV ρ4: An observation vector is added to help the UAV understand its own state; At each time step, the UAV can select from up to five possible actions according to its current position. These actions correspond to the directions of north, south, east, west, and stationary. Any possible action that may cause the UAV to cross the boundary in the next step is discarded from the set of actionable actions; In summary, the observation space of the UAV consists of the target existence probability map, the grid access frequency map, the neighbor and threat map, and its own position information; O i = [ρ 1,i , ρ 2,i , ρ 3,i , p4, i (8) Step 3: Design an extensible multi-agent deep reinforcement learning UAV swarm collaborative search reward function. The reward function includes exploration and cognition reward, target discovery reward, time consumption reward, area coverage reward, and collaborative security reward. The final reward function is a linear coupling of the above five parts; The specific definition of the reward function is as follows: (1) Explore cognitive rewards: This reward function is used to guide the drone to find more targets. If the target existence probability is greater than the predefined threshold ξ, it is considered that the target has been found. It should be noted that only when the first agent finds the target for the first time can the target reward be obtained. The discovery target reward at time step is defined as: where α is a positive constant, is the information entropy at time t; (2) Target discovery reward: This reward function is used to guide the UAV to discover more targets. Therefore, in order to encourage the UAV to discover new targets at time t, the UAV i will be given a positive reward: where β is a positive constant; (3) Time consumption reward: This reward function aims to guide the UAV to complete the search task in the shortest time. The time consumption reward for each time step is given by: r time,t = -λ (11) where λ is a positive constant; (4) Area coverage reward: This reward function is used to guide the UAV to cover more unexplored grid cells. Therefore, the unknown area coverage reward is defined as: where μ is a positive constant used to adjust the reward, and N u is the number of drones N new,t,i is the number of newly covered grid cells of drone i at time step t, and S tot is the total number of grid cells in the entire area; (5) Collaborative security reward: This reward function aims to guide the UAV to stay away from threats and other UAVs. Therefore, the collision penalty is defined as: where γ, η are positive constants; In summary, the complete reward function at time step t is expressed as: r t =r cog,t +r tar,t +r time,t +r cov,t +r safe,t (14) Step 4: Obtain the observations of the UAV swarm collaborative search problem in the simulation environment, input them into the policy network of SAMARL, generate the UAV detection actions and obtain the environmental feedback rewards. Subsequently, continuously train the SAMARL model until the reward function converges stably, and finally generate the UAV swarm collaborative search decision-making model; For ease of reading and understanding, in the following narrative, random variables are represented by uppercase letters, the specific realizations of random variables are represented by lowercase letters, and the vector forms of random variables are represented by boldface; SAMARL adopts the policy gradient method, where the policy π receives the environmental input s, outputs a probability distribution of an action, samples an action from this distribution, and executes the selected action in the environment. Subsequently, the agent receives a reward r and transfers to the next state s′. Let π θ denote the policy parameterized by θ, and J(π θ ) represent the long-term cumulative reward under this policy. The gradient update rule of the policy gradient method is as follows: where τ represents the trajectory of the agent, and A πθ represents the generalized advantage estimation based on the γ discount factor and the λ weighting factor, which is given by the following formula: To improve the data utilization efficiency in the policy, the importance sampling technique is introduced, using data sampled from the previous policy instead of being restricted to the data under the current policy. Assume the previous policy is π ω , and the updated policy is π θ . The gradient update formula can be expressed as: where r(a|s) is the probability ratio, given by: Then the loss function is expressed as: Although the modified gradient formula improves the data correlation, the significant difference between the old and new policy distributions may lead to inaccurate estimation of the gradient expectation. Therefore, the difference between the old and new policies must be restricted. To ensure that the difference between the old and new policy distributions is not too large, SAMARL adopts a clipping operation. The Actor network is parameterized by θ and is trained by maximizing the following objective: where clip(·) represents the truncation function, ε represents the hyperparameter used to measure the gap between the old and new policy distributions. This formula constrains the ratio of the old and new policy distributions to be within the range of (1 - ε, 1 + ε). SAMARL uses a "twin training, distributed execution, and continuous evolution" framework to train the Actor network and the Critic network. In the centralized training stage, the Critic network evaluates the Actor network of each UAV based on the global state. In the distributed execution stage, each UAV makes real-time decisions using its Actor network according to the local observations. The Critic network is represented by the parameter φ, and its loss function is defined as follows: Among them, B represents the batch size, N represents the number of agents, represents the discounted reward; With the continuous increase in the number of drones, the complexity of the model has increased significantly, resulting in a decline in the performance of traditional fully connected networks and a slow convergence speed. Therefore, SAMARL integrates a multi-head attention mechanism into the Actor network, enabling the agent to focus on more important observation data from neighbors. Specifically, after weighted aggregation of the suggestions according to the confidence, the local cognition of each drone is enhanced to collect intentions: Among them, f u represents the features of other agents except the current agent i. Regarding l and v as cognition and intention respectively, the confidence is the dot product between l of other agents and its own query q. g, l, and v are linear projections of the local observation o i , where l j represents the cognition of agent j, and v j represents the suggestion of agent j, and dx is the scaling coefficient; Then, combine the collective intention of the agent with the individual information to determine the final behavior value, where is: where f(·) is a fully connected neural network layer, based on the collective intention z i and the historical information of the agent Use the MLP network to carefully extract insightful information and further refine the details of individual agents. The personal information of the agent is as follows: where GRU(·) represents the gated recurrent unit; after obtaining the loss function, all parameters are trained through scalable multi-agent deep reinforcement learning using backpropagation.

Citation Information

Cited By

  • Multi-unmanned aerial vehicle cooperative search and rescue intelligent decision-making method and device based on hierarchical intention

    CN120672085A

  • Dynamic platform unmanned aerial vehicle autonomous landing method and system based on deep reinforcement learning and digital twinning

    CN121143445A

  • Distributed photovoltaic scheduling method based on multi-agent consensus optimization

    CN121436615A

  • Modularized construction dynamic hoisting scheduling decision-making system

    CN122264339A

  • A modular construction dynamic hoisting scheduling decision system

    CN122264339B