Unmanned aerial vehicle cluster cooperation method and device based on local interpretability reinforcement learning

By introducing local interpretability reinforcement learning methods in the coordinated control of drone clusters, using local joint action value functions and reliability value computer system, the problem of insufficient utilization of drone cluster collaborative control efficiency and local interactive information in the prior art is solved, and more efficient drone cluster collaboration and automated task execution is achieved.

CN119960468APending Publication Date: 2025-05-09SHANXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411952859.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing collaborative control method for drone clusters based on multi-agent reinforcement learning has insufficient overall collaboration efficiency and the utilization of local team interactive information, resulting in limited intelligence and autonomous operation capabilities of drone clusters.

Method used

Using a method based on local interpretability reinforcement learning, the local joint action value function and reliability value computer system are introduced to enhance the local interactive information utilization ability of the drone group and improve the calculation and training process of the global joint value function.

Benefits of technology

It improves the overall collaboration capability of the drone cluster and the efficiency of automated execution of tasks, enhances the collaboration capability of a single drone with the surrounding neighbors drones, provides an explanation of the drone behavior, and guides the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960468A_ABST
    Figure CN119960468A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle control, in particular to an unmanned aerial vehicle cluster cooperation method and device based on local interpretability reinforcement learning. Comprising the following steps that individual local observation is obtained according to a current environment state, combined observation is formed, and an unmanned aerial vehicle cluster forms a combined action based on a current behavior strategy; collecting neighbor information in real time to form an information tuple and putting the information tuple into an experience playback pool; when the data volume collected in the experience playback pool reaches a data collection threshold value, collecting a part of data from the experience playback pool, calculating a global joint value function, then calculating a target global joint value function by using a Double DQN method, and training and updating hybrid network and behavior strategy network parameters by using a TD-Loss method; and after the single training is finished, the latest information tuple data is repeatedly collected for training, and the training is finished until the number of times of the single training reaches the total training threshold, and the overall cooperation capability and the automatic task execution efficiency of the unmanned aerial vehicle group can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drone control technology, and in particular to a drone cluster collaboration method and device based on local interpretable reinforcement learning. Background Art

[0002] Nowadays, drone swarms are increasingly used and play an increasingly important role in agriculture, military affairs, and communications. However, in the process of executing tasks, the scheduling of drone swarms still faces a series of challenges, such as changes in complex environments, dynamic changes in tasks, and handling of emergencies, which prompt drone swarms to be more intelligent and have stronger autonomous action capabilities in order to adapt to the changing needs of tasks and environments. The development of existing drone cooperative control technology can be divided into three stages: distributed collaboration, swarm intelligence collaboration, and multi-agent collaboration. Distributed collaboration requires setting goals and methods before executing tasks and finding the optimal method in advance. It cannot adapt to changes in the environment and goals during task execution, so it is not intelligent. Swarm intelligence collaboration is based on the behavior of biological groups, such as ant colonies, bee colonies, etc., to simulate the mutual influence and cooperation between biological groups to solve optimization problems, so that drone groups have certain intelligence. Classical swarm intelligence algorithms such as particle swarm optimization PSO (particle swarm optimization) and ant colony optimization ACO (ant colony optimization) both use heuristic information search optimization problems. Swarm intelligence collaboration enhances the distributed collaborative agent, but is limited by specific optimization algorithms and its intelligence is still limited. Multi-agent collaboration is an important direction for the future development of UAV collaboration technology. With the breakthrough progress of deep reinforcement learning, research on multi-agent reinforcement learning is gaining more and more attention. UAV collaborative control based on multi-agent reinforcement learning can autonomously make optimal actions according to the state of the environment, has stronger adaptability to the dynamic changes of the environment and tasks, and makes the UAV swarm more intelligent. UAV collaborative control methods based on multi-agent reinforcement learning can be divided into two categories: value-based methods and policy-based methods. These methods make UAV swarms more efficient in completing tasks than the first two methods.

[0003] Existing collaborative control methods based on multi-agent reinforcement learning, such as VDN, model the global joint value function as the sum of individual value functions; QMIX designs a neural network to ensure the monotonicity between individual value functions and global joint value functions; Qtran achieves the full representation of the global joint value function, but its performance is affected by the complexity of calculation; Qplex uses a dual adversarial network to achieve the decomposition of the joint value function, etc. However, when these algorithms are installed on drone clusters, they are trained from the dimension of a single drone, lacking the use of local team interaction information of the drone group, resulting in low overall collaboration efficiency of the drone group. At the same time, the local team behavior of the drone group cannot be explained during the training process, and there is a lack of accurate guidance for drone training. Summary of the invention

[0004] In order to solve the technical problems in the prior art of complex drone cluster control algorithms and low overall collaboration efficiency, the present invention proposes a drone cluster collaboration method and device based on local interpretable reinforcement learning to improve the overall collaboration capability of the drone cluster.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a UAV cluster collaboration method based on local interpretable reinforcement learning, comprising the following steps:

[0006] Step 1: Each single UAV obtains individual local observations based on the current environmental state, and the UAV cluster forms a joint observation based on the local observation information of each UAV. The UAV cluster forms a joint action based on the current behavior strategy;

[0007] Step 2: At each time step, each drone collects neighbor information in real time, forms an information tuple and puts it into the experience playback pool. The information tuple includes the current environment state, joint observation, adjacency matrix, joint action, joint reward, and the next environment state;

[0008] Step 3: When the amount of data collected in the experience replay pool reaches the data collection threshold, collect a portion of data from the experience replay pool, calculate the global joint value function, and then use the Double DQN method to calculate the target global joint value function, and use the TD-Loss method for training to update the parameters of the hybrid network and the behavior strategy network;

[0009] The calculation formula of the global joint value function is:

[0010]

[0011] Among them, Q tot (s,u) represents the global joint value function, s represents the environment state, u represents the joint action, β i represents the confidence value of a single UAV i, Indicates the drone team i The reliability value, Q i (τ i ,u i ) represents the state action value function of a single UAV i, Indicates the drone team i The local joint action value function, τ i ,u i They represent the observation history and actions of a single UAV i respectively; and Respectively represent team e i The joint observation history and joint actions of n are the number of drones; the drone team e is i is composed of drone i and its neighbors;

[0012] Step 4: After a single training is completed, repeat steps 1 and 2 to collect the latest information tuple data and put it into the experience replay pool to replace the old data, and then repeat step 3 to continue single training. Repeat the above steps to update the parameters of the hybrid network and the behavior strategy network, and update the parameters of the target behavior strategy network and the target hybrid network every certain number of training times until the number of single training times reaches the total training threshold, and then end the training;

[0013] Step 5: Each drone takes action according to the trained behavior strategy to complete drone cluster collaboration.

[0014] In step 1, the joint observation formed is:

[0015]

[0016] Among them, o represents joint observation, represents the local observation information of drones 1, 2, ..., n obtained by sensors on drones;

[0017] The method for drone swarms to form joint actions based on the current behavior strategy is:

[0018]

[0019] Among them, u represents the joint action formed at time t, u i Indicates the action of drone i, They represent the actions of drones 1, 2, …, n at time t respectively; represents the action observation history of drone i, Γ represents the action observation history space, U represents the action space, Indicates the current behavior strategy. Indicates the time step t, in the action observation history τ i Under the condition that drone i takes action u iprobability.

[0020] In step 2, the expression of the information tuple formed is:

[0021] G=(s t ,o t ,A t ,u t ,r,s t+1 );

[0022] Among them, G represents the information tuple, s t represents the environmental state at time t; o t A represents the joint observation at time t, which is formed by the local observation information of all UAVs; t represents the adjacency matrix at time t, which is formed by the neighbor relationships of all drones; u t represents the joint action at time t, which is formed by the actions taken by all drones according to the current behavior strategy; r represents the joint reward; s t+1 Represents the environmental state at time t+1.

[0023] In step 3, the calculation method of the local joint action value function of each team is as follows:

[0024] The action value function matrix Q is composed of the individual action value functions of all drones. t =R n×n , each row element of the action value function matrix is ​​the action value of all drones, and the row elements are copied to expand it into an n×n shape, and the adjacency matrix A in the information tuple is t And the action value function matrix Q t Perform the Hadamard product operation as follows:

[0025] Q tmp =A t ⊙Q t ,

[0026] Where Q tmp represents a temporary value, ⊙ represents the Hadamard product operation, and the temporary value Q tmp As input, the local joint action value vector Q is obtained through a three-layer convolutional network. e , from the local joint action value vector Q e The local joint action function of each team is extracted.

[0027] In step 3, the calculation formula of the confidence value of a single drone i is:

[0028]

[0029] Where K represents the number of Monte Carlo sampling. represents the coalition C of drone i in the jth samplingj The marginal contribution value in ;

[0030] Drone Team i The reliability value calculation formula is:

[0031]

[0032] in, represents the alliance C sampled at the jth time j Middle Team i The marginal contribution value of .

[0033] In step 3, the calculation formula of the target global joint value function is:

[0034] y=r+γQ′ tot (τ′,argmax u′ Q tot (τ′,u′)),

[0035] Where y represents the target global joint value function, r represents the joint reward, γ represents the discount factor, and Q′ tot (τ′,argmax u′ Q tot (τ′,u′)) represents the joint observation history and joint action of the next step, τ′ and τargmax respectively u′ Q tot (τ′,u′), the next step global joint value function is calculated in the target hybrid network, τ′ and u′ represent the next step joint observation history and joint action respectively, Q tot (τ′,u′) represents the global joint value function for calculating the next step in the hybrid network, argmax u′ Q tot (τ′,u′) means let Q tot The value of u′ that gives (τ′,u′) the maximum value.

[0036] In step 3, the parameter update formula of the hybrid network and the behavior strategy network is:

[0037]

[0038] Among them, L TD (θ ω ) represents the TD-Loss value, Q tot (τ,u) represents the simulated global joint value function, y represents the target global joint value function, represents the square of L2 norm;

[0039] The parameter update formula of the target behavior strategy network and the target hybrid network is:

[0040]

[0041] in, and Respectively represent the updated target behavior strategy network parameters and target hybrid network parameters, and denote the target behavior strategy network parameters and the target hybrid network parameters, θ p and θ w They represent the behavior strategy network parameters and the hybrid network parameters respectively, and λ represents the learning rate.

[0042] In step 4, when the latest information tuple data is collected and put into the experience replay pool to replace the old data, the total amount of data in the experience replay pool is kept unchanged.

[0043] In addition, the present invention also provides a UAV cluster collaboration device based on local interpretable reinforcement learning, comprising:

[0044] Perception module: includes a combination of sensors to obtain the current state of the drone's environment and its own location information;

[0045] Centralized training module: used to execute steps 1 to 4 of claim 1 according to the environmental state and its own position information obtained by the sensor combination;

[0046] Execution module: For specific environmental conditions, the drone takes actions based on the behavioral strategies learned in the centralized training module.

[0047] The sensor combination includes an infrared sensor, a laser sensor, and a visual sensor.

[0048] During the movement collaboration process of a drone swarm, a single drone often regards itself as an independent individual, and the probability of collaboration with the drones around it is low, resulting in low overall collaboration efficiency of the drone swarm, which affects the automated execution of tasks by the drone swarm. In order to improve the overall collaboration efficiency of the drone swarm, the present invention draws on the behavior of local collaboration in human society, proposes a drone cluster collaboration method based on local interpretable reinforcement learning, and designs a local joint action value function, the purpose of which is to incorporate a local interaction idea into the overall collaboration of the drone swarm, to make up for the shortcomings of insufficient utilization of local interaction information in the global joint value function, to promote a single drone to improve its collaboration ability with surrounding neighboring drones, and to improve the efficiency of the drone swarm's automated execution of tasks. Specifically, the present invention includes at least the following improvements:

[0049] 1. The present invention provides a representation of the joint value function. In addition to the state value function of a single drone, the expression also includes a local joint value function The local joint value function is used to represent the cooperative relationship between a single UAV i and its neighboring UAVs. The local interactive information is used to improve the local cooperative ability of a single UAV, thereby improving the overall cooperative ability of the UAV swarm.

[0050] 2. In addition, the present invention utilizes the adjacency matrix A t To generate the local joint value function, the neighbors of a single drone can be dynamically adjusted in real time. The position parameters of the drone group are constantly changing during flight, and the adjacency matrix can enable a single drone to adaptively adjust the scope of local collaboration to achieve the effect of automated learning.

[0051] 3. The present invention utilizes individual reliability value and team credibility Explanations are provided for the behaviors of single UAVs and local UAV teams respectively. Explanatory behaviors can provide precise guidance for UAV training and inspire UAVs to generate targeted behavioral strategies to cope with changes in the current environmental state during training.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] The present invention proposes a drone cluster collaboration method and device based on local interpretable reinforcement learning, which adds a local joint action value function to the global joint value function, and realizes the update of the behavior strategy network through the global joint value function. It can integrate local interaction information into the overall collaboration of the drone group, make up for the shortcoming of insufficient utilization of local interaction information in the global joint value function, promote a single drone to improve its collaboration ability with surrounding neighboring drones, and then improve the overall collaboration ability of the drone group and the efficiency of the drone group's automated task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A schematic diagram of a network structure adopted by a drone cluster collaboration method based on local interpretable reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0056] Embodiment 1

[0057] Embodiment 1 of the present invention provides a drone cluster collaboration method based on local interpretable reinforcement learning, comprising the following steps:

[0058] Step 1: Each single UAV obtains individual local observations based on the current environmental state, and the UAV cluster forms a joint observation based on the local observation information of each UAV. The UAV cluster forms a joint action based on the current behavior strategy.

[0059] In step 1, the joint observation formed is:

[0060]

[0061] Among them, o represents joint observation, represents the local observation information of drones 1, 2, ..., n obtained by sensors on drones;

[0062] The method for drone swarms to form joint actions based on the current behavior strategy is:

[0063]

[0064] Among them, u represents the joint action formed at time t, u i Indicates the action of drone i, They represent the actions of drones 1, 2, …, n at time t respectively; represents the action observation history of drone i, Γ represents the action observation history space, U represents the action space, Indicates the current behavior strategy. Indicates the time step t, in the action observation history τ i Under the condition that drone i takes action u i probability.

[0065] (1) Environmental state: During simulation, the simulation environment consists of a drone cluster and a target point. The drone positions and target points are randomly initialized. The environmental state at time t is given by s t Indicates that the global environment status information is provided by the simulation environment.

[0066] (2) Joint observation: The local observation information of all UAVs i at time t is obtained through various sensors on the UAVs. Joint Observation in, represents the local observation information of UAVs 1, 2, …, n at time t, i represents the UAV ID, and n represents the number of UAVs.

[0067] (3) Drone behavior strategy: At time t, drone i follows the current behavior strategy Make a move Forming joint actions in represents the action of drones 1, 2, ..., n at time t, τ irepresents the action observation history of drone i, Γ represents the action observation history space, and U represents the action space.

[0068] Step 2: At each time step, each drone collects neighbor information in real time to form an information tuple and put it into the experience replay pool. The information tuple includes the current environment state, joint observation, adjacency matrix, joint action, joint reward, and the next environment state.

[0069] In step 2, the expression of the information tuple formed is:

[0070] G=(s t ,o t ,A t ,u t ,r,s t+1 ); (4)

[0071] Among them, G represents the information tuple, s t represents the environmental state at time t; o t A represents the joint observation at time t, which is formed by the local observation information of all UAVs; t represents the adjacency matrix at time t, which is formed by the neighbor relationships of all drones; u t represents the joint action at time t, which is formed by the actions taken by all drones according to the current behavior strategy; r represents the joint reward; s t+1 Represents the environmental state at time t+1.

[0072] In this embodiment, at time t, each drone obtains neighbor information of itself and other drones through sensors, and stipulates that other drones within the detection range of the drone sensor form a neighbor relationship with itself, which is expressed by the following formula:

[0073]

[0074] Among them A ij Indicates the adjacency relationship between drones, B i Represents the neighbor set of drone i. By collecting the neighbor relationships of all drones, the adjacency matrix A at time t is formed t .

[0075] In this embodiment, the reward value is returned by the simulation environment at each time step t, and the calculation formula is:

[0076] r=S×U n →R, (6)

[0077] Among them, r represents the joint reward, r is related to the environment state s and the joint action u, S represents the state space, U nrepresents the action space, and R represents the reward value space. The joint reward is a single value and is shared by all drones.

[0078] After obtaining the above information tuple, the information tuple G at time t = (s t ,o t ,A t ,u t ,r,s t+1 ) into the experience replay pool.

[0079] Step 3: When the amount of data collected in the experience replay pool reaches the data collection threshold, collect some data from the experience replay pool, calculate the global joint value function, and then use the Double DQN method to calculate the target global joint value function, and use the TD-Loss method for training to update the parameters of the hybrid network and the behavior strategy network. Figure 1 FIG. 1 is a schematic diagram of the network structure used in this embodiment.

[0080] In this embodiment, local interaction information is introduced in the process of generating the global joint value function to provide an explanation for the local cooperative behavior of drones. The calculation formula of the global joint value function is:

[0081]

[0082] Among them, Q tot (s,u) represents the global joint value function, s represents the environment state, u represents the joint action, β i represents the confidence value of a single UAV i, Indicates the drone team i The reliability value, Q i (τ i ,u i ) represents the state action value function of a single UAV i, Indicates the drone team i The local joint action value function, τ i ,u i They represent the observation history and actions of a single UAV i respectively; and Respectively represent team e i The joint observation history and joint actions of n are the number of drones; the drone team e is i is composed of drone i and its neighbors;

[0083] In this embodiment, the calculation method of the local joint action value function of each team is as follows:

[0084] The action value function matrix Q is composed of the individual action value functions of all drones. t =R n×n, each row element of the action value function matrix is ​​the action value of all drones, and the row elements are copied to expand it into an n×n shape, and the adjacency matrix A in the information tuple is t And the action value function matrix Q t Perform the Hadamard product operation as follows:

[0085] Q tmp =A t ⊙Q t ; (8)

[0086] Where Q tmp represents a temporary value, ⊙ represents the Hadamard product operation, and the temporary value Q tmp As input, the local joint action value vector Q is obtained through a three-layer convolutional network. e , from the local joint action value vector Q e The local joint action function of each team is extracted.

[0087] Specifically, in step 3, the Shapley value theory is used to calculate the individual credibility value for each drone, specifically:

[0088]

[0089] Where C represents the number of drones in a coalition of any number of drones but does not include i, V(C) represents the joint reward value of the drone coalition C, and V(C∪{i}) represents the joint reward value of the coalition C and drone i. represents the marginal contribution value of drone i, and N is the total number of drones in the group; represents the Shapley value of drone i in the drone group N, for Approximate value of .

[0090]

[0091] Where K represents the number of Monte Carlo sampling. represents the coalition C of drone i in the jth sampling j The marginal contribution value in ;

[0092] Specifically, in step 3, the Shapley value theory is used to calculate the credibility value for each local drone team, specifically:

[0093]

[0094] Q C (τ C ,u C )=Q(τ,u)―Q N\C (τN\C ,u N\C ); (13)

[0095]

[0096] in Indicates team i The marginal contribution of Represents alliance C and team e i The marginal contribution of the combination of , Q(τ,u) represents the global joint value function of the drone group, Exclude alliance C and team e from N i The joint value function after the combination of C (τ C ,u C ) represents the marginal contribution of alliance C, and the Shapley value of team i is further calculated as follows:

[0097]

[0098] Using Monte Carlo sampling The reliability value of team i is approximately obtained as follows:

[0099]

[0100] in, Indicates the alliance C sampled at the jth time j Middle Team i The marginal contribution value of .

[0101] In step 3, the calculation formula of the target global joint value function is:

[0102] y=r+γQ′ tot (τ′,argmax u′ Q tot (τ′,u′)); (17)

[0103] Among them, y represents the target global joint value function, r represents the joint reward, γ represents the discount factor, which is generally 0.99, and Q′ tot (τ′,argmax u′ Q tot (τ′,u′)) represents the joint observation history and joint action of the next step, τ′ and argmax respectively. u′ Q tot (τ′,u′), the next step global joint value function is calculated in the target hybrid network, τ′ and u′ represent the next step joint observation history and joint action respectively, Q tot(τ′,u′) represents the global joint value function for calculating the next step in the hybrid network, argmax u′ Q tot (τ′,u′) means let Q tot The value of u′ that gives (τ′,u′) the maximum value.

[0104] Specifically, in step 3, the TD-Loss method is used for training to update the parameters of the hybrid network and the behavior strategy network. The parameter update formula of the hybrid network and the behavior strategy network is:

[0105]

[0106] Among them, L TD (θ ω ) represents the TD-Loss value, which is simulated by the global joint value function Q tot The squared L2 norm of the difference between (τ,u) and the target global joint value function y, θ ω represents the hybrid network parameters;

[0107] Step 4: After a single training is completed, repeat steps 1 and 2 to collect the latest information tuple data and put it into the experience replay pool to replace the old data, and then repeat step 3 to continue single training. Repeat the above steps to update the parameters of the hybrid network and the behavior strategy network, and update the parameters of the target behavior strategy network and the target hybrid network every certain number of training times until the number of single training times reaches the total training threshold, and then end the training.

[0108] In this embodiment, the update step threshold of the target behavior strategy network and the target hybrid network is set. When the training step reaches the update step threshold, the parameters of the target behavior strategy network and the target hybrid network are updated as shown below:

[0109]

[0110] in, and Respectively represent the updated target behavior strategy network parameters and target hybrid network parameters, and denote the target behavior strategy network parameters and the target hybrid network parameters, θ p and θ w They represent the behavior strategy network parameters and the hybrid network parameters respectively, λ represents the learning rate, λ∈(0,1).

[0111] Specifically, in this embodiment, when repeating steps 1 and 2 to collect the latest data and put it into the experience replay pool to replace the old data, the total amount of data in the experience replay pool remains unchanged. Then continue to sample part of the data for training. When the number of training steps reaches the total threshold t max, the training process ends and the drone cluster performs distributed execution.

[0112] Step 5: Each drone takes action according to the trained behavior strategy to complete drone cluster collaboration.

[0113] In this embodiment, the entire training process is centralized training distributed execution. The total threshold of training steps is set. According to the method of sampling-training-updating experience replay pool data-sampling-training, when the training steps reach the total threshold, the drone swarm performs distributed execution.

[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by means of software plus necessary general hardware platforms. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of software products, such as systems. The system can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in the embodiments of the present invention.

[0115] Embodiment 2

[0116] Embodiment 2 of the present invention provides a drone cluster collaboration device based on local interpretable reinforcement learning, including:

[0117] Perception module: includes a combination of sensors to obtain the current state of the drone's environment and its own location information;

[0118] Centralized training module: used to execute steps 1 to 4 of the first embodiment according to the environmental state and its own position information obtained by the sensor combination;

[0119] Execution module: For specific environmental conditions, the drone takes actions based on the behavioral strategies learned in the centralized training module.

[0120] Specifically, the sensor combination includes an infrared sensor, a laser sensor, and a visual sensor.

[0121] In this embodiment, the perception module collects the current environmental status information and its own position information through various external sensors of the drone, processes the collected information using the onboard data processing module, and then inputs it into the centralized training module. The centralized training module includes a neural network module for executing the training algorithm, which trains the input data using the algorithm of the present invention to obtain actions and transmits them to the execution module. The execution module parses the action into the speed and flight angle of the drone rotor motor and executes it, obtains a reward, and transfers the environment to the next state. Determine whether the target point has been reached. If not, repeat the above process until the target point is reached; if the target point is reached or the execution step threshold is reached, stop execution, and reinitialize the environmental state and drone state. Training ends until the total training step threshold is reached.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A UAV swarm collaboration method based on local interpretable reinforcement learning, characterized in that: The following steps are involved: Step 1: Each single UAV obtains individual local observations based on the current environmental state, and the UAV cluster forms a joint observation based on the local observation information of each UAV. The UAV cluster forms a joint action based on the current behavior strategy; Step 2: At each time step, each drone collects neighbor information in real time, forms an information tuple and puts it into the experience playback pool. The information tuple includes the current environment state, joint observation, adjacency matrix, joint action, joint reward, and the next environment state; Step 3: When the amount of data collected in the experience replay pool reaches the data collection threshold, collect a portion of data from the experience replay pool, calculate the global joint value function, and then use the Double DQN method to calculate the target global joint value function, and use the TD-Loss method for training to update the parameters of the hybrid network and the behavior strategy network; The calculation formula of the global joint value function is: Among them, Q tot (s,u) represents the global joint value function, s represents the environment state, u represents the joint action, β i represents the confidence value of a single UAV i, Indicates the drone team i The reliability value, Q i (τ i ,u i ) represents the state action value function of a single UAV i, Indicates the drone team i The local joint action value function, τ i ,u i They represent the observation history and actions of a single UAV i respectively; and Respectively represent team e i The joint observation history and joint actions of n are the number of drones; the drone team e is i is composed of drone i and its neighbors; Step 4: After a single training is completed, repeat steps 1 and 2 to collect the latest information tuple data and put it into the experience replay pool to replace the old data, and then repeat step 3 to continue single training. Repeat the above steps to update the parameters of the hybrid network and the behavior strategy network, and update the parameters of the target behavior strategy network and the target hybrid network every certain number of training times until the number of single training times reaches the total training threshold, and then end the training; Step 5: Each drone takes action according to the trained behavior strategy to complete drone cluster collaboration.

2. The UAV cluster collaboration method based on local interpretable reinforcement learning according to claim 1 is characterized in that: In step 1, the joint observation formed is: Among them, o represents joint observation, represents the local observation information of drones 1, 2, ..., n obtained by sensors on drones; The method for drone swarms to form joint actions based on the current behavior strategy is: Among them, u represents the joint action formed at time t, u i Indicates the action of drone i, They represent the actions of drones 1, 2, …, n at time t respectively; represents the action observation history of drone i, Γ represents the action observation history space, U represents the action space, Indicates the current behavior strategy. Indicates the time step t, in the action observation history τ i Under the condition that drone i takes action u i probability.

3. The UAV cluster collaboration method based on local interpretable reinforcement learning according to claim 1 is characterized in that: In step 2, the expression of the information tuple formed is: G=(s t ,o t ,A t ,u t ,r,s t+1 ); Among them, G represents the information tuple, s t represents the environmental state at time t; o t A represents the joint observation at time t, which is formed by the local observation information of all UAVs; t represents the adjacency matrix at time t, which is formed by the neighbor relationships of all drones; u t represents the joint action at time t, which is formed by the actions taken by all drones according to the current behavior strategy; r represents the joint reward; s t+1 Represents the environmental state at time t+1.

4. The UAV cluster collaboration method based on local interpretable reinforcement learning according to claim 1, characterized in that: In step 3, the calculation method of the local joint action value function of each team is as follows: The action value function matrix Q is composed of the individual action value functions of all drones. t =R n×n , each row element of the action value function matrix is ​​the action value of all drones, and the row elements are copied to expand it into an n×n shape, and the adjacency matrix A in the information tuple is t And the action value function matrix Q t Perform the Hadamard product operation as follows: Q tmp =A t ⊙Q t , Where Q tmp represents a temporary value, ⊙ represents the Hadamard product operation, and the temporary value Q tmp As input, the local joint action value vector Q is obtained through a three-layer convolutional network. e , from the local joint action value vector Q e The local joint action function of each team is extracted.

5. The UAV swarm collaboration method based on local interpretable reinforcement learning according to claim 1, characterized in that: In step 3, the calculation formula of the confidence value of a single drone i is: Where K represents the number of Monte Carlo sampling. represents the coalition C of drone i in the jth sampling j The marginal contribution value in ; Drone Team i The reliability value calculation formula is: in, Indicates the alliance C sampled at the jth time j Middle Team i The marginal contribution value of .

6. The UAV cluster collaboration method based on local interpretable reinforcement learning according to claim 1 is characterized in that: In step 3, the calculation formula of the target global joint value function is: y=r+γQ′ tot (τ′,argmax u′ Q tot (τ′,u′)), Where y represents the target global joint value function, r represents the joint reward, γ represents the discount factor, and Q′ tot (τ′,argmax u′ Q tot (τ′,u′)) represents the joint observation history and joint action of the next step, τ′ and argmax respectively. u′ Q tot (τ′,u′), the next step global joint value function is calculated in the target hybrid network, τ′ and u′ represent the next step joint observation history and joint action respectively, Q tot (τ′,u′) represents the global joint value function for calculating the next step in the hybrid network, argmax u′ Q tot (τ′,u′) means let Q tot The value of u′ that gives (τ′,u′) the maximum value.

7. The UAV cluster collaboration method based on local interpretable reinforcement learning according to claim 1 is characterized in that: In step 3, the parameter update formula of the hybrid network and the behavior strategy network is: Among them, L TD (θ ω ) represents the TD-Loss value, Q tot (τ,u) represents the simulated global joint value function, y represents the target global joint value function, represents the square of L2 norm; The parameter update formula of the target behavior strategy network and the target hybrid network is: in, and Respectively represent the updated target behavior strategy network parameters and target hybrid network parameters, and denote the target behavior strategy network parameters and the target hybrid network parameters, θ p and θ w They represent the behavior strategy network parameters and the hybrid network parameters respectively, and λ represents the learning rate.

8. The UAV swarm collaboration method based on local interpretable reinforcement learning according to claim 1, characterized in that: In step 4, when the latest information tuple data is collected and put into the experience replay pool to replace the old data, the total amount of data in the experience replay pool is kept unchanged.

9. A drone swarm collaboration device based on local interpretable reinforcement learning, characterized in that: include: Perception module: includes a combination of sensors to obtain the current environmental status of the drone and its own location information; Centralized training module: used to execute steps 1 to 4 of claim 1 according to the environmental state and its own position information obtained by the sensor combination; Execution module: For specific environmental conditions, the drone takes actions based on the behavioral strategies learned in the centralized training module.

10. The UAV cluster collaboration device based on local interpretable reinforcement learning according to claim 9, characterized in that: The sensor combination includes an infrared sensor, a laser sensor, and a visual sensor.

Citation Information

Cited By

  • Low-altitude economic trusted data processing method and storage medium

    CN120430532A