Unmanned aerial vehicle cluster consistency decision-making system and method based on deep reinforcement learning
Through the drone cluster consistency decision-making system based on deep reinforcement learning, the problems of inconsistent decision-making and low coordination efficiency in the drone cluster collaborative tasks are solved, and efficient collaborative decision-making and good scalability in complex environments are achieved.
Patent Information
- Application Number
- CN202510048398.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-27
AI Technical Summary
In the collaborative tasks of existing drone clusters, decision-making is inconsistent and synergistic efficiency is low, especially in complex dynamic environments.
The drone cluster consistency decision-making system based on deep reinforcement learning is adopted, and the overall decision-making consistency of the drone cluster is achieved through modules such as experience playback pool, observation conversion module, consistency decision-making adjustment module, multi-drone hybrid network and network optimizer.
Rapidly form group consensus in complex and changeable environments, achieve efficient collaborative decision-making, improve the overall performance and adaptability of the drone cluster, and have good scalability.
Smart Images

Figure CN120044964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a UAV swarm consensus decision-making system and method, in particular to a UAV swarm consensus decision-making system and method based on deep reinforcement learning. Background Art
[0002] The information provided in this section is only background information related to the present disclosure and is not necessarily prior art.
[0003] With the rapid development of technology, unmanned systems are increasingly widely used in various fields, and the improvement of intelligent decision-making ability has become the key focus. Traditional unmanned systems mostly rely on preset programs and manual remote control, which often cannot face the ever-changing complex environment. Therefore, there is an increasing urgent need for intelligent systems with autonomous adaptation ability and flexible decision-making mechanisms to cope with the changing situation and improve the task success rate.
[0004] Early intelligent attempts mainly focused on introducing simple heuristic algorithms and basic machine learning models. Although these methods perform well in specific scenarios, their limitations become increasingly prominent when facing highly uncertain and dynamically changing environments. They are difficult to effectively handle complex multi-variable decision-making problems and cannot cope with new situations that have not been encountered before, resulting in insufficient adaptability of the system in practical applications.
[0005] With the progress of artificial intelligence technology, single-agent reinforcement learning algorithms have been introduced into unmanned systems, aiming to optimize decision-making strategies through continuous interaction with the environment. Such methods have achieved certain results in improving the decision-making ability of a single agent and can cope with environmental uncertainties to a certain extent. However, when applied to multi-UAV cooperative tasks, single-agent methods are difficult to fully consider the integrity and coordination of group behavior, often leading to problems such as decision-making conflicts, uneven resource allocation, and low overall efficiency.
[0006] In recent years, the research on multi-agent systems has made remarkable progress in the field of artificial intelligence, providing new solutions to complex cooperative decision-making problems. However, existing multi-agent methods mainly focus on improving the performance and decision-making ability of individual agents and lack consideration of the consistency of group behavior. Although this individual-oriented optimization strategy enhances the autonomy of a single agent, it exposes obvious defects in scenarios that require high coordination such as UAV swarm cooperative operations. In addition, current multi-agent systems often face difficulties in achieving real-time decision-making due to the sharp increase in computational complexity when dealing with large-scale swarms. How to provide an efficient cooperative decision-making mechanism for large-scale UAV swarms under limited computational resources and time constraints has become a thorny challenge.
[0007] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] Object of the Invention: The technical problem to be solved by the present invention is to provide a drone swarm consensus decision-making system and method based on deep reinforcement learning in view of the deficiencies of the prior art.
[0009] To solve the above technical problem, the present invention discloses a drone swarm consensus decision-making system and method based on deep reinforcement learning, including:
[0010] An experience replay pool, an observation conversion module, a consensus decision adjustment module, a multi-drone hybrid network, a network optimizer, and a drone task simulation platform; wherein,
[0011] The experience replay pool is used to store the observation information, actions, states, and rewards generated by the drone swarm during the task execution.
[0012] The observation conversion module converts the observation information of the current environment where the drone is located into collaborative observation information input to the consensus decision adjustment module.
[0013] The consensus decision adjustment module is used to coordinate the decision-making behaviors of different drones after collaborative observation.
[0014] The multi-drone hybrid network combines the local information of each drone, jointly optimizes the decisions of each drone, and outputs the overall decision of the drone swarm.
[0015] The network optimizer optimizes and updates the parameters of the multi-drone hybrid network according to the feedback information obtained from the drone task simulation platform.
[0016] The drone task simulation platform is used to provide a training and testing environment, simulate the actual task scenarios of the drone swarm, and provide state feedback to the multi-drone hybrid network and the network optimizer.
[0017] Further, the consensus decision adjustment module includes:
[0018] A decision consensus evaluation network corresponding to the number of drones, which calculates the local decision values of each drone respectively and coordinates the decisions among drones through a consensus constraint mechanism.
[0019] Further, the multi-drone hybrid network includes:
[0020] A drone value network corresponding to the number of drones and one overall value aggregation network; wherein,
[0021] The UAV value network generates corresponding local action values according to the local observation information of the current UAV;
[0022] The overall value aggregation network calculates the total value of the UAV swarm based on the local action values of each UAV, and is used to guide the overall decision-making of the UAV swarm.
[0023] Furthermore, the UAV value network generates corresponding local action values according to the local observation information of the current UAV, including:
[0024] The observation information of UAV i is input into the corresponding UAV value network Q (i) to obtain the value of each action a k under the current observation of this UAV, which is expressed as follows:
[0025]
[0026] where A represents the preset action space, and a k is any one of the actions;
[0027] According to the preset greedy strategy, a random selection decision is made for exploration with a probability of ε, and the action with the maximum value is selected with a probability of 1 - ε, which is expressed as follows:
[0028]
[0029] Finally, the action selected by UAV i at the current moment t is obtained
[0030] where the UAV value network Q (i) adopts a four-layer feedforward neural network structure, namely: an input layer, a 64-dimensional hidden layer, a 32-dimensional hidden layer, and a 6-dimensional output layer.
[0031] Furthermore, the UAV mission simulation platform provides state feedback, including:
[0032] According to the action selected by UAV i at the current moment t the reward R is settled by comparing with the preset mission objective t .
[0033] Furthermore, the multi-UAV hybrid network jointly optimizes the decisions of each UAV, including:
[0034] The observation information in the experience replay pool is input into each UAV value network Q (i) to re-evaluate the value of the selected action of It is shown as follows:
[0035]
[0036] Evaluate the action values of all drones and the global state information S t Input them into the multi - drone hybrid network M to evaluate the overall action value of the drone swarm It is shown as follows:
[0037]
[0038] where I is the number of drones in the drone swarm;
[0039] Input the observation information at the next moment into each drone value network Q (i) and select the action with the maximum value It is shown as follows:
[0040]
[0041] Obtain the value evaluation at the next moment Input it and the global state information S at the next moment t+1 into the multi - drone hybrid network M to evaluate the overall action value of the drone swarm at the next moment, which is shown as follows:
[0042]
[0043] Among them, the multi - drone hybrid network M is composed of an encoder layer, a 4 - head attention layer, and two fully - connected layers. Among them, the encoder layer maps the input to a 128 - dimensional feature space, the 4 - head attention layer processes the interaction between drones, and the two fully - connected layers are 128 - dimensional and 64 - dimensional respectively, and output the estimated value of the final action value.
[0044] Furthermore, the consistency decision adjustment module coordinates the decision - making behaviors of different drones after collaborative observation, that is, inputs the collaborative observation information into the consistency evaluation network of each drone to obtain the action values of different drones under different collaborative observation information, and selects the corresponding action decisions, specifically including:
[0045] Set the collaborative observation information of any drone i and different drones j as Input it into the corresponding consistency evaluation network C of drone i (i) to obtain the action values of different drones under different collaborative observation information, which is shown as follows:
[0046]
[0047] Select the action decision with the greatest value under different collaborative observations It is expressed as follows:
[0048]
[0049] Obtain the action value It is expressed as follows:
[0050]
[0051] Among them, the consistency evaluation network C (i) Adopts a five-layer feedforward neural network structure, including an input layer, a 128-dimensional hidden layer, a 64-dimensional hidden layer, a 32-dimensional hidden layer, and a 6-dimensional output layer.
[0052] The present invention also proposes a method for drone swarm consensus decision-making based on deep reinforcement learning, which uses the aforementioned system to achieve drone swarm consensus decision-making, including the following steps:
[0053] Step 1, initialize the drone swarm in the drone mission simulation platform and configure the initial state of each drone; the drone mission simulation platform simulates the real mission scenario, sets the mission objectives, environmental conditions and the behavior patterns of the target drones of the other drones of the drone swarm;
[0054] Step 2, through the drone mission simulation platform, obtain the state information and the observation information of each drone in the current environment; among them, the state information includes: the position information, speed information, sensor data, the position of the other drones and the relevant environmental information of the drones;
[0055] Step 3, each drone evaluates the current local state by using the corresponding drone value network through the deep reinforcement learning method, outputs the local action value, and selects the corresponding action decision;
[0056] Step 4, in the drone mission simulation platform, the action decisions selected by each drone according to their respective observation information in Step 3 are sequentially executed by the drone swarm; the drone mission simulation platform records the execution results of each drone in real time, obtains the state information and the observation information of each drone at the next moment, settles the rewards according to the mission objectives, and stores the above information in the experience replay pool;
[0057] Step 5, check whether the data in the experience replay pool meets the preset conditions and is sufficient for sampling. If not, repeat Steps 1-4. Otherwise, batch collect the observation information, state information, actions and rewards generated by the drone swarm during the mission execution from the experience replay pool;
[0058] Step 6: Each UAV value network re-evaluates the value of the action selected by the UAV according to the observation information collected from the experience replay pool, and inputs it and the state information into the multi-UAV hybrid network to evaluate the overall action value of the UAV cluster;
[0059] Step 7: Perform observation conversion on the observation information of each UAV collected from the experience replay pool to obtain the collaborative observation information of different UAVs; wherein, the collaborative observation information includes the collaborative observation information between any two UAVs;
[0060] Step 8: Input the collaborative observation information in Step 7 into the decision consistency evaluation network of each UAV to obtain the action values of different UAVs under different collaborative observation information, and select the corresponding action decisions;
[0061] Step 9: Input the value of the action selected by each UAV in Step 6 and the action values obtained by each UAV in Step 8 into the consistency decision adjustment module to calculate the action non-consistency penalty;
[0062] Step 10: Design a loss function, that is, obtain a loss term using the temporal difference method, and minimize the loss function through the backpropagation method to update the system parameters;
[0063] Step 11: Determine whether the current policy has converged by setting the completion rate of consecutive multiple tasks in the scenario; if the policy has not converged, repeat Steps 1-10, otherwise stop training and updating to complete the consensus decision-making of the UAV cluster based on deep reinforcement learning.
[0064] Further, the data in the experience replay pool is sampled through a prioritized experience replay mechanism, that is, the state-action pair data in the saved data is sorted according to the error, and selected in order of priority.
[0065] Further, the consistency decision adjustment module transmits local state and action decision information between UAVs through a communication-based mechanism.
[0066] Beneficial effects:
[0067] The present invention introduces an effective collaborative consensus decision-making mechanism, which can quickly form a group consensus in a complex and changeable environment, achieve efficient collaborative decision-making, and at the same time maintain good scalability to meet the needs of large-scale clusters. This can not only improve the overall performance and adaptability of the UAV cluster, but also open up new possibilities for its wider application fields. Description of the Drawings
[0068] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0069] Figure 1 It is the architecture diagram of the UAV swarm consistency decision-making system of the present invention.
[0070] Figure 2 It is the flow chart of the UAV swarm consistency decision-making method of the present invention.
[0071] Figure 3 It is the schematic diagram of the UAV swarm collaborative task scenario of the present invention. Detailed implementation manners
[0072] In order to overcome the technical problems of inconsistent decision-making and low collaborative efficiency in the existing UAV swarm collaborative tasks, the present invention provides a UAV collaborative system based on deep reinforcement learning and its application method. In the prior art, the decision-making process of the UAV swarm is often limited by insufficient local information and lack of effective collaborative mechanisms, resulting in low task efficiency of the swarm in complex dynamic environments and difficulty in achieving the expected goals in highly adversarial tasks.
[0073] To solve the above problems, the present invention designs a UAV collaborative system integrating a multi-agent local network and a hybrid network by introducing a deep reinforcement learning algorithm, aiming to improve the decision-making consistency and collaborative task ability of the UAV swarm in dynamic environments. The core technology of the present invention lies in the design of a consistency decision adjustment module and a reward function based on group interests, guiding individual UAVs to maintain the consistency of collective actions when making autonomous decisions, and thus optimizing the tactical decisions of the entire swarm.
[0074] The method includes the following steps:
[0075] Step 1: Initialize the UAV swarm in the UAV task simulation platform and configure the initial state of each UAV. This platform simulates the real task scenario, sets the task objectives, environmental conditions and the target behavior patterns of the opponent UAVs, providing a simulation environment for training and testing the UAV swarm.
[0076] Step 2: Obtain the state information in the current simulation environment and the observation information of each UAV through the simulation platform. The state information includes the position information, speed information, sensor data, the position of the opponent UAV units and other relevant environmental information of the UAVs.
[0077] Step 3: Each UAV evaluates the current local state using the deep reinforcement learning method through its independent value network, outputs the local action value, and selects the corresponding action decision.
[0078] Step 4: The action decisions selected by each drone in step 3 based on their respective observation information are sequentially executed by the drone swarm in the simulation platform. The simulation platform records the execution results of each drone in real time to obtain the state information and the observation information of each drone at the next moment. At the same time, the reward is settled according to the task objective, and this information is stored in the experience replay pool.
[0079] Step 5: Check whether the data in the experience replay pool is sufficient for sampling. If not, repeat steps 1-4 to continue collecting replay data. After the data in the experience replay pool is sufficient, batch collect the observation, state, action, reward and other data generated by the drone swarm during the task execution from the experience replay pool.
[0080] Step 6: The value network of each drone re-evaluates the value of the selected action according to the observation information in the experience replay pool and inputs it into the multi-drone hybrid network to evaluate the overall action value of the drone swarm.
[0081] Step 7: Input the observation information of each drone collected in the experience replay pool into the observation conversion module to obtain the collaborative observation information of different drones. The collaborative observation information includes the integration of the observation information between any two drones.
[0082] Step 8: Input the collaborative observation information of each drone with different drones into the consistency evaluation network of each drone to obtain the action value of different drones under different collaborative observation information, and select the corresponding action decision.
[0083] Step 9: Input the value of the action decision selected by each drone according to its own observation information in step 6 and the value of the action decision selected by each drone according to the collaborative observation information from different drones in step 8 into the consistency evaluation module to calculate the action non-consistency penalty.
[0084] Step 10: Use the temporal difference algorithm to obtain the loss term, and minimize the loss function through the backpropagation algorithm to update the parameters of each network, so that the strategies of each drone gradually approach the optimal strategy, ensuring that the overall task efficiency of the drone swarm is continuously improved.
[0085] Step 11: Determine whether the current strategy has converged through the task completion rate of multiple consecutive scenarios. If the strategy has not converged, the system will repeat steps 1-10. After the strategy converges, stop training and updating, and save the models of the local networks of each drone to obtain the decision-making models of multiple drones applicable to the current scenario.
[0086] Example:
[0087] According to the appendix Figure 1As shown in the figure, this embodiment provides a method for UAV swarm consensus decision-making based on deep reinforcement learning. The system involved in this method mainly consists of multiple core modules such as a simulation platform, an observation conversion module, a UAV decision network, a consensus decision adjustment module, an experience replay pool, and a network optimizer.
[0088] The simulation platform module is used to construct and update a task environment similar to the real task scenario. In this simulation environment, each UAV in the UAV swarm obtains the state information in the current environment through the observation conversion module, including position information, speed information, the position of the target of other UAVs, etc. After conversion, these information form local observation data suitable for input into the multi-UAV hybrid network. In addition, the observation conversion module also integrates the observation information of each UAV to generate various collaborative observation information suitable for the cooperation of each UAV, which is used for collaborative decision-making consensus matching to constrain the decisions of UAVs.
[0089] The UAV decision network module uses a deep reinforcement learning algorithm to evaluate the input local observation data. Each UAV calculates the action value of its local state through an independent value network to generate a preliminary local action decision. The multi-UAV hybrid network module then evaluates the overall action value of the swarm based on the local action values of each UAV.
[0090] The consensus decision adjustment module is used to coordinate the action decisions among UAVs. By introducing a consensus constraint mechanism, the UAV swarm can maintain cooperation and consistency during the task execution process, thereby improving the task efficiency.
[0091] The experience replay pool module is responsible for storing information such as observation data, action data, and reward signals generated during the task execution for use by the network optimizer. Through the backpropagation algorithm, the network optimizer can continuously optimize the parameters of the multi-UAV hybrid network based on the training data extracted from the experience replay pool to ensure the continuous improvement of the overall strategy of the swarm.
[0092] This embodiment adopts a modular design to ensure the flexibility and efficient operation of the system in a complex task environment. Through this design, the system can continuously optimize the cooperation strategy of the UAV swarm in a changing scenario, effectively improving the overall efficiency and performance of task execution.
[0093] Combined with Figure 2 As shown in the figure, the adaptive multi-UAV reinforcement learning method for UAV swarm consensus decision-making in this embodiment is further described. The method specifically includes:
[0094] Step 1: Initialize the UAV swarm in the UAV task simulation platform and configure the initial state of each UAV. In this embodiment, a UAV swarm task simulation platform developed based on QT is used, such as Figure 3As shown in the figure, the task scenario is that 4 of our drones detect 5 opponent drones. The radar detection ranges of the drones on both sides are configured identically, and resources such as fuel are also configured identically. The actions that the drones can take include six basic actions: level flight, acceleration, deceleration, turning, climbing, and diving. That is, the action space A = {a 1 , a 2 , a 3 , a 4 , a 5 , a 6}.
[0095] Step 2: Through the simulation platform, obtain the state information S in the simulation environment at the current moment t t and the observation information of each drone. The observation information of each drone includes the position information (lat i , lon i , alt i ) of each drone itself, speed (v e , v n , v u ), fuel quantity k i , and information such as the radar detection list; the state information integrates the observation information of all drones and includes relevant information about the opponent drones.
[0096] Step 3: Input the observation information of each drone into its respective value network Q (i) to obtain the value of each action under the current observation of each drone According to the ε-greedy strategy, make a random selection decision for exploration with a probability of ε, and select the action with the maximum value with a probability of 1 - ε Finally, obtain the action selected by drone i at the current moment t
[0097] In this embodiment, the drone value network Q (i) adopts a four-layer feedforward neural network structure: an input layer, a 64-dimensional hidden layer, a 32-dimensional hidden layer, and a 6-dimensional output layer.
[0098] Step 4: The action decisions selected by each drone according to their respective observation information in Step 3 are sequentially executed by the drone swarm in the simulation platform. The simulation platform records the execution results of each drone in real time to obtain the state information S t+1 and the observation information of each drone At the same time, calculate the reward R according to the task objective t , and store this information in the experience replay pool.
[0099] In this embodiment, the rewards used by the simulation task platform consider the discovery target reward, the task completion reward, the detection penalty, the penalty for flying out of the task area, the additional fuel consumption penalty, and the task failure penalty. The reward function is as follows:
[0100]
[0101] Step 5: Check whether the data in the experience replay pool is sufficient for sampling. If not, repeat Steps 1-4 to continue collecting replay data. After the data in the experience replay pool is sufficient, batch-collect the data such as observations, states, actions, and rewards generated by the UAV cluster during task execution from the experience replay pool. The collected data is denoted as where i = 1, 2, 3, 4.
[0102] Step 6: Input the observation information in the experience replay pool into each UAV value network Q (i) to re-evaluate the value of the selected action and input the action value evaluations of each UAV and the global state information S t into the multi-UAV hybrid network M to evaluate the overall action value of the UAV cluster Input the observation information at the next moment into each UAV value network Q (i) and select the action with the maximum value Similarly, obtain the value evaluation Input it and the global state information S t+1 into the multi-UAV hybrid network M to evaluate the overall action value of the UAV cluster at the next moment
[0103] In this embodiment, the multi-UAV hybrid network M consists of an encoder layer, an attention layer, and a fully connected layer: the encoder maps the input to a 128-dimensional feature space, the 4-head attention layer processes the interaction between UAVs, and two fully connected layers (128-dimensional and 64-dimensional) output the final estimated value.
[0104] Step 7: Input the observation information of each UAV collected in the experience replay pool into the observation conversion module to obtain the cooperative observation information of different UAVs represents the combined observation information integrated by UAV i after obtaining the observation information of UAV j.
[0105] Step 8: Input the cooperative observation information of each UAV with different UAVs into the consistency evaluation network C of each UAV (i), obtain the action values of different drones under different cooperative observation information Select the action decision with the maximum value under different cooperative observations Action value
[0106] In this embodiment, the consistency evaluation network C (i) Adopts a five-layer feedforward neural network structure: an input layer, a 128-dimensional hidden layer, a 64-dimensional hidden layer, a 32-dimensional hidden layer, and a 6-dimensional output layer.
[0107] Step 9: Input the values of the action decisions selected by each drone according to its respective observation information in Step 6 and the values of the action decisions selected by each drone according to the cooperative observation information from different drones in Step 8 into the consistency evaluation module to calculate the action non-consistency penalty P (P < 0). In this embodiment, the mean square error is used as the calculation formula for the penalty, and the specific calculation formula is shown in Formula (1):
[0108]
[0109] Step 10: Use the temporal difference algorithm to obtain the loss term L. The calculation formula for the loss function L is shown in Formula (2), where γ represents the attenuation factor. In this embodiment, γ is set to 0.95. Minimize the loss function L through the backpropagation algorithm to update the parameters of each network, making the strategies of each drone gradually approach the optimal strategy, and ensuring the continuous improvement of the overall task efficiency of the drone cluster.
[0110]
[0111] Step 11: Determine whether the current strategy has converged by whether the task completion rate of the scenario exceeds 90% for 50 consecutive times. If the strategy has not converged, the system will repeat Steps 1-10. After the strategy converges, stop training and updating, and save the models of the local networks of each drone to obtain the decision-making models of multiple drones applicable to the current scenario.
[0112] Specifically, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the present invention of a drone cluster consistency decision-making system and method based on deep reinforcement learning and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0113] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium, including several instructions to enable a device (which can be a personal computer, a server, a single-chip microcomputer, an MCU or a network device, etc.) containing a data processing unit to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0114] The present invention provides an idea and method for a UAV swarm consensus decision-making system and method based on deep reinforcement learning. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. A drone cluster consensus decision system based on deep reinforcement learning, characterized in that: include: Experience replay pool, observation conversion module, consistency decision adjustment module, multi-UAV hybrid network, network optimizer and UAV mission simulation platform; among them, The experience replay pool is used to store the observation information, actions, states and rewards generated by the drone cluster during the mission execution; The observation conversion module converts the observation information of the current environment where the UAV is located into the collaborative observation information input into the consistency decision adjustment module; The consistency decision adjustment module is used to coordinate the decision-making behaviors of different UAVs after collaborative observation; The multi-UAV hybrid network combines the local information of each UAV, jointly optimizes the decision of each UAV, and outputs the overall decision of the UAV cluster; The network optimizer optimizes and updates the parameters of the multi-UAV hybrid network according to the feedback information obtained from the UAV mission simulation platform; The UAV mission simulation platform is used to provide a training and testing environment, simulate the actual mission scenarios of the UAV cluster, and provide status feedback to the multi-UAV hybrid network and network optimizer.
2. The UAV cluster consensus decision system based on deep reinforcement learning according to claim 1 is characterized in that: The consistency decision adjustment module includes: A decision consistency evaluation network corresponding to the number of drones is used to calculate the local decision value of each drone and coordinate the decisions among the drones through a consistency constraint mechanism.
3. The UAV cluster consistency decision system based on deep reinforcement learning according to claim 2 is characterized in that: The multi-UAV hybrid network includes: The drone value network corresponding to the number of drones and an overall value summary network; among them, The drone value network generates corresponding local action values according to the local observation information of the current drone; The overall value aggregation network calculates the total value of the drone cluster based on the local action value of each drone, which is used to guide the overall decision of the drone cluster.
4. The UAV cluster consistency decision system based on deep reinforcement learning according to claim 3 is characterized in that: The drone value network generates the corresponding local action value according to the local observation information of the current drone, including: Observation information of drone i Input to the corresponding drone value network Q (i) In the above example, we can get each action a of the drone under current observation. k The value is expressed as follows: Among them, A represents the preset action space, a k For any of these actions; According to the preset greedy strategy, a random selection decision is made with probability ε for exploration, and the action with the greatest value is selected with probability 1-ε, which is expressed as follows: Finally, we get the action selected by drone i at the current time t Among them, the drone value network Q (i) , a four-layer feedforward neural network structure is adopted, namely: input layer, 64-dimensional hidden layer, 32-dimensional hidden layer and 6-dimensional output layer.
5. The UAV cluster consistency decision system based on deep reinforcement learning according to claim 4 is characterized in that: The UAV mission simulation platform provides status feedback, including: According to the action selected by drone i at the current time t Compare the preset mission objectives to calculate the reward R t .
6. The UAV cluster consistency decision system based on deep reinforcement learning according to claim 5 is characterized in that: The multi-UAV hybrid network jointly optimizes the decisions of each UAV, including: Replay observation information in the experience pool Input to each drone value network Q (i) Re-evaluate the selected action The value of It is expressed as follows: Evaluate the action value of all drones and global state information S t Input to the multi-UAV hybrid network M to evaluate the overall action value of the UAV cluster It is expressed as follows: Where I is the number of drones in the drone cluster; The observation information of the next moment Input to each drone value network Q (i) And select the action with the greatest value It is expressed as follows: Get the value assessment at the next moment Combine it with the global state information S at the next moment t+1 Input to the multi-UAV hybrid network M to evaluate the overall action value of the UAV cluster at the next moment, which is expressed as follows: Among them, the multi-UAV hybrid network M is composed of an encoder layer, a 4-head attention layer and 2 fully connected layers. Among them, the encoder layer maps the input to a 128-dimensional feature space, the 4-head attention layer processes the interaction between UAVs, and the 2 fully connected layers, which are 128 dimensions and 64 dimensions respectively, output the estimated value of the final action value.
7. The UAV cluster consistency decision system based on deep reinforcement learning according to claim 6 is characterized in that: The consistency decision adjustment module coordinates the decision-making behaviors of different drones after collaborative observation, that is, inputs the collaborative observation information into the consistency evaluation network of each drone, obtains the action values of different drones under different collaborative observation information, and selects the corresponding action decision, which specifically includes: The collaborative observation information of any UAV i and different UAV j is set as Input to the consistency evaluation network C of the corresponding drone i (i) In the above equation, the action values of different UAVs under different collaborative observation information are obtained, which are expressed as follows: Select the action decision with the greatest value under different collaborative observations It is expressed as follows: Get action value It is expressed as follows: Among them, the consistency assessment network C (i) A five-layer feedforward neural network structure is adopted, including an input layer, a 128-dimensional hidden layer, a 64-dimensional hidden layer, a 32-dimensional hidden layer and a 6-dimensional output layer.
8. A drone cluster consensus decision method based on deep reinforcement learning, characterized in that: Using the system described in any one of claims 1 to 8 to implement drone cluster consensus decision-making, the method comprises the following steps: Step 1, initialize the drone cluster in the drone mission simulation platform and configure the initial state of each drone; the drone mission simulation platform simulates the real mission scenario, sets the mission objectives, environmental conditions and behavior patterns of the other drone cluster; Step 2, obtaining the status information of the current environment and the observation information of each drone through the drone mission simulation platform; wherein the status information includes: the location information, speed information, sensor data, the location of the other drone and related environmental information of the drone; Step 3: Each drone uses the corresponding drone value network to evaluate the current local state using deep reinforcement learning methods, outputs the local action value, and selects the corresponding action decision; Step 4: The action decisions selected by each drone in step 3 according to its own observation information are recorded in the drone mission simulation platform and executed in sequence by the drone cluster; the drone mission simulation platform records the execution results of each drone in real time, obtains the state information and observation information of each drone at the next moment, settles the reward according to the mission goal, and stores the above information in the experience replay pool; Step 5: Check whether the data in the experience replay pool meets the preset conditions and is sufficient for sampling. If not, repeat steps 1 to 4. Otherwise, batch collect the observation information, state information, actions, and rewards generated by the drone cluster during the task execution from the experience replay pool; Step 6: Each drone value network re-evaluates the value of the action selected by the drone based on the observation information collected from the experience replay pool, and inputs it and the state information into the multi-drone hybrid network to evaluate the overall action value of the drone cluster; Step 7, performing observation conversion on the observation information of each drone collected in the experience playback pool to obtain collaborative observation information of different drones; wherein the collaborative observation information includes: collaborative observation information between any two drones; Step 8: Input the collaborative observation information in step 7 into the decision consistency evaluation network of each UAV, obtain the action value of different UAVs under different collaborative observation information, and select the corresponding action decision; Step 9, input the value of the action selected by each drone in step 6 and the action value obtained by each drone in step 8 into the consistency decision adjustment module to calculate the action inconsistency penalty; Step 10, design the loss function, that is, use the time difference method to obtain the loss term, minimize the loss function through the back propagation method, and update the system parameters; Step 11, by setting the completion rate of multiple consecutive tasks in the scene, determine whether the current strategy has converged; if the strategy has not converged, repeat steps 1 to 10, otherwise stop training and update to complete the drone cluster consistency decision based on deep reinforcement learning.
9. The method for unmanned aerial vehicle cluster consistency decision-making based on deep reinforcement learning according to claim 8 is characterized in that: The data in the experience replay pool is sampled through a priority experience replay mechanism, that is, the states and actions in the saved data are sorted according to the errors and are prioritized in turn.
10. The method for unmanned aerial vehicle cluster consistency decision-making based on deep reinforcement learning according to claim 9, characterized in that: The consistency decision adjustment module transmits local state and action decision information between UAVs through a communication-based mechanism.