Multi-unmanned aerial vehicle coordination control method based on dynamic directed graph communication structure
By adopting dynamic directed graph communication structure and graph collapse network in multi-UAV systems, the scalability problems and global state dependence when the number of drones changes are solved, and efficient information integration and collaborative decision-making are achieved.
Patent Information
- Application Number
- CN202510136886.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-07
Smart Images

Figure CN120103868A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-UAV cooperative control, and in particular relates to a multi-UAV coordinated control method based on a dynamic directed graph communication structure. Background Art
[0002] In recent years, collaborative multi-UAV reinforcement learning technology has emerged as an important framework for solving complex collaborative tasks in real-world scenarios, attracting a lot of attention and showing great potential in practical applications and business. The centralized collaborative MARL method treats the multi-UAV system as a whole and then uses a single-UAV method to learn the overall strategy. However, this method faces the problem of insufficient scalability and the limitation of setting up a central controller. The distributed collaborative MARL method treats UAVs as separate individuals and uses a single-UAV method to learn individual strategies. Although it solves the limitations of the central controller, it leads to non-stationarity and credit allocation problems. In order to solve these problems, existing multi-UAV reinforcement learning algorithms mainly adopt the centralized training with distributed execution (CTDE) paradigm. This paradigm requires access to the global state, and the setting that can access the global state is usually an idealized assumption, and there is usually no central trainer in the real world. In contrast, studying communication between UAVs is more practical.
[0003] In multi-UAV systems, communication is particularly important for information sharing, learning, and collaboration to achieve common goals, especially in partially observable environments. Effective communication is the basis for cooperation for complex tasks such as coordinating autonomous vehicles, sensor networks, and multi-robotic systems. Inspired by the way humans cooperate, researchers have integrated communication into multi-UAV reinforcement learning to enhance the ability to share information between UAVs. Early research results mainly disseminated information through broadcasting, but this approach resulted in high communication costs and information redundancy. Later research aimed to reduce communication overhead by selectively determining communication times and eliminating redundant information through directed peer-to-peer communication. However, existing methods usually model communication for each UAV separately, which lacks scalability when the number of UAVs changes. There is still a lack of a method that can solve five key communication problems, including how to determine the communication object, communication time, communication content, how to integrate the received information, and how to use communication to avoid dependence on the global state. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a multi-UAV coordination control method based on a dynamic directed graph communication structure.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] On one hand, the present invention provides a multi-UAV coordination control method based on a dynamic directed graph communication structure, characterized in that it includes the following steps:
[0007] The drone obtains local observations from the environment and obtains historical hidden states from other drones through communication;
[0008] In the distributed execution stage, based on the initial topological structure of the drones, the adjacency trajectory matrix used to represent the communication relationship between drones at the current moment is obtained, and the pre-communication information is calculated based on the local observation value and the historical hidden state. The current adjacency trajectory matrix and the pre-communication information are input to the multi-key gated communication network to calculate the adjacency trajectory matrix at the next moment. Each drone calculates the local action value function of each drone based on the pre-communication information and the adjacency trajectory matrix at the next moment through the Transformer-based decoder, and updates the hidden state. According to the local action value function and the local observation value, the final control decision is generated and executed through the graph collapse network and the hybrid network, and then enters the next cycle;
[0009] In the centralized training phase, the adjacency trajectory matrix at each moment in the distributed execution phase is recorded, and the local observation value and the adjacency trajectory matrix at each moment are input into the graph collapse network to generate the global state at each moment. Based on the global state at each moment and the local action value function of each drone at each moment, a global action value function at each moment is generated. Based on the global action value function, the drone swarm interacts with the environment and obtains rewards. Based on the rewards, the parameters of the hybrid network are updated through the first loss function.
[0010] After the centralized training is completed, the trained graph collapse network and the hybrid network are deployed in each drone.
[0011] Furthermore, the adjacency trajectory matrix represents the communication structure of multiple UAVs. The adjacency trajectory matrix is a Boolean matrix, and the value of the matrix element indicates whether there is communication between the UAVs.
[0012] Furthermore, the calculation of pre-communication information based on local observation values and historical hidden states specifically includes:
[0013] Through an encoding network MLP, local observations and historical hidden states are converted into pre-communication information C 0 :
[0014]
[0015] Among them, C 0 For pre-communication information, is the local observation value, Hidden state for history.
[0016] Furthermore, the input of the current adjacency trajectory matrix and the pre-communication information to the multi-key door control communication network to calculate the adjacency trajectory matrix at the next moment specifically includes:
[0017] The current adjacency trajectory matrix A t Pre-communication information 0 Enter the multi-key door control communication network for processing. The calculation formula of the multi-key door control communication network is as follows:
[0018]
[0019] Among them, K 0 is the output of the key, which represents the updated adjacency trajectory matrix in this communication. Gumbel-softmax is a discrete sampling technique used to sample from discrete communication structures to determine whether each pair of drones communicates. v is the linear layer weight of the output dimension, used to generate the final communication judgment, tanh is the activation function, C 0 For pre-communication information, W q is the query weight, W k is the weight of the key;
[0020] After processing through the multi-key door control communication network, a total of i keys are obtained at time t, denoted as k t|i , the adjacency trajectory matrix at the next moment is expressed as:
[0021]
[0022] in, Indicates rounding up, A t+1 is the adjacency trajectory matrix at the next moment.
[0023] Furthermore, each drone calculates the local action value function of each drone based on the pre-communication information and the adjacent trajectory matrix at the next moment through a Transformer-based decoder, and updates the hidden state, specifically including:
[0024] The pre-communication information is repeatedly processed by the multi-layer perceptron MLP and the adjacent trajectory matrix of the next moment is input into the Transformer-based decoder. The Transformer-based decoder includes a self-attention mechanism, a position feedforward network, a residual connection and a normalization layer. The input information is processed by a multi-head attention mechanism, a residual connection and normalization, a point-by-point feedforward neural network, a residual connection and normalization. The processed information is added to the pre-communication information, and the updated hidden state of drone i is obtained through the fully connected layer. With the local action value function Q i(s,a).
[0025] Furthermore, the graph collapse network includes a graph convolutional network, a self-attention pooling network and a reading mechanism.
[0026] Furthermore, the global state acquisition process includes:
[0027] The local observation values of each UAV As the feature input of each node in the graph convolutional network, feature extraction and node representation are performed. The formula is:
[0028]
[0029] in, is the local observation value of UAV i at time t, is the representation of a node, is the value of the i-th node of the adjacency trajectory matrix at the next moment, represents the degree matrix of the ith UAV at time t+1, and σ is a nonlinear activation function;
[0030] The learned node representation Input to the self-attention pooling network, the self-attention pooling network selects important nodes in the graph and weights the features of important nodes through feature weighting and node importance evaluation, and outputs the feature weighted nodes of the graph. The formula is:
[0031]
[0032] in, is the weighted feature node representation of node i, concat is the feature fusion method, tanh is the activation function, and ⊙ represents element-by-element multiplication;
[0033] Aggregate the weighted feature node representations of each node through the reading mechanism to generate the global state s t , the formula is:
[0034]
[0035] Where N is the number of nodes, max is the aggregation function, and i and n are both positive integers.
[0036] Furthermore, the generating of the global action value function at each moment based on the global state at each moment and the local action value function of each drone at each moment specifically includes:
[0037] The global state s t And the local action value function Q of each drone i (s,a) is input into the hybrid network, and the drone swarm learns the hybrid network Q tot(τ, a, Am, s; θ), using the QMIX network, the global state is used to determine the weight of each local action value function, and the local action value function of each drone is weighted and synthesized into a global action value function Q tot (s,a).
[0038] Furthermore, the first loss function is:
[0039]
[0040] in, is the first loss function, b is the batch size, which indicates the number of samples sampled from the replay buffer, is the target Q value of drone i, Q tot (τ t ,α t ,m t ,s t ; θ) is the Q value of the hybrid network at time t, τ t is the historical observation value at time t, α t is the action at time t, m t is the information of the drone group at time t, s t is the global state at time t, θ is the parameter of the hybrid network, r is the reward obtained by the drone swarm interacting with the environment, γ is the discount factor, represents the maximization action a t+1 Q tot value.
[0041] Furthermore, the distributed execution phase and the centralized training phase are bridged via an adjacency trajectory matrix.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] (1) A new communication-cooperative multi-UAV reinforcement learning paradigm is implemented, which solves the scalability problem when the number of UAVs changes: For the multi-UAV system, a graph collapse network is used to achieve bridging and synchronous communication between the training and execution phases. The multi-UAV system is modeled as a dynamic directed graph, which effectively represents the communication structure of the system at any time, avoiding the problem of traditional methods that require each UAV to be modeled separately, and improving the scalability of the method.
[0044] (2) The global state information aggregation is realized by using the topological structure of the dynamic directed graph, which reduces the system's dependence on the global state: For the multi-UAV system, a multi-key gated communication network is designed. The dynamic directed graph structure is learned in the distributed execution phase and combined with a Transformer-based decoder to achieve communication and feature extraction. The dynamic directed graph structure learned by the multi-key gated communication network is input into the graph collapse network in the centralized training phase to aggregate information and generate approximate global state information.
[0045] (3) The present invention dynamically adjusts the adjacency trajectory matrix between drones through a multi-key gated communication network, so that the drone group can flexibly adjust the communication structure in a complex environment, thereby improving the efficiency and reliability of information transmission. Compared with the communication method with a fixed topology structure, this method can adapt to different task requirements and achieve more efficient collaborative control.
[0046] (4) The present invention uses a Transformer-based decoder to process the local observation values and communication information of the UAV, fully utilizes the self-attention mechanism to enhance the information interaction capability, improves the decision-making accuracy of each UAV, and avoids the decision-making errors caused by insufficient local information in traditional methods.
[0047] (5) The present invention extracts the overall characteristics of the drone swarm through a graph convolutional network and a self-attention pooling network, and uses a reading mechanism to aggregate the information of each drone to generate an accurate global state. This method can effectively reduce noise interference, improve the expression ability of global information, and enhance the collaborative perception ability of the drone swarm.
[0048] (6) The present invention adopts the QMIX hybrid network to adaptively adjust the weight of each drone's local action value function according to the global state to achieve the global optimal decision of the drone group. Compared with the traditional independent Q-learning method, this method can effectively solve the non-stationary problem in multi-agent reinforcement learning and improve the convergence speed and control accuracy.
[0049] (7) In the training phase, the hybrid network parameters are optimized through centralized training, so that the drone swarm can learn the global optimal strategy. In the execution phase, each drone only needs local computing and limited communication to make independent decisions, thus achieving efficient distributed execution. This method reduces communication overhead and improves the robustness and real-time performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of the overall module in the present invention;
[0051] Figure 2 It is a schematic diagram of the graph collapse network module;
[0052] Figure 3 It is a system relationship diagram of graph collapse network modules and methods;
[0053] Figure 4 The figure is a flowchart of an example of a multi-UAV reinforcement learning method based on a dynamic directed graph communication structure. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0055] Embodiment 1:
[0056] One aspect of the present invention provides a multi-UAV reinforcement learning method based on a dynamic directed graph communication structure, which is applied to multi-UAV collaboration. The method specifically includes the following steps:
[0057] Step S1: The drone obtains local observations from the environment and obtains historical hidden states from other drones through communication. Refers to the state information obtained by each drone from the environment at the current moment (t moment). This information includes the environmental features and self-state perceived by the drone itself, such as position, speed, attitude, surrounding obstacles, target position, etc. Local observations are usually obtained directly from sensors (such as vision, radar, lidar, etc.) when each drone performs a task, and only reflects the local environment around the drone; the historical hidden state refers to the representation of the internal state of each drone at the past moment (t-1 moment). It contains the decision-making process of the drone in the past period of time, the observed environmental information, the actions performed, and the communication history with other drones. The historical hidden state is used to save and transmit important memories of the drone in the process of performing the task, so as to facilitate reference in future decisions. The historical hidden state encodes the drone's observation of the environment (local observations) at the past moment and the decisions made based on these observations. The historical hidden state records the changes in the drone's state over time and contains dynamic information about how it changes during the execution of the task. The historical hidden state also includes interaction information between drones, communication information from other drones, and how to make coordinated decisions based on this information.
[0058] Step S2: In the distributed execution stage, based on the initial topological structure of the drones, the current adjacency trajectory matrix used to represent the communication relationship between drones at the current moment is obtained, and the pre-communication information is calculated based on the local observation value and the historical hidden state. The current adjacency trajectory matrix and the pre-communication information are input to the multi-key gate control communication network to calculate the adjacency trajectory matrix at the next moment. Each drone calculates the local action value function of each drone based on the pre-communication information and the adjacency trajectory matrix at the next moment through a Transformer-based decoder, and updates the hidden state. According to the local action value function and the local observation value, the final control decision is generated and executed through the graph collapse network and the hybrid network, and then enters the next cycle;
[0059] Step S3: In the centralized training stage, the adjacency trajectory matrix at each moment in the distributed execution stage is recorded, and the local observation value and the adjacency trajectory matrix at each moment are input into the graph collapse network to generate the global state at each moment. Based on the global state at each moment and the local action value function of each drone at each moment, a global action value function at each moment is generated. Based on the global action value function, the drone swarm interacts with the environment and obtains rewards. Based on the rewards, the parameters of the hybrid network are updated through the first loss function.
[0060] Step S4: After the centralized training is completed, the trained graph collapse network and the hybrid network are deployed in each drone.
[0061] As a preferred technical solution, the multi-UAV system is modeled as a dynamic directed graph. Further, each UAV is a node in the graph, and the adjacency trajectory matrix A t Represents the communication structure of the system. Specifically, A t It is a Boolean matrix, and the value of the matrix element indicates whether there is communication between the drones.
[0062] As a preferred technical solution, the Transformer-based graph collapse network consists of a graph collapse network and a Transformer-based multi-key gated communication network. Furthermore, the graph collapse network includes a graph convolutional network, a self-attention pooling network and a reading mechanism; the Transformer-based multi-key gated communication network includes an encoder based on a dynamic directed graph and a decoder based on a Transformer, following the encoder-decoder architecture.
[0063] As a preferred technical solution, the local observation values are provided by drones. Specifically, at each time step t, each drone receives a local observation value from the observation value function.
[0064] As a preferred technical solution, the process of obtaining the adjacency trajectory matrix and fusion information specifically includes initializing and preprocessing the local observations and historical information; performing pre-communication and obtaining pre-communication information by combining the adjacency trajectory matrix of the previous time step and the hidden information of the drone at the previous time step; and using a multi-key gated communication network to process the pre-communication information and the initialized adjacency trajectory matrix to obtain the adjacency trajectory matrix and fusion information of the next time step.
[0065] As a preferred technical solution, the Transformer-based decoder consists of four parts, specifically, a self-attention mechanism, a position feedforward network, a residual connection and a normalization layer.
[0066] As a preferred technical solution, the process of generating the local action value function and hidden information specifically includes inputting the adjacency trajectory matrix and fusion information output by the multi-key door control communication network into a Transformer-based decoder; using the output of the Transformer-based decoder to obtain the local action value function and the hidden information of the next time step.
[0067] As a preferred technical solution, the process of obtaining the global action value function specifically includes inputting the local action value function and the adjacency trajectory matrix into the graph collapse network; normalizing the adjacency trajectory matrix; then using the local observation value as the feature input of each node in the graph convolution network to perform graph convolution iteration; further, based on the graph convolution iteration result, performing self-attention pooling; generating an approximate global state function by performing aggregation operations on all nodes; and finally, based on the approximate global state function, outputting the global action value function through a hybrid network.
[0068] As a preferred technical solution, there is a connection between the distributed execution stage and the centralized training stage, specifically, the connection is bridged by a dynamic directed graph.
[0069] Another aspect of the present invention provides a new reading mechanism. Specifically, an aggregation operation is performed on all nodes to generate a global collapsed representation of the graph, which is close to the global state.
[0070] The present invention designs a method for a multi-UAV system that simultaneously includes functions such as communication, information extraction, and decision-making. It supports a multi-UAV system with dynamic adjustment of the number of UAVs, provides support for the coordinated control of the multi-UAV system, and strengthens the communication and decision-making of UAVs.
[0071] like Figure 1 As shown, the overall system model of this embodiment includes a multi-key gated communication network, a graph collapse network, and a hybrid network, and the multi-key gated communication network and the graph collapse network are connected to the drone.
[0072] The multi-key gated communication network is used to receive local observations from the drone via the communication portion of the network Historical Information and the adjacency trajectory matrix A at the current moment t , in the distributed execution phase, the adjacency trajectory matrix A of the next moment is generated t+1 , local action value function Q i And hidden information The local action value function and the local observation value are input into the graph collapse network. The multi-key gated communication network is also used to bridge the centralized training and distributed execution through the adjacency trajectory matrix of the next moment. The input information of the multi-key gated communication network includes local observation values, historical information and the adjacency trajectory matrix of the current moment. After the input information is initialized and pre-processed, pre-communication is performed, and the pre-communication information C is processed by the multi-key gated communication network. 0 and the adjacency trajectory matrix A 0 The communication frequency between drones is determined by the number of keys, and the output of each key represents the updated adjacency trajectory matrix in this communication. The specific calculation formula is shown in formula (1):
[0073]
[0074] Among them, K 0 is the output of the key, A 0 is the adjacency trajectory matrix, gumbel-softmax sampling is a technique for sampling from a discrete distribution, w v is a linear layer for two output dimensions (corresponding or not), tanh is the activation function, C 0 For pre-communication information, W q is the query weight, W k is the weight of the key. After passing through the multi-key gated communication network, a total of i keys are obtained at time t, denoted as k t|i The adjacency trajectory matrix at the next moment is expressed as shown in formula (2):
[0075]
[0076] in, Indicates rounding up, A t+1 is the adjacency trajectory matrix at time t+1. The transmission and reception of information is completed by using the adjacency trajectory matrix updated in the next time step combined with the pre-communication information.
[0077] Graph Collapse Network for Local Observations via UAVs Next moment adjacency trajectory matrix A t+1, all inputs are aggregated during the centralized training phase to generate a global collapsed representation of the dynamic directed graph, which is approximately the global state function s t The input of the graph collapse network includes the local observations o of each drone. t and the adjacency trajectory matrix A at the next moment t+1 The input adjacency trajectory matrix is normalized, the local observation value of the drone is used as the input of the graph convolution network, the result after three graph convolution iterations is input into the self-attention pooling network, and the pooled result is aggregated through the reading mechanism to generate the global state, which is then input into the hybrid network.
[0078] The hybrid network is used to aggregate the local action value function Q of each drone. i , forming a global action value function Q tot , to guide the overall strategy optimization of the multi-UAV system. The input of the hybrid network is the local action value function of each UAV and the global state function s output by the graph collapse network t , the output is the global action value function, which evaluates the overall value of a multi-UAV system taking a certain coordinated action in a certain state. The hybrid network can reflect the collaborative effect between UAVs and guide the UAV strategy update.
[0079] like Figure 2 As shown, the graph collapse network module of the present invention includes a graph convolutional network, a self-attention pooling network and a reading mechanism, the reading mechanism is connected to the hybrid network, and the graph convolutional network is connected to the multi-key gated communication network.
[0080] The graph convolutional network transforms the local observation value o of the drone into t As the feature input of each node in the graph convolutional network, feature extraction and node representation are performed, and the learned node representation is input into the self-attention pooling network. The input of the convolutional network is the local observation, and the output is the representation of all nodes in the graph. The specific network structure is shown in formula (3):
[0081]
[0082] in, is the local observation value of UAV i at time t, is the representation of a node, is the value of the i-th node of the adjacency trajectory matrix at the next moment, represents the degree matrix of the i-th UAV at time t+1, σ is a nonlinear activation function; the graph convolutional network is mainly responsible for feature extraction and information propagation in the data processing of dynamic directed graphs.
[0083] The function of the self-attention pooling network is to select important nodes in the graph and weight the features of important nodes through feature weighting and node importance evaluation. The input is the representation of all nodes in the graph, and the output is the feature weighted nodes of the graph. The structure of the self-attention pooling network in the present invention is shown in formula (4):
[0084]
[0085] in, is the weighted feature node of the graph, is the representation of the node, concat is the feature fusion method, tanh is the activation function, is the self-circulating adjacency matrix of the ith node at time t+1, represents its degree matrix.
[0086] The reading mechanism aggregates all nodes to generate a global collapsed representation of the graph and an approximate global state function. The input of the reading mechanism is all nodes in the graph, and the output is a global collapsed representation of the graph. The reading mechanism uses the Mean aggregation function and the Max aggregation function to take the average and maximum values of the nodes respectively. The Mean aggregation function smoothly integrates the information of all nodes and reduces the influence of noise. The Max aggregation function highlights the significant features of specific nodes and retains the most important information in the graph. The expression of the reading mechanism is shown in formula (5):
[0087]
[0088] Among them, s t is the global state of the system at time t, is the weighted feature node of the graph, N is the number of nodes, max is the aggregation function, and i and n are both positive integers.
[0089] Specifically, the step of obtaining the fusion information and the adjacency trajectory matrix of the next time step through the encoder based on the dynamic directed graph is to obtain input information through the communication part of the network module, including the local observation value of the drone, historical information, and the adjacency trajectory matrix A at the current moment. t ; Embed and repeatedly process the input information to convert it into a continuous vector; input the processed data into the multi-key door control unit, and control the transmission of information through multiple keys; based on the output of the multi-key door control unit, update the adjacency matrix A through graph processing t+1 , represents the graph structure of the next time step; the adjacency trajectory matrix A t+1 The information after repeated processing is used as fusion information and is input into the Transformer-based decoder together with the information after embedding processing.
[0090] Specifically, based on the value decomposition algorithm, the relationship between the local action value function and the global action value function is shown in formula (6):
[0091]
[0092] Among them, Q tot (τ,a) is the global action value function, τ is the historical local observation value, a is the action, argmax means finding the maximum point, Q i is the local action value function of UAV i.
[0093] Specifically, the steps of generating the local action value function and hidden information through the Transformer-based decoder are to obtain fusion information from the encoder based on the dynamic directed graph; the input information is processed by a multi-head attention mechanism, residual connection and normalization, a point-by-point feedforward neural network, residual connection and normalization, and the processed information is added to the embedded information output by the encoder, and the hidden information is obtained through the fully connected layer, and the drone i learns Q i (τ,a,Am,s) network to obtain the local action value function Q of the drone i (s,a).
[0094] Specifically, the steps of obtaining the global action value function through the graph collapse network and the hybrid network are to obtain input information through the communication part of the network module, including local observations, the adjacency matrix A of the next time step t+1 And the local action value function of the drone; for the adjacency matrix A t+1 Normalize, combine the adjacency trajectory matrix and local observation to obtain the communication structure of the node at the current time step; input the local observation value into the graph convolution network; perform three graph convolution iterations; perform self-attention pooling based on the iterated information; aggregate the pooled information through the reading mechanism to generate the approximate global state of the system; input the global state and the local action value function into the hybrid network, and the drone swarm learns the hybrid network Q tot (τ, a, Am, s; θ), using the QMIX network, the global state is used to determine the weight of the local action value function, and the local action value function of each drone is weighted and synthesized into a global action value function Q tot (s,a).
[0095] Specifically, the hybrid network Q tot The θ parameters of are updated by the loss function shown in formula (7):
[0096]
[0097] in, is the first loss function, b is the batch size sampled from the replay buffer, is the target Q value of drone i, Q tot (τ t ,α t ,m t ,s t ; θ) is the Q value of the hybrid network at time t, τ t is the historical observation value at time t, α t is the action at time t, m t is the information of the drone group at time t, s t is the global state at time t, θ is the parameter of the hybrid network, r is the reward obtained by the interaction between the drone swarm and the environment, γ is the discount factor, represents the maximization action a t+1 Q tot value.
[0098] Specifically, the information transmitted by the drone in the present invention is a local observation value, and the relationship between the local observation value and the global state is shown in formula (8):
[0099]
[0100] Among them, s t is the global state of the system at time t, is the local observation of UAV j at time t, Agg represents aggregation, v represents vertex, N(v i ) represents the neighboring node set of node i, σ and W represent the parameters of the feedforward neural network, and i and j are both positive integers.
[0101] Specifically, the global state value function Q tot (s,a) is a function of the state vector s and the joint action a. According to the implicit function theorem, Q tot It can also be viewed as a function of the local action value Q tot (s,Q i ), we can get the expression of the global action value function as shown in formula (9):
[0102]
[0103] Among them, s t is the global state of the system at time t, is the local observation of drone j at time t, Agg represents aggregation, v represents vertex, N(v i ) represents the neighboring node set of node i, σ and W represent the parameters of the feedforward neural network, Q i is the local action value function of the i-th UAV, and i and j are both positive integers.
[0104] Specifically, the parameters of the graph collapse network are updated through the process described in formula (10):
[0105]
[0106] in, is the first loss function, W is the weight, b is the bias, and α is the learning rate.
[0107] Embodiment 2:
[0108] The parts not mentioned in this embodiment are the same as those in Embodiment 1.
[0109] like Figure 3 , Figure 4 As shown, the multi-rotor UAV is a six-rotor industrial UAV with high stability and high scalability using a full carbon fiber fuselage. This embodiment will be implemented with the multi-rotor UAV as the intelligent body and based on the technical solution of the present invention.
[0110] The multi-UAV reinforcement learning module based on the dynamic directed graph communication structure will distinguish two types of task execution for multi-rotor UAVs according to different execution strategies, providing support for the whole process logic analysis and diagnosis. The specific situation is as follows:
[0111] Case 1: When a multirotor drone needs to perform a task that can access global state
[0112] Case 2: When a multirotor drone needs to perform tasks that do not have access to the global state.
[0113] like Figure 4 As stated, Figure 4 a is a flowchart of the tasks that the multi-rotor drone needs to perform and can access the global state. Figure 4 b is a flowchart of the multi-rotor drone that needs to perform tasks that cannot access the global state. When the multi-rotor drone faces situation 1, the specific steps of performing related tasks through the TGCNet module and the multi-agent reinforcement learning method include:
[0114] Step S101: constructing a comprehensive equipment platform including a drone control system, and arranging sensors and a multi-agent reinforcement learning network module based on a dynamic directed graph communication structure;
[0115] Step S102: The UAV control system sends a collaborative task instruction of the multi-rotor UAV cluster to the multi-agent system;
[0116] Step S103: inputting the mission instructions of the UAV control system and the data captured by the sensor into the TGCNet module through the wireless receiving device, and the TGCNet module analyzes the data;
[0117] Step S104: All input information is processed through the multi-key door control communication network; the processed data and the global state of the system are passed as input to the trained graph collapse network, the action of the current multi-rotor drone cluster is determined according to the generated global action value function, and the action strategy is sent to the drone control system;
[0118] Step S105: The UAV control system receives the action strategy issued by the TGCNet module for evaluation and analysis. If the execution action of the current multi-rotor UAV swarm meets the task instruction requirements issued by the UAV control system, the multi-rotor UAV swarm is allowed to execute the current action by default; if the execution action of the current multi-rotor UAV swarm does not meet the task instruction requirements issued by the UAV control system, the UAV control system will send an instruction to stop the current multi-rotor UAV swarm action, and require the integrated module to re-collect all data and repeat steps S3 to S5 until the execution action of the multi-rotor UAV swarm meets the task instruction requirements issued by the UAV swarm control system, then the process is stopped.
[0119] When the multi-rotor drone faces situation 2, the specific steps of performing related tasks through the TGCNet module and the multi-agent reinforcement learning method include:
[0120] Step S201: constructing a comprehensive equipment platform including a drone control system, and arranging sensors and a multi-agent reinforcement learning network module based on a dynamic directed graph communication structure;
[0121] Step S202: The UAV control system sends a collaborative task instruction of the multi-rotor UAV cluster to the multi-agent system;
[0122] Step S203: inputting the mission instructions of the UAV control system and the data captured by the sensor into the TGCNet module through the wireless receiving device, and the TGCNet module analyzes the data;
[0123] Step S204: All input information is processed through the multi-key gated communication network, and the processed data is passed as input to the graph collapse network; the information is completed through the graph collapse network and the global state of the current multi-rotor drone cluster is generated; the generated global state and input information are passed as input to the trained hybrid network to generate a global action value function; the action of the current multi-rotor drone cluster is determined according to the global action value function, and the action strategy is sent to the drone control system;
[0124] Step S205: The UAV control system receives the action strategy issued by the TGCNet module for evaluation and analysis. If the execution action of the current multi-rotor UAV swarm meets the task instruction requirements issued by the UAV control system, the multi-rotor UAV swarm is allowed to execute the current action by default; if the execution action of the current multi-rotor UAV swarm does not meet the task instruction requirements issued by the UAV control system, the UAV control system will send an instruction to stop the current multi-rotor UAV swarm action, and require the integrated module to re-collect all data and repeat steps S3 to S5 until the execution action of the multi-rotor UAV swarm meets the task instruction requirements issued by the UAV swarm control system, then the process is stopped.
[0125] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0126] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A multi-UAV coordination control method based on a dynamic directed graph communication structure, characterized in that: The following steps are involved: The drone obtains local observations from the environment and obtains historical hidden states from other drones through communication; In the distributed execution stage, based on the initial topological structure of the drones, the current adjacency trajectory matrix used to represent the communication relationship between drones at the current moment is obtained, and the pre-communication information is calculated based on the local observation value and the historical hidden state. The current adjacency trajectory matrix and the pre-communication information are input to the multi-key gated communication network to calculate the adjacency trajectory matrix at the next moment. Each drone calculates the local action value function of each drone based on the pre-communication information and the adjacency trajectory matrix at the next moment through the Transformer-based decoder, and updates the hidden state. According to the local action value function and the local observation value, the final control decision is generated and executed through the graph collapse network and the hybrid network, and then enters the next cycle; In the centralized training phase, the adjacency trajectory matrix at each moment in the distributed execution phase is recorded, and the local observation value and the adjacency trajectory matrix at each moment are input into the graph collapse network to generate the global state at each moment. Based on the global state at each moment and the local action value function of each drone at each moment, a global action value function at each moment is generated. Based on the global action value function, the drone swarm interacts with the environment and obtains rewards. Based on the rewards, the parameters of the hybrid network are updated through the first loss function. After the centralized training is completed, the trained graph collapse network and the hybrid network are deployed in each drone.
2. According to claim 1, a multi-UAV coordinated control method based on a dynamic directed graph communication structure is characterized in that: The adjacency trajectory matrix represents the communication structure of multiple UAVs. The adjacency trajectory matrix is a Boolean matrix, and the value of the matrix element indicates whether there is communication between the UAVs.
3. The multi-UAV coordination control method based on a dynamic directed graph communication structure according to claim 1 is characterized in that: The calculation of pre-communication information based on local observation values and historical hidden states specifically includes: The local observations and historical hidden states are converted into pre-communication information C0 through an encoding network MLP: Among them, C0 is the pre-communication information, is the local observation value, Hidden state for history.
4. The multi-UAV coordination control method based on a dynamic directed graph communication structure according to claim 1 is characterized in that: The process of inputting the current adjacency trajectory matrix and pre-communication information to the multi-key door control communication network and calculating the adjacency trajectory matrix at the next moment specifically includes: The current adjacency trajectory matrix A t The pre-communication information C0 is input into the multi-key door control communication network for processing. The calculation formula of the multi-key door control communication network is as follows: Among them, K0 is the output of the key, which represents the updated adjacency trajectory matrix in this communication. Gumbel-softmax is a discrete sampling technique used to sample from discrete communication structures to determine whether each pair of drones communicates. v is the linear layer weight of the output dimension, which is used to generate the final communication judgment, tanh is the activation function, C0 is the pre-communication information, and W q is the query weight, W k is the weight of the key; After processing through the multi-key door control communication network, a total of i keys are obtained at time t, denoted as k t|i , the adjacency trajectory matrix at the next moment is expressed as: in, Indicates rounding up, A t+1 is the adjacency trajectory matrix at the next moment.
5. The multi-UAV coordination control method based on a dynamic directed graph communication structure according to claim 1 is characterized in that: Each drone calculates the local action value function of each drone based on the pre-communication information and the adjacent trajectory matrix at the next moment through a Transformer-based decoder, and updates the hidden state, specifically including: The pre-communication information is repeatedly processed by the multi-layer perceptron MLP and the adjacent trajectory matrix of the next moment is input into the Transformer-based decoder. The Transformer-based decoder includes a self-attention mechanism, a position feedforward network, a residual connection and a normalization layer. The input information is processed by a multi-head attention mechanism, a residual connection and normalization, a point-by-point feedforward neural network, a residual connection and normalization. The processed information is added to the pre-communication information, and the updated hidden state of drone i is obtained through the fully connected layer. With the local action value function Q i (s,a).
6. The multi-UAV coordination control method based on a dynamic directed graph communication structure according to claim 1 is characterized in that: The graph collapse network includes a graph convolutional network, a self-attention pooling network and a reading mechanism.
7. The multi-UAV coordination control method based on a dynamic directed graph communication structure according to claim 1 or 6, characterized in that: The global status acquisition process includes: The local observation values of each UAV As the feature input of each node in the graph convolutional network, feature extraction and node representation are performed. The formula is: in, is the local observation value of UAV i at time t, is the representation of a node, is the value of the i-th node of the adjacency trajectory matrix at the next moment, represents the degree matrix of the ith UAV at time t+1, σ is a nonlinear activation function; The learned node representation Input to the self-attention pooling network, the self-attention pooling network selects important nodes in the graph and weights the features of important nodes through feature weighting and node importance evaluation, and outputs the feature weighted nodes of the graph. The formula is: in, is the weighted feature node representation of node i, concat is the feature fusion method, tanh is the activation function, and ⊙ represents element-by-element multiplication; Aggregate the weighted feature node representations of each node through the reading mechanism to generate the global state s t , the formula is: Where N is the number of nodes, max is the aggregation function, and i and n are both positive integers.
8. The method for coordinated control of multiple UAVs based on a dynamic directed graph communication structure according to claim 1, characterized in that: The generating of the global action value function at each moment based on the global state at each moment and the local action value function of each drone at each moment specifically includes: The global state s t And the local action value function Q of each drone i (s,a) is input into the hybrid network, and the drone swarm learns the hybrid network Q tot (τ, a, Am, s; θ), using the QMIX network, the global state is used to determine the weight of each local action value function, and the local action value function of each drone is weighted and synthesized into a global action value function Q tot (s,a).
9. The method for coordinated control of multiple UAVs based on a dynamic directed graph communication structure according to claim 1, characterized in that: The first loss function is: in, is the first loss function, b is the batch size, which indicates the number of samples sampled from the replay buffer, is the target Q value of drone i, Q tot (τ t ,α t ,m t ,s t ; θ) is the Q value of the hybrid network at time t, τ t is the historical observation value at time t, α t is the action at time t, m t is the information of the drone group at time t, s t is the global state at time t, θ is the parameter of the hybrid network, r is the reward obtained by the drone swarm interacting with the environment, γ is the discount factor, represents the maximization action a t+1 Q tot value.
10. The method for coordinated control of multiple UAVs based on a dynamic directed graph communication structure according to claim 1, characterized in that: The distributed execution phase is bridged with the centralized training phase through an adjacency trajectory matrix.
Citation Information
Patent Citations
Unmanned cooperation system based on deep reinforcement learning and application method
CN118244774A
Unmanned aerial vehicle cooperative game behavior decision-making method based on improved QMIX algorithm
CN119356399A
Apparatus, system, method and computer-implemented storage media to implement radio resource management policies using machine learning
US20220377614A1