Unmanned aerial vehicle group coordination control method and system based on communication information completion

By adopting a coordination control method based on communication information completion in the multi-UAV reinforcement learning system, using the information-level weight network and the adaptive generation network to generate global state and action value functions, the problems of low efficiency of the UAV collaborative decision-making and low accuracy of the communication model in the existing technology are solved, and more efficient information exchange and collaboration capabilities are achieved.

CN120122718AActive Publication Date: 2025-06-10TONGJI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510136889.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-10
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing multi-UAV reinforcement learning algorithms have problems such as low efficiency in collaborative decision-making and information exchange, low accuracy of communication models, and large gaps between training and execution, making it difficult to effectively deal with the scalability problem of increasing number of drones.

Method used

The UAV cluster coordination control method based on communication information completion is adopted, and global state and action value functions are generated in the centralized training stage through information-level weight network, adaptive generation network and hybrid network, and coordinated decisions are made through the Transformer decoder in the distributed execution stage to reduce dependence on the global state.

Benefits of technology

It improves the decision-making and collaboration capabilities of the drone, reduces the system's dependence on the global state, enhances the accuracy and computing efficiency of information exchange, and effectively bridges the gap between centralized training and distributed execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122718A_ABST
    Figure CN120122718A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle group coordination control method and system based on communication information completion, and the method comprises the steps: calculating an information weight based on a local observation value and the observation information of other unmanned aerial vehicles through an information-level weight network, obtaining the weighted information of the unmanned aerial vehicles, and generating a global state through a self-adaptive generation network. Then, based on the global state and the local action value function, the local action value function of the unmanned aerial vehicle is integrated through a hybrid network, and a global action value function is obtained; and in the distributed execution stage, the unmanned aerial vehicle obtains a local action value function through a Transform-based decoder according to the local observation value and the information weight, and carries out coordination decision making by using a strategy network. Compared with the prior art, information completion is realized through the generative adversarial network in the training stage, decision making is carried out only by depending on local information in the execution stage, and the calculation overhead is effectively reduced. According to the method, the cooperation and decision-making capabilities of the unmanned aerial vehicle group are improved, and meanwhile, the method has relatively good expandability and application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV cooperative control, and particularly relates to a method and system for coordinated control of a UAV swarm based on communication information completion. Background Art

[0002] In recent years, multi-UAV reinforcement learning has received extensive attention due to its potential for modeling and solving complex cooperation tasks in real-world applications. Different from single-UAV reinforcement learning, multi-UAV reinforcement learning faces the challenge of promoting UAV cooperation to achieve a common goal. The increased cooperation dimension brings many difficulties because both centralized and distributed methods cannot effectively handle the non-stationarity generated by concurrent learning of joint strategies and the scalability problem with the increase in the number of UAVs. Existing multi-UAV reinforcement learning algorithms mainly rely on the paradigm of centralized training with decentralized execution (CTDE). However, this paradigm has certain limitations. In centralized training, it is idealized to assume that UAVs can access the global state, while in decentralized execution, the decision-making of UAVs can only rely on their own local observations, which may affect their ability to make decisions and cooperate.

[0003] To improve the decision-making and cooperation ability of UAVs, more attention has been focused on the research of integrating communication mechanisms into the multi-agent reinforcement learning (MARL) framework. By allowing UAVs to exchange information such as local observations, beliefs, or intentions, the inference of the global state and coordinated decision-making can be promoted. However, the existing methods face three main problems. First, the communication model accuracy of the existing methods is not high. These models usually pre-define the communication mode and communicate through broadcast or peer-to-peer mechanisms. These methods clarify the timing of UAV communication and the objects that need to communicate, but do not penetrate communication into the information level. Second, many methods have low computational efficiency. The communication model and the policy network can only be jointly learned through sparse reinforcement learning rewards, and the communication module cannot be separated from the training process, resulting in a decline in training efficiency. Finally, communication is usually limited to the decentralized execution stage. During the centralized training process, these methods still rely on access to the global state and fail to bridge the gap between training and execution through communication. Therefore, there is currently a lack of a pluggable multi-UAV reinforcement learning method that can achieve information-level modeling to solve these problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for coordinated control of a UAV swarm based on communication information completion to overcome the defects of the above-mentioned existing technologies.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] On the one hand, the present invention provides a method for coordinated control of an unmanned aerial vehicle (UAV) swarm based on communication information completion, including the following steps:

[0007] Step S1: Each UAV obtains local observations of each UAV from the environment, and the local observations include environmental information that a single UAV can perceive;

[0008] Step S2: In the centralized training stage, based on the local observations, obtain information weights through an information-level weight network; based on the information weights, obtain weighted information of each UAV; based on the weighted information of each UAV, generate a global state through an adaptive generation network; based on the global state and the local action value function of the UAV, integrate the local action value function of the UAV through a hybrid network to obtain a global action value function; wherein, the local action value function is calculated by each UAV through an independent Q network based on the local observations; based on the global action value function, the UAV swarm interacts with the environment to obtain rewards; based on the rewards, update the parameters of the hybrid network through a first loss function; based on the information weights, local observations, and global state, update the parameters of the adaptive generation network and the information-level weight network through a generative adversarial network loss function;

[0009] Step S3: Re-obtain information weights according to the updated information-level weight network and send the obtained information weights to each UAV;

[0010] Step S4: In the distributed execution stage, each UAV obtains the local action value function of the UAV through a Transformer-based decoder based on the local observations and information weights, and based on the local action value function, makes a coordinated decision through a policy network to guide the UAV to execute corresponding actions.

[0011] Further, obtaining information weights through an information-level weight network based on the local observations specifically includes:

[0012] Perform position encoding on the local observations of UAV i and the information obtained from other UAVs to convert them into a vector representation containing position information;

[0013] The vector after position encoding is repeatedly processed and converted into the input format of the attention mechanism, including query, key, and value. The position encoding vector of the local observations of UAV i is repeatedly processed by a multi-layer perceptron (MLP) to generate a query vector The information obtained from other UAVs is processed to generate a key vector Sum vector

[0014] Calculate the query vector And the key vector Take the dot product to obtain the attention score, apply the softmax function to the attention score for normalization, and obtain the information weights between each drone. The calculation formula is:

[0015]

[0016] Where, M t Represents the set of local observations of all drones at time t, Is the information weight of drone i at time t, d is the dimension of the query vector and the key vector, σ is the ReLU activation function, and m is the number of repeated processes.

[0017] Furthermore, based on the information weights, obtain the weighted information of each drone. The formula is:

[0018]

[0019] Where, Is the weighted information of drone i at time t, Is the information weight of drone i at time t, Represents the complete information obtained by drone i at time t, Represents the set of local observations of other drones except drone i at time t, ∈ t Is white noise, Is the normal distribution, and I is the identity matrix.

[0020] Furthermore, the adaptive generation network includes an information completion network and a global discriminator network. The information completion network is connected to the information weight network, and the information completion network adopts a U-Net structure and follows an encoder-decoder structure.

[0021] Furthermore, based on the weighted information of each drone, generate the global state through the adaptive generation network, specifically including:

[0022] Input the weighted information of each drone Through downsampling processing, input the data after downsampling processing into the encoder and decoder of the information completion network, and output the global state:

[0023]

[0024] Where, s t Is the global state at time t generated by the information completion network, and n is the number of drones in the drone swarm.

[0025] Furthermore, based on the global state and the local action value function of the UAV, the local action value functions of the UAVs are integrated through a hybrid network to obtain a global action value function, which specifically includes:

[0026] All UAVs learn the hybrid network Q tot (τ, a, m, s; θ), utilize the additive structure of the QMIX network, use the global state to determine the weights of the local action value functions, and synthesize the local action value functions of each UAV into a global action value function through the determined weights of the local action value functions.

[0027] Furthermore, the first loss function is:

[0028]

[0029] where is the first loss function, n is the number of UAVs, is the target Q value of UAV i, Q tot (τ t , α t , m t , s t ; θ) is the hybrid network Q value at time t, τ t is the historical observation value at time t, α t is the action at time t, m t is the information of the UAV swarm at time t, s t is the global state at time t, θ is the parameter of the hybrid network, r is the reward obtained by the UAV swarm interacting with the environment, γ is the discount factor, represents maximizing the Q t+1 value of action a tot .

[0030] Furthermore, the loss function of the generative adversarial network is:

[0031]

[0032] where is the second loss function, ‖.‖ is the Euclidean norm, o t is the actual local observation, O(s t ) is the predicted local observation value according to the global state, D is the global discriminator network in the adaptive generative network, G is the information completion network in the adaptive generative network, E represents calculating the expectation, o ∼ O means o is sampled from the distribution O, s ∼ S means s is sampled from the distribution S, s t is the global state at time t, D(s t ) is the global state s tThe probability of whether the judgment global state s output by the input global discriminator network is true, o t is the local observation value at time t, and α is the weight hyperparameter. t

[0033] Furthermore, based on the local observation value and information weight, through a Transformer-based decoder, the local action value function of the UAV is obtained, specifically including:

[0034] According to the local observation value and information weight, the weighted information of the UAV is obtained, and the obtained weighted information and the hidden state of the UAV at time step t - 1 are used as inputs to the embedding layer. The embedding layer converts the input data into a continuous vector representation, and the output of the embedding layer is used as the input of the Transformer-based decoder. The Transformer-based decoder processes the input through the self-attention mechanism, extracts useful features from it, and outputs the hidden state of the UAV at time step t and its local action value function.

[0035] On the other hand, the present invention provides a pluggable multi-UAV reinforcement learning system for communication information completion, including:

[0036] An information-level weight network for processing based on the local observation value of each UAV and the information of other UAVs to generate information weights;

[0037] An adaptive generation network, including an information completion network and a global discriminator network. The information completion network is connected to the information-level weight network and generates a global state through an encoder-decoder structure. The global discriminator network is used to judge the authenticity of the global state;

[0038] A hybrid network for integrating the local action value functions of each UAV and generating a global action value function through an additive structure;

[0039] A Transformer-based decoder for calculating the local action value function of the UAV based on the local observation value and information weight during the distributed execution stage;

[0040] A policy network for generating the final execution action according to the local action value function and coordination decision;

[0041] A loss function module, including a first loss function and a generative adversarial network loss function, for updating the parameters of the hybrid network, the adaptive generation network, and the information-level weight network;

[0042] A communication module for transmitting weighted information and decision information between different UAVs;

[0043] An environment interaction module for simulating the interaction process between the UAV and the environment and collecting reward signals.​

[0044] Compared with the prior art, the present invention has the following advantages:

[0045] (1) Completing communication information and reducing the system's dependence on accessing the global state: For multi-UAV coordinated control, a pluggable module is used to combine the generation network with the communication module of multi-UAV reinforcement learning, allowing the UAVs to generate the global state using weighted local observations and make coordinated decisions using these states and the learned communication weights, thereby reducing the system's dependence on accessing the global state.

[0046] (2) Implementing information-level communication modeling and reducing the computational requirements related to the joint training of the communication and policy networks: The pluggable module can be pre-trained and seamlessly integrated into various MARL algorithms. By fusing the generative model loss function, the training of the communication module is decoupled from the sparse reinforcement learning rewards, thereby improving the efficiency of the policy network learning process.

[0047] (3) The pluggable multi-UAV reinforcement learning method for communication information completion of the present invention designs a flexible structure to separate the communication and policy training, thereby improving the computational efficiency and effectively bridging the gap between centralized training and distributed execution.

[0048] (4) By introducing an information-level weight network in the centralized training stage, the present invention calculates the information weights between each UAV based on the local observations of each UAV and the observation information of other UAVs, thereby effectively capturing the information correlation between UAVs. This technical means enhances the information exchange and collaboration capabilities between UAVs and overcomes the problem of low accuracy of the communication model in traditional methods.

[0049] (5) By introducing a generative adversarial network, the present invention can process global information in the centralized training stage, while only relying on local observations and communication information in the distributed execution stage, thereby effectively narrowing the gap between training and execution. This technical means improves the flexibility and scalability of the system and solves the limitation of relying on global information in centralized training in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic diagram of the pluggable module in the present invention;

[0051] Figure 2 It is a schematic diagram of the adaptive generation network;

[0052] Figure 3 It is a system relationship diagram of the pluggable module and the method;

[0053] Figure 4Schematic flowchart of an example of a pluggable multi - UAV reinforcement learning method for communication information completion. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Embodiment 1:

[0056] In this embodiment, a multi - UAV reinforcement learning method for communication information completion is provided, which is applied to multi - UAV cooperation. The method specifically includes the following steps:

[0057] Step 1: The UAV obtains local observations from the environment;

[0058] Step 2: In the centralized training stage, based on the local observations, obtain information weights through the information - level weight network; based on the information weights, obtain the weighted information of each UAV; based on the weighted information of each UAV, generate a global state through the adaptive generation network; based on the global state and the local action value function of the UAV, integrate the local action value function of the UAV through the hybrid network to obtain the global action value function; wherein, the local action value function is calculated by each UAV through an independent Q - network based on the local observations; based on the global action value function, the UAV swarm interacts with the environment to obtain rewards; based on the rewards, update the parameters of the hybrid network through the first loss function; based on the information weights, local observations and global state, update the parameters of the adaptive generation network and the information - level weight network through the generative adversarial network loss function;

[0059] Step 3: Re - obtain the information weights according to the updated information - level weight network and send the obtained information weights to each UAV;

[0060] Step 4: In the distributed execution stage, based on the local observations and information weights, obtain the local action value function of the UAV through a Transformer - based decoder, and make coordinated decisions based on the weighted information through the policy network of the UAV.

[0061] As a preferred technical solution, the pluggable module of the present invention consists of four parts. Specifically, as Figure 1As shown, it includes an information-level weight network for generating information weights, an adaptive generation network for generating global states, a Transformer-based decoder for integrating information and extracting features, and a hybrid network for generating global action value functions.

[0062] As a preferred technical solution, the local observation value is provided by the unmanned aerial vehicle. Specifically, at each time step t, each unmanned aerial vehicle receives a local observation value from the observation value function.

[0063] As a preferred technical solution, the information-level weight network has an attention mechanism. Specifically, the local observation values of each unmanned aerial vehicle are used as query vectors, and the local observation values of all unmanned aerial vehicles are used as keys and values.

[0064] As a preferred technical solution, the process of obtaining information weights through the information-level weight network specifically includes performing position encoding on the local observation values and then calculating the information-level weights through the dot-product attention mechanism.

[0065] As a preferred technical solution, the process of generating a global state through the adaptive generation network specifically includes inputting weighted information into the information completion network and performing information completion modeling through U-Net; then outputting the generated global state; further, inputting the generated global state into the global discriminator network and performing fusion through a cascade layer using the sigmoid activation function to output the probability of the global state being real, which is used to evaluate the performance of the information completion network.

[0066] As a preferred technical solution, the process of making a coordinated decision specifically includes the policy network updating the policy through the action value function.

[0067] The present invention designs a pluggable multi-unmanned aerial vehicle reinforcement learning module unit that simultaneously includes functions such as communication, information completion, decision-making, and adaptive weights for communication information completion. The module supports seamless integration into the multi-unmanned aerial vehicle reinforcement learning network, providing support for the collaborative control of multi-unmanned aerial vehicle systems and enhancing the communication and decision-making of unmanned aerial vehicles.

[0068] As Figure 1 described, the pluggable multi-unmanned aerial vehicle reinforcement learning module of the present invention includes an information-level weight network, an adaptive generation network, a Transformer-based decoder, a hybrid network, unmanned aerial vehicles, and an environment.

[0069] The information-level weight network 2 is used to receive the local observations of the UAVs through the communication part of the network, generate information-level weights in the centralized training stage, and input the information-level weights into the adaptive generation network. The information-level weight network is also used to obtain the weights in the generation model through the communication part of the network, implement weighted extraction of information in the distributed execution stage, filter redundant information, and input the weighted-extracted information into the policy network of the UAVs.

[0070] The communication information input to the information-level weight network is the local observations of each UAV. After feature extraction, the local observations can be represented as a one-dimensional vector of length l, denoted as where represents the observation value. In the scenario of full communication, the complete information obtained by each UAV at time t is defined as where the symbol - represents all other UAVs except UAV i. The weighted information includes the time, target, and content of communication as well as the weight. The calculation formula of the weighted information is shown in formula (1):

[0071]

[0072] where, is the weighted information of UAV i at time t, is the information weight of UAV i at time t, represents the complete information obtained by UAV i at time t, represents the set of local observation values of other UAVs except UAV i at time t, ∈ t is white noise, is the normal distribution, and I is the identity matrix. The information-level weight network is jointly trained with the adaptive generation network in the centralized training stage to update the parameters together.

[0073] The adaptive generation network is used to receive the local observations of the UAVs through the communication part of the network, assist in generating information-level weights by the information-level weight network and conduct joint training, complete the communication information and output the generated global state. The global state is used as the input of the hybrid network part, which helps to enhance the expression ability of the global action value function. The adaptive generation network includes n input channels and 1 output channel, integrating the local observations of the UAVs into the communication process, ensuring scalability. The adaptive generation network updates the parameters using the MSE loss and GAN loss in the centralized training stage.

[0074] The hybrid network is used to aggregate the local action value functions of individual drones to form a global action value function, guiding the overall policy iteration of the multi-drone system. The input of the hybrid network is the local action value function of each drone and the global state output by the adaptive generation network part, and the output is the global action value function, evaluating the overall value of the multi-drone system taking a certain cooperative action in a certain state. The drone swarm interacts with the environment according to the global action value function and obtains rewards, and updates the parameters of the hybrid network based on the rewards.

[0075] As Figure 2 described, the adaptive generation network module of the present invention includes an information completion network and a global discriminator network, and the information completion network is connected to the information-level weight network.

[0076] The information completion network receives the local observations of the drones and the weighted information generated by the information-level weight network through the communication part of the network, completes the information and generates the global state. The information completion network consists of repeated one-dimensional convolutional residual blocks, uses U-Net for information completion modeling, and follows the encoder-decoder structure. Memory usage and computation time are reduced by using downsampling before further processing the information, and then the output is restored to the length of the global state through upsampling.

[0077] The role of the global discriminator network is to identify whether the global state is real or generated from local observations. The global discriminator network is based on a one-dimensional convolutional neural network, compressing the global state into a compact feature vector. The output of the network is fused through a connection layer, and the connection layer uses the sigmoid activation function to predict a continuous value between 0 and 1, representing the probability that the global state corresponds to the real situation rather than being caused by information completion.

[0078] Specifically, the steps to obtain the information weights through the information-level weight network are to obtain the input information through the communication part of the pluggable module, that is, the local observations of the drones; perform position encoding on the input information to convert the input data into a vector representation containing position information; repeat the position-encoded vector to convert it into the input format of the attention mechanism (including query, key, and value); the attention mechanism generates attention weights by calculating dot-product attention; where the query of the attention mechanism is the local observation The calculation formula is as shown in formula (2):

[0079]

[0080] where is the i-th query at time t, MLP is the multi-layer perceptron, and m is the number of times of repeated processing

[0081] The local observations of all drones are used as keys and values, and the calculation formula is as shown in formula (3):

[0082]

[0083] Among them, M t represents the set of local observations of all UAVs at time t.

[0084] The calculation formula of the information-level weight is shown in formula (4):

[0085]

[0086] Among them, d is the dimension of the query vector and the key vector, σ is the ReLU activation function, and softmax is a normalization method. is the information weight of UAV i at time t.

[0087] Specifically, the steps to generate the global state through the adaptive generation network are to obtain the input information from the communication part of the pluggable module, including the local observations and information-level weights of the UAVs; before further processing the information, perform preprocessing through downsampling to reduce memory occupancy and improve computational efficiency; restore the preprocessed information to the length of the global state through upsampling; input the weighted information into the information completion network to implement the function of information completion, and send the generated global state to the global discriminator network. The calculation formula of the weighted information is shown in formula (1), and the output global state is mathematically defined as:

[0088]

[0089] Among them, s t is the global state at time t generated by the information completion network, and n is the number of UAVs in the UAV swarm.

[0090] Specifically, the steps to determine whether the global state is real are to input information from the communication part of the adaptive network, including the generated global state; extract local features from the input information through a one-dimensional convolutional neural network, and then perform pooling to compress the global state into a compact feature vector; the output is fused through a cascade layer using the sigmoid activation function, and the probability that the output global state corresponds to the real situation rather than being generated by information completion is output, so as to evaluate the performance of the adaptive generation network and improve the performance of the adaptive generation network through generative adversarial training.

[0091] Specifically, the steps to obtain the approximate global action value function are to input information from the communication part of the pluggable module, including the local action value function of each UAV and the generated global state; in the centralized training stage, all UAVs learn through the hybrid network Q tot$Q(\tau, a, m, s; \theta)$, using the additive structure of the QMIX network, determines the weights of the local action value function using the global state, and synthesizes the local action value functions of the UAVs into the global action value function $Q$ through the determined weights. tot $(s, a)$, where the parameter $\theta$ is updated by minimizing the expected temporal difference (TD) error:

[0092]

[0093] where, is the TD error, is the target Q value, $Q$ tot $(\tau$ t , $\alpha$ t , $m$ t , $s$ t ; $\theta)$ is the Q value of the mixing network at time $t$, $\tau$ t is the historical observation at time $t$, $\alpha$ t is the action at time $t$, $m$ t is the information of the UAV swarm at time $t$, $s$ t is the global state at time $t$.

[0094] The calculation process of is shown in Equation (7):

[0095]

[0096] where, $\theta$ - is the parameter of the target mixing network, $r$ is the reward obtained by executing the action, $\gamma$ is the discount factor, represents maximizing the Q t+1 value of action $a$ tot .

[0097] Specifically, the steps to obtain the approximate local action value function are to input information from the communication part of the pluggable module, including the local observations of the UAVs; in the distributed execution stage, the UAVs learn a Q network $Q$ i $(\tau, a, m, s)$, thereby estimating the action value function $Q$ i $(s, a)$.

[0098] Specifically, in the process of simultaneously training the adaptive generation network and the information-level weight network and updating their parameters, the present invention uses two loss functions: the MSE loss for stability and the GAN loss for enhancing the authenticity of the results. Combining the two loss functions, the generation network and the discriminative network are trained adversarially, and the optimization method becomes:

[0099]

[0100] where, is the second loss function, ‖.‖ is the Euclidean norm, and o t is the actual local observation, and O(s t ) is the local observation value predicted according to the global state, D is the global discriminator network in the adaptive generation network, G is the information completion network in the adaptive generation network, E represents calculating the expectation, o ~ O means o is sampled from the distribution O, s ~ S means s is sampled from the distribution S, and s t is the global state at time t, and D(s t ) is the probability of judging whether the global state s t input to the global discriminator network is true, and o t is the local observation value at time t, and α is the weight hyperparameter. t is the local observation value at time t, and α is the weight hyperparameter.

[0101] Embodiment 2:

[0102] As Figure 3 , Figure 4 described, the multi-rotor UAV is a six-rotor industrial UAV with a high-stability and high-expandability full carbon fiber fuselage. In this embodiment, this multi-rotor UAV is used as the UAV and the technical solution of the present invention is used as a premise for implementation.

[0103] The pluggable multi-UAV reinforcement learning module will distinguish two task execution situations for the multi-rotor UAV according to different execution strategies, providing support for the full-process logic analysis and diagnosis. The specific situations are as follows:

[0104] Situation 1: When the multi-rotor UAV needs to execute a task that can access the global state

[0105] Situation 2: When the multi-rotor UAV needs to execute a task that cannot access the global state.

[0106] As Figure 4 shown in a of Figure 4 , b of

[0107] is the flowchart of the task that the multi-rotor UAV needs to execute to access the global state, and

[0108] is the flowchart of the task that the multi-rotor UAV needs to execute without accessing the global state. When the multi-rotor UAV faces Situation 1, the specific steps for executing related tasks through the pluggable module and the multi-UAV reinforcement learning method include:

[0109] Step S101: Build an integrated equipment platform including the UAV control system, and arrange sensors and a pluggable multi-UAV reinforcement learning network module for communication information completion; Step S102: The UAV control system issues a collaborative task instruction for the multi-rotor UAV cluster to the multi-UAV system;Step S103: Input the task instructions of the UAV control system and the data captured by the sensors into the pluggable module through the wireless receiving device, and analyze the data by the pluggable module;

[0110] Step S104: Process all the input information through the information-level weight network; Take the processed data and the global state of the system as inputs and feed them into the policy network of the UAV. Determine the actions of the current multi-rotor UAV cluster according to the policy function, and send the action policy to the UAV control system;

[0111] Step S105: The UAV control system receives the action policy sent by the pluggable module and conducts an evaluation and analysis. If the execution actions of the current multi-rotor UAV cluster meet the requirements of the task instructions issued by the UAV control system, it is default to allow the multi-rotor UAV cluster to execute the current actions; If the execution actions of the current multi-rotor UAV cluster do not meet the requirements of the task instructions issued by the UAV control system, the UAV control system will send an instruction to stop the actions of the current multi-rotor UAV cluster, and require the reinforcement learning network module to collect all the data again and repeat steps S3 to S5 until the execution actions of the multi-rotor UAV cluster meet the requirements of the task instructions issued by the UAV cluster control system, then stop this process.

[0112] When the multi-rotor UAV faces situation 2, the specific steps for performing related tasks through the pluggable module and the multi-UAV reinforcement learning method include:

[0113] Step S201: Build an integrated equipment platform including the UAV control system, and arrange sensors and a pluggable multi-UAV reinforcement learning network module for communication information completion;

[0114] Step S202: The UAV control system issues collaborative task instructions for the multi-rotor UAV cluster to the multi-UAV system;

[0115] Step S203: Input the task instructions of the UAV control system and the data captured by the sensors into the pluggable module through the wireless receiving device, and analyze the current situation by the pluggable module;

[0116] Step S204: Process all the input information through the information-level weight network, and take the processed data and the information-level weight as inputs and feed them into the adaptive generation network; Complete information complementation through the adaptive generation network and generate the global state of the current multi-rotor UAV cluster. Take the generated global state and the input information as inputs and feed them into the policy network of the UAV. Determine the actions of the current multi-rotor UAV cluster according to the policy function, and send the action policy to the UAV control system;

[0117] Step S205: The UAV control system receives the action strategy sent by the pluggable module for evaluation and analysis. If the actions performed by the current multi-rotor UAV group meet the requirements of the task instructions sent by the UAV control system, the multi-rotor UAV group is allowed to perform the current actions by default; if the actions performed by the current multi-rotor UAV group do not meet the requirements of the task instructions sent by the UAV control system, the UAV control system will send an instruction to stop the actions of the current multi-rotor UAV group, and require the reinforcement learning network module to collect all data again and repeat steps S3 to S5 until the actions performed by the multi-rotor UAV group meet the requirements of the task instructions sent by the UAV group control system, then the process is stopped.

[0118] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0119] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A coordinated control method for a drone group based on communication information completion, characterized in that: The following steps are involved: Step S1: Each drone obtains local observation values ​​of each drone from the environment, and the local observation values ​​are obtained in real time through sensors set on the drone; Step S2: In the centralized training phase, based on the local observations, information weights are obtained through the information-level weight network; based on the information weights, weighted information of each drone is obtained; based on the weighted information of each drone, a global state is generated through an adaptive generative network; based on the global state and the local action value function of the drone, the local action value function of the drone is integrated through a hybrid network to obtain a global action value function; wherein the local action value function is calculated by each drone through an independent Q network based on the local observations; based on the global action value function, the drone swarm interacts with the environment to obtain a reward; based on the reward, the parameters of the hybrid network are updated through a first loss function; based on the information weights, local observations and global states, the parameters of the adaptive generative network and the information-level weight network are updated through a generative adversarial network loss function; Step S3: reacquire the information weight according to the updated information level weight network, and send the acquired information weight to each UAV; Step S4: In the distributed execution phase, each drone obtains the local action value function of the drone based on the local observation value and information weight through the Transformer-based decoder, and makes coordinated decisions based on the local action value function through the strategy network to guide the drone group to perform corresponding actions.

2. The coordinated control method of a drone group based on communication information completion according to claim 1 is characterized in that: The obtaining of information weights through an information level weight network based on the local observation value specifically includes: The local observation value of UAV i and information obtained from other drones Perform position encoding and convert it into a vector representation containing position information; The position-encoded vector is repeatedly processed and converted into the input format of the attention mechanism, including query, key and value, and the local observation value of drone i. The position encoding vector is repeatedly processed by the multi-layer perceptron MLP to generate the query vector Information obtained by other drones After processing, a key vector is generated Sum value vector Calculating the query vector and key vector The dot product of is used to obtain the attention score. The softmax function is applied to the attention score for normalization to obtain the information weight between each drone. The calculation formula is: Among them, M t represents the set of local observation values ​​of all UAVs at time t, is the information weight of drone i at time t, d is the dimension of query vector and key vector, σ is the ReLU activation function, and m is the number of repeated processing.

3. The method for coordinated control of a drone group based on communication information completion according to claim 1 is characterized in that: Based on the information weight, the weighted information of each drone is obtained, and the formula is: in, is the weighted information of UAV i at time t, is the information weight of UAV i at time t, represents the complete information obtained by UAV i at time t, represents the local observation value set of other UAVs except UAV i at time t, ∈ t is white noise, is a normal distribution and I is the identity matrix.

4. The method for coordinated control of a drone group based on communication information completion according to claim 1, characterized in that: The adaptive generation network includes an information completion network and a global discriminator network. The information completion network is connected to the information weight network. The information completion network adopts a U-Net structure and follows an encoder-decoder structure.

5. The method for coordinated control of a drone group based on communication information completion according to claim 1 or 4, characterized in that: The method of generating a global state through an adaptive generation network based on weighted information of each drone specifically includes: The weighted information of each drone Through downsampling, the downsampled data is input into the encoder and decoder of the information completion network, and the global state is output: Among them, s t is the global state at time t generated by the information completion network, and n is the number of drones in the drone swarm.

6. The method for coordinated control of a drone group based on communication information completion according to claim 1, characterized in that: The method of integrating the local action value function of the drone based on the global state and the local action value function of the drone through a hybrid network to obtain a global action value function specifically includes: All drones learn the hybrid network Q tot (τ, a, m, s; θ), using the additive structure of the QMIX network, the global state is used to determine the weight of the local action value function, and the local action value functions of each drone are synthesized into a global action value function through the determined weight of the local action value function.

7. The method for coordinated control of a drone group based on communication information completion according to claim 1, characterized in that: The first loss function is: in, is the first loss function, n is the number of drones, is the target Q value of drone i, Q tot (τ t ,α t ,m t ,s t ; θ) is the Q value of the hybrid network at time t, τ t is the historical observation value at time t, α t is the action at time t, m t is the information of the drone group at time t, s t is the global state at time t, θ is the parameter of the hybrid network, r is the reward obtained by the interaction between the drone swarm and the environment, γ is the discount factor, represents the maximization action a t+1 Q tot value.

8. The method for coordinated control of a drone group based on communication information completion according to claim 1, characterized in that: The generative adversarial network loss function is: in, is the second loss function, ‖.‖ is the Euclidean norm, o t is the actual local observation, O(s t ) is the local observation value predicted according to the global state, D is the global discriminator network in the adaptive generation network, G is the information completion network in the adaptive generation network, E represents the calculation expectation, o~O represents o is sampled from distribution O, s~S represents s is sampled from distribution S, s t is the global state at time t, D(s t ) is to convert the global state s t Input the global discriminator network output to judge the global state s t The probability of being true, o t is the local observation value at time t, and α is the weight hyperparameter.

9. The method for coordinated control of a drone group based on communication information completion according to claim 1, characterized in that: The method of obtaining the local action value function of the drone based on the local observation value and the information weight through a Transformer-based decoder specifically includes: According to the local observation value and information weight, the weighted information of the UAV is obtained, and the obtained weighted information and the hidden state of the UAV at time step t-1 are passed as input to the embedding layer. The embedding layer converts the input data into a continuous vector representation, and the output of the embedding layer is used as the input of the Transformer-based decoder. The Transformer-based decoder processes the input through the self-attention mechanism, extracts useful features from it, and outputs it as the hidden state of the UAV at time step t and its local action value function.

10. A drone group coordination control system based on communication information completion, characterized in that: include: The information-level weight network is used to process each UAV’s local observations with the information of other UAVs to generate information weights; An adaptive generation network, including an information completion network and a global discriminator network, wherein the information completion network is connected to the information level weight network, generates a global state through an encoder-decoder structure, and the global discriminator network is used to judge the authenticity of the global state; A hybrid network is used to integrate the local action-value functions of each drone and generate a global action-value function through an additive structure; A Transformer-based decoder is used to calculate the local action value function of the drone based on local observations and information weights during the distributed execution phase; The policy network is used to generate the final execution action based on the local action-value function and coordination decision; A loss function module, including a first loss function and a generative adversarial network loss function, for updating parameters of a hybrid network, an adaptive generative network, and an information-level weight network; Communication module, used to transmit weighted information and decision information between different UAVs; The environment interaction module is used to simulate the interaction process between the drone and the environment and collect reward signals.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster collaborative confrontation method combined with self-attention mechanism

    CN117608315A

  • Unmanned aerial vehicle cooperative game behavior decision-making method based on improved QMIX algorithm

    CN119356399A

  • Qmix reinforcement learning algorithm-based ship welding spots collaborative welding method using multiple manipulators

    WO2022095278A1