Unmanned aerial vehicle group coordination control method for implicit communication in offline scene

By adopting implicit communication and global state prediction methods in offline MARL, the problem of limited decision-making capabilities of drone groups in offline scenarios is solved, and more efficient strategy execution and task completion are achieved.

CN120223159APending Publication Date: 2025-06-27TONGJI UNIV

Patent Information

Application Number
CN202510361526.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing multi-agent reinforcement learning (MARL) algorithms face problems such as external distributed actions, partial observability and global state dependence in offline scenarios, resulting in insufficient strategy generalization capabilities and limited decision-making capabilities.

Method used

Using an implicit communication method for offline scenarios, local observation information is processed through the Transformer-based encoder-decoder architecture, the global state is inferred, and the timing modeling of local action value functions is combined with the multi-layer perceptron and gated loop unit to calculate the global action value function.

Benefits of technology

It realizes effective implicit communication and global state prediction of the drone swarm in offline environments, improves the optimization of policy execution and task completion rate, and enhances the collaborative decision-making capabilities of the drone swarm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223159A_ABST
    Figure CN120223159A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle group coordination control method for implicit communication in an offline scene, and the method comprises the steps: converting the local state and communication information of an unmanned aerial vehicle into a discrete index sequence through the sequence modeling based on a vocabulary in a centralized training stage, and predicting the global state through a Transform encoder-decoder structure, so as to enhance the information fusion capability. In the distributed execution stage, the unmanned aerial vehicle only depends on local observation and historical information, calculates a local action value function through a gating circulation unit in combination with a global state, optimizes a hybrid network by using a deep Q network, finally obtains an optimal global action value function, and realizes efficient cooperative decision making. According to the method, the information sharing capability of the unmanned aerial vehicle group can be enhanced without explicit communication, the decision stability is improved, the training cost and the safety risk of online execution are reduced, and the method is suitable for unmanned aerial vehicle group tasks in a communication limited or complex dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV control, and particularly relates to a method for coordinated control of UAV swarms for implicit communication in an offline scenario. Background Art

[0002] In recent years, Multi-Agent Reinforcement Learning (MARL) has become an important framework for solving complex cooperation problems in multi-agent systems and is widely applied in fields such as vehicle autonomous driving, UAV swarm cooperation, and smart grid control. Existing MARL algorithms mainly rely on the paradigm of Centralized Training with Decentralized Execution (CTDE). However, this paradigm has significant limitations in practical applications: In the centralized training phase, the algorithm usually assumes that agents can access the global state, but in the decentralized execution phase, agents can only rely on local states for decision-making, resulting in limited decision-making capabilities. This partial observability problem seriously affects the cooperation efficiency of agents in complex environments.

[0003] Although online MARL algorithms perform well in some controllable environments, they face double challenges of safety and efficiency in real-world scenarios. For example, online training requires a large amount of real-time interaction, which is not only costly but also may cause safety problems due to the uncertainty during the exploration process. To overcome these problems, offline MARL emerged. It learns from a fixed dataset, avoiding high-risk online exploration and significantly improving the scalability and safety of the algorithm. However, offline MARL still faces the following technical problems: Out-of-distribution action problem: In a multi-agent environment, the action space grows exponentially with the number of agents, resulting in policies being prone to selecting actions not covered in the dataset, leading to evaluation errors and performance degradation. Partial observability problem: Online MARL usually alleviates partial observability through explicit communication mechanisms, but it is difficult to directly apply these mechanisms in the offline scenario. Existing methods lack an effective implicit communication framework and it is difficult to achieve information complementarity among agents without the global state. Global state dependence: Traditional offline MARL algorithms rely on the global state during centralized training, while only local observations can be obtained during actual execution, resulting in insufficient policy generalization ability.

[0004] In the prior art, Chinese Patent CN113641192A discloses a path planning method for a swarm intelligence perception task of unmanned aerial vehicles (UAVs) based on reinforcement learning. By adding a multi-head attention mechanism and fitting the strategies of other UAVs in the actor-critic architecture, when a UAV makes a decision, it fully considers the states and strategies of other UAVs. When the data collection volume of a UAV is greater than the average level, an additional reward value is given to accelerate the task completion. When the paths of UAVs overlap, it is judged whether it belongs to cooperation or competition according to the signal point data volume, and their reward values are corrected accordingly to promote their cooperation. The n-step return temporal difference is used to calculate the target value of the critic network, making the UAVs more far-sighted. Finally, to enable the UAVs to better explore and maximize the data collection volume, a distributed architecture is used to add noises with different variances to the actions output by the decision-making networks of UAVs in different virtual scenarios. However, this patent is mainly for online training, and the UAVs need to share local observation values in real time and achieve cooperation through explicit communication (such as directly transmitting observation vectors), which is not feasible in actual offline scenarios. Offline MARL cannot obtain real-time data through environmental interaction, resulting in the failure of its communication mechanism. Although the multi-head attention is used to fuse observation information, an implicit state completion mechanism is not designed, and the agents still rely on local observations during distributed execution and cannot predict the global state, resulting in limited decision-making. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a UAV swarm coordination control method for implicit communication in an offline scenario.

[0006] The purpose of the present invention can be achieved by the following technical solutions: The present invention provides a UAV swarm coordination control method for implicit communication in an offline scenario, including the following steps: In the centralized training stage, the UAVs obtain local state information and communication information from the offline dataset; based on the local state and communication information, a source sequence is obtained by using vocabulary-based sequence modeling; the source sequence is input into the encoder, the encoder processes and fuses the information, the information processed by the encoder is input into the decoder, the decoder extracts features and predicts the global state; the local action value function is obtained from the offline dataset, the global state and the local action value function are input into the hybrid network, the global action value function is calculated, based on the global action value function, the actions and rewards are obtained from the offline dataset, and the parameters of the hybrid network are updated according to the rewards; In the distributed execution phase, the UAVs obtain the local state from the environment, obtain the communication information from other UAVs, based on the local state and communication information, obtain the global state through the encoder and decoder, obtain the hidden state of the previous time step from the gated recurrent unit, combine the global state and the hidden state, obtain the local action value function through the multi-layer perceptron and the gated recurrent unit, input the local action value functions of each UAV and the global state into the updated hybrid network to obtain the global action value function, and the UAV swarm takes actions according to the global action value function.

[0007] Further, the offline dataset includes: Local state information: the observation data of the UAV itself, including position, speed, attitude, sensor data, and environmental perception information; Communication information obtained from other UAVs: the observation information, historical decisions, and task status shared by other UAVs, as the input of implicit communication; Local action value function: the value of each UAV taking different actions in a specific state estimated by the policy network, used for decision optimization; The global action value function of the UAV swarm at each moment based on the local state information, the actions taken based on the global action value function, and the corresponding rewards.

[0008] Further, the process of obtaining the offline dataset is as follows: In the data collection phase, the UAV swarm interacts with the environment through different strategies to collect multiple trajectory information. Each trajectory information contains a series of actions, states, rewards, and the state of the next moment, specifically including: Obtain its own local state information through the UAV's sensors, navigation system, and environmental perception module, and obtain the observation information, historical decisions, and task status shared by other UAVs through the wireless communication network; calculate the local action value function using the policy network, and aggregate the local action value functions through the hybrid network to calculate the global action value function; select the optimal joint action based on the global action value function, and record the reward value after executing this action and the global state of the next time step; store the collected local state, communication information, global state, action, and reward data in the offline dataset.

[0009] Further, based on the local state and communication information, a source sequence is obtained by using vocabulary-based sequence modeling, specifically including: Set the size of the vocabulary , which is used to store the discretized values of the local state and communication information; Extract the local state and communication information of all UAVs from the offline dataset to form a data set D: Among them, Represents the local state information of UAV i at time step t, which represents the communication information received by UAV i at time step t through the communication network; According to the set vocabulary size , divide the data set D into intervals and calculate the threshold of each interval : wherein, is the discrete index of interval ; Arrange the local states and communication information of n UAVs at time step t in chronological order to obtain continuous values , for each continuous value , find the interval it belongs to, calculate the discrete index I, and construct the source sequence based on the discrete index I of the local state and communication information : wherein, represents the discrete local state index of the Nth dimension of the nth UAV at time step t, represents the discrete communication information index of the Nth dimension of the nth UAV at time step t.

[0010] Furthermore, the source sequence is input into the encoder. The encoder processes and fuses the information. The information processed by the encoder is input into the decoder, and the decoder extracts features and predicts the global state, specifically including: Input the source sequence into the Transformer-based encoder. The encoder extracts the correlation between different time steps through stacked multi-head self-attention layers. After attention calculation, generate the encoded fused information : wherein, is the encoder operation; Take the fused information output by the encoder as the input and input it into the Transformer-based decoder. The decoder combines the fused information of the current time step and the predicted global state of the previous time step for optimization calculation to generate the target sequence, where the target sequence consists of discrete indexes of the global states of multiple time steps; Through the discrete index in the target sequence, look up the corresponding interval threshold in the vocabulary, and use the interval midpoint approximation method to convert the discrete index into a continuous value to obtain the global state at time step t: Among them, is the i-th dimensional information of the global state at time step t, and are the interval thresholds corresponding to this index, respectively.

[0011] Furthermore, the loss function of the encoder and the decoder is: Among them, is the loss function of the encoder and the decoder, is the number of time steps, M is the dimension of the global state, and represent the i-th dimensional information of the global states at time steps t and t + 1, respectively, is the conditional probability of the encoder and the decoder predicting the global state controlled by the parameter , represents time step t and is the set of all local states and communication information of the drones; By minimizing this loss function, the parameters of the encoder and the decoder are optimized.

[0012] Furthermore, obtaining actions and rewards from the offline dataset based on the global action value function specifically includes: In the offline dataset, according to the global state and the global action value function , select the action with the highest action value: Extract the reward value after executing the action according to the task completion situation in the offline dataset, indicating the feedback of the drone swarm executing this action in this state: Among them, is the reward function value pre-stored in the offline dataset.

[0013] Furthermore, updating the parameters of the hybrid network according to the reward specifically includes: Calculate the target value based on the reward at time step t and the global action value function at the future time step t + 1: Among them, is the discount factor, which is used to control the influence of future rewards on the current decision, are possible actions for the next time step, is the global action value function in the global state The maximum estimate of all possible actions; Calculate the current global action value function With target value The error between them is: in, is the loss function of the hybrid network, is the total number of time steps in the centralized training phase; The gradient descent method is used to minimize the loss function of the hybrid network and optimize the parameters of the hybrid network.

[0014] Furthermore, the combination of the global state and the hidden state to obtain the local action value function through a multi-layer perceptron and a gated recurrent unit specifically includes: In the distributed execution stage, the global state of the current time step and the hidden state of the previous time step are combined and input into the multi-layer perceptron for feature extraction and preliminary estimation of the local action value function; the output of the multi-layer perceptron is used as input and passed to the gated recurrent unit, and the state is modeled and updated through the gating mechanism to further optimize the estimation of the local action value function; the local action value function of the current drone is obtained through the output of the gated recurrent unit; The local action value function is the expected reward value of the action selected by each drone according to the current strategy network under a given local state and communication information, reflecting the value of the action in the current state.

[0015] Furthermore, the hybrid network is a deep Q network, and the hybrid network is optimized by conservative strategy constraints.

[0016] Compared with the prior art, the present invention has the following advantages: (1) The present invention uses the Transformer encoder-decoder architecture to process local observation information and infer the global state through historical information. This technical means solves some observability problems, allowing drones to use historical experience to restore global information during the execution phase, thereby optimizing strategy execution and improving task completion rate.

[0017] (2) The present invention adopts a vocabulary discretization method to convert the local state and communication information of the drone into a discrete index sequence and input it into the encoder for processing. This method can reduce the computational complexity brought by continuous values ​​and enhance the effectiveness of information sharing, so that drones can establish an implicit communication mechanism even in an offline environment, thereby improving the collaborative decision-making ability of drone groups.

[0018] (3) The present invention combines a multi-layer perceptron (MLP) with a gated recurrent unit (GRU) to perform temporal modeling on the local action value function, enabling the drone to more stably estimate the long-term return values of each action. This method reduces noise interference and improves the accuracy of action value estimation, making the decision-making of the drone more reliable in complex environments.

[0019] (4) The present invention uses a deep Q-network (DQN) to calculate the global action value function and introduces a conservative policy constraint to prevent unstable updates of the policy on out-of-distribution data. This technical means effectively reduces the policy deviation problem caused by unseen actions in the offline dataset and improves the stability and safety of drone swarm control.

[0020] (5) The present invention is trained based on an offline dataset, enabling the drone to learn the optimal policy in a simulated environment without the need for frequent trial and error in actual tasks. This method not only reduces the loss risk during the exploration of the drone but also reduces the expensive training cost, making multi-agent reinforcement learning easier to deploy and apply.

[0021] (6) Implement sequence-to-sequence communication modeling and adopt a sliding window technique to solve the problem of the increasing length of the source sequence caused by the increase in time steps and the number of agents, reducing the space requirements related to the joint training of the communication and policy networks: The sequence-to-sequence communication model can be pre-trained and seamlessly integrated into various offline MARL algorithms. By using the sliding window technique, the length of the context token is fixed, thereby improving the scalability of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the coordinated control method for a drone swarm with implicit communication in the present invention; Figure 2 It is a schematic diagram of the global state acquisition method in the present invention; Figure 3 It is a system relationship diagram of the observed state communication module and method; Figure 4 It is a schematic flowchart of an example of the multi-agent reinforcement learning method for implicit communication in an offline scenario, DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Embodiment 1: This embodiment provides a method for coordinated control of an unmanned aerial vehicle (UAV) swarm for implicit communication in an offline scenario. This method includes two main parts: a centralized training stage and a distributed execution stage.

[0025] Step 1: Construction and acquisition of an offline dataset In this embodiment, the offline dataset contains the following key information: local state information, communication information, local action value function, global action value function, actions, and rewards. The process of obtaining the offline dataset is as follows: In the data collection phase, first, a UAV swarm consisting of 5 UAVs is deployed to perform collaborative tasks in a simulated environment. Each UAV is equipped with sensing devices such as a GPS positioning system, an inertial measurement unit, a lidar, and a camera to obtain local state information such as its own position, speed, and attitude. The UAVs share observation information, historical decisions, and task status through a wireless communication network to form communication information.

[0026] Each UAV calculates the local action value function through a pre-trained policy network. This function represents the expected return of performing different actions in the current state. The UAV swarm aggregates their respective local action value functions through a hybrid network, calculates the global action value function, and selects the optimal joint action based on this. After executing the selected action, the system records the obtained reward value and the global state at the next time step.

[0027] During the data collection process, the UAV swarm interacts with the environment using different strategies, including rule-based strategies, random exploration strategies, and pre-trained reinforcement learning strategies, to ensure that the collected data is diverse and representative. In this way, the system collects 1000 trajectory information, and each trajectory contains a state-action-reward sequence of 200 time steps, forming a complete offline dataset.

[0028] Step 2: Sequence modeling based on a vocabulary To effectively process the local state and communication information of UAVs, this embodiment adopts a sequence modeling method based on a vocabulary to discretize the continuous state space for subsequent sequence processing. The specific implementation steps are as follows: First, set the vocabulary size |V| = 1024 to store the discretized values of local state and communication information. Extract the local state and communication information of all UAVs from the offline dataset to form a data set D: where represents the local state information of UAV i at time step t, represents the communication information received by UAV i at time step t through the communication network; According to the set vocabulary size |V| = 1024, the data in each dimension of the data set D is divided into 1024 intervals, and the threshold of each interval is calculated. Among them, is the discrete index of the interval . The interval division adopts the equal-frequency binning method to ensure that the number of data points in each interval is approximately equal, thereby improving the efficiency of discretization.

[0029] Arrange the local states and communication information of 5 drones at time step t in chronological order to obtain a sequence of continuous values. For each continuous value, find the interval it belongs to and calculate the discrete index I. Construct the source sequence based on the discrete index I of the local state and communication information : Among them, represents the discrete local state index of the Nth dimension of the nth drone at time step t, represents the discrete communication information index of the Nth dimension of the nth drone at time step t.

[0030] Step 3: Global state prediction of the encoder-decoder architecture In this embodiment, an encoder-decoder architecture based on Transformer is adopted to predict the global state by processing the source sequence. The specific implementation steps are as follows: Input the source sequence into the Transformer-based encoder through a sliding window of length 32. The encoder consists of 6 stacked multi-head self-attention layers, each layer contains 8 attention heads, and the hidden layer dimension is 512. The encoder extracts the correlation between different time steps through the multi-head self-attention mechanism. After the attention calculation, the encoded fusion information is generated : Among them, is the encoder operation; Take the fusion information output by the encoder as the input, and input it into the Transformer-based decoder. The decoder combines the fusion information at the current time step and the predicted global state at the previous time step for optimization calculation to generate the target sequence, where the target sequence consists of the discrete indices of the global states at multiple time steps; Through the discrete index in the target sequence, look up the corresponding interval threshold in the vocabulary, and adopt the interval midpoint approximation method to convert the discrete index into a continuous value to obtain the global state at time step t: Among them, is the i-th dimensional information of the global state at time step t, and are the interval thresholds corresponding to this index respectively.

[0031] To optimize the parameters of the encoder and decoder, the loss function is defined as follows: Among them, is the loss function of the encoder and decoder, is the number of time steps, M is the dimension of the global state, and represent the i-th dimensional information of the global state at time steps t and t + 1 respectively, is the conditional probability of the encoder and decoder predicting the global state controlled by the parameter , represents time step t and is the set of all local state and communication information of the drones; Minimize the loss function through the Adam optimizer, set the learning rate to 0.0001, and the number of training iterations is 100,000 times to optimize the parameters of the encoder and decoder.

[0032] Step 4: Construction and training of the hybrid network In this embodiment, a hybrid network is used to aggregate the local action value functions of each drone, calculate the global action value function, and perform decision optimization based on this. The specific implementation steps are as follows: The hybrid network adopts a deep Q-network structure, including 3 fully connected layers, and the dimensions of the hidden layers are 256 and 128 respectively. The network input is the global state and the local action value functions of each drone, and the output is the global action value function.

[0033] In the offline dataset, according to the global state and the global action value function , select the action with the highest action value: According to the task completion situation in the offline dataset, extract the reward value after executing the action , indicating the feedback of the drone swarm executing this action in this state: Among them, is the reward function value pre-stored in the offline dataset.

[0034] Updating the parameters of the hybrid network according to the reward specifically includes: Reward at time step t and the global action-value function at the future time step t+1 to calculate the target value : where is the discount factor, used to control the influence of future rewards on the current decision, is the possible action at the next time step, is the maximum estimated value of the global action-value function for all possible actions in the global state ; Calculate the current global action-value function and the error between the target value The formula is: where is the loss function of the hybrid network, is the total number of time steps in the centralized training phase; Use the gradient descent method to minimize the loss function of the hybrid network and optimize the hybrid network parameters.

[0035] Step Five: Calculation of the local action-value function in the distributed execution phase In the distributed execution phase, the UAVs need to calculate the local action-value function based on local observations and communication information, and obtain the global action-value function through the hybrid network for decision-making. The specific implementation steps are as follows: Each UAV obtains the local state from the environment, including position, speed, attitude, and sensor data. At the same time, it obtains communication information from other UAVs through the communication network, including observation information, historical decisions, and task status.

[0036] Based on the local state and communication information, obtain the global state through the trained encoder and decoder in Step Two and Step Three. The parameters of the encoder and decoder have been optimized in the centralized training phase and remain fixed in the distributed execution phase.

[0037] Obtain the hidden state of the previous time step from the gated recurrent unit (GRU). The GRU unit contains 128 hidden neurons, used to capture temporal information. Combine the global state and the hidden state, and obtain the local action-value function through the multi-layer perceptron and the gated recurrent unit.

[0038] The multi-layer perceptron consists of 3 fully connected layers, and the dimensions of the hidden layers are 128 and 64 respectively. The input is the concatenation of the global state and the hidden state, and the output is the preliminary estimation of the local action value function. The output of the multi-layer perceptron is used as the input and passed to the GRU unit, which models the time series of the state and updates the information through the gating mechanism to further optimize the estimation of the local action value function.

[0039] The output of the GRU unit passes through a fully connected layer to generate the local action value function of the current drone. The local action value function represents the expected return value of the action selected according to the current policy network for each drone under the given local state and communication information, reflecting the value of the action in the current state.

[0040] Step 6: Decision-making and Actions in the Distributed Execution Phase After obtaining the local action value functions of each drone, the global action value function is calculated through the hybrid network and decisions are made based on this. The specific implementation steps are as follows: The local action value functions of each drone and the global state are input into the updated hybrid network to obtain the global action value function. The parameters of the hybrid network have been optimized in the centralized training phase and remain fixed in the distributed execution phase.

[0041] Based on the global action value function, select the action with the highest action value: The drone swarm acts according to the selected action Perform corresponding operations, including adjusting the flight speed, direction, altitude, etc. After performing the action, the drone obtains the new local state and communication information and enters the next decision-making cycle.

[0042] In the distributed execution phase, the drone swarm can complete complex tasks, such as area coverage, target tracking, and obstacle avoidance, in a coordinated manner based on local observations and implicit communication. Even in the case of limited or interrupted communication, the drone swarm can still maintain a certain degree of coordination ability and demonstrate strong robustness.

[0043] As a preferred technical solution, the implicit communication framework consists of four parts, specifically, including a sliding window, a Transformer-based decoder and encoder, and a feedback mechanism.

[0044] As a preferred technical solution, the sequence-to-sequence communication technology specifically includes the following steps: Model the sequence data with stacked self-attention layers with residual connections; embed positional encoding before the sequence is input into the encoder and decoder; at each self-attention layer, the model generates and outputs several embedded tokens; each token is mapped to keys, values, and queries using a linear transformation.

[0045] As a preferred technical solution, the local state is provided by the unmanned aerial vehicle (UAV). Specifically, at each time interval it is obtained from the environment.

[0046] As a preferred technical solution, the sequence modeling method is applied to the source sequence and the target sequence. Specifically, the source sequence consists of local states at [number] time steps, and the observations at each time step are serialized across all UAVs. The target sequence consists of global states at [number] time steps.

[0047] As a preferred technical solution, for the vocabulary building method, specifically, a discretization method is used to standardize the input and output formats. Each element in the sequence is regarded as a token, and a vocabulary is constructed separately for the source sequence and the target sequence.

[0048] As a preferred technical solution, the encoder is based on the Transformer architecture. Specifically, it includes two identical stacked layers, and each stacked layer is composed of two sub-layers. The first sub-layer is the multi-head self-attention mechanism, and the second sub-layer is the position-wise feed-forward network.

[0049] As a preferred technical solution, the decoder is based on the Transformer architecture. Specifically, it includes two stacked layers and additional sub-layers, which are inserted between the two sub-layers of the encoder. In this sub-layer, the query comes from the output of the previous decoder layer, and the keys and values come from the output of the entire encoder.

[0050] As a preferred technical solution, for the process of obtaining the local action value function, specifically, it is obtained by learning the independent Q-network of each UAV.

[0051] As a preferred technical solution, for the process of obtaining the global action value function, specifically, it is obtained by learning a hybrid network.

[0052] According to another aspect of the present invention, a unique sliding window mechanism is provided. Specifically, due to the inherent Markov property of O2S, that is, memorylessness, predicting the global state is only related to the local states at the current time step and the previous time step. Therefore, a sliding window with a default window size of 2 can be used to control the number of context tokens.

[0053] The present invention designs a flexible multi-UAV reinforcement learning architecture that simultaneously includes functions such as implicit communication, information completion, and decision-making for the offline MARL scenario, supports seamless integration into the offline multi-UAV reinforcement learning network, and provides support for the cooperative control of multi-UAV systems and enhancing the communication and decision-making of UAVs.

[0054] Such as Figure 1As described above, the implicit communication offline multi-UAV reinforcement learning method of the present invention includes UAVs, an O2S network, a feedback mechanism, a hybrid network, a multi-layer perceptron and a gated recurrent unit, a policy network, a communication network, and the O2S network is connected to the UAVs and the communication network.

[0055] The O2S network is used to receive local states and shared information from the UAVs and the communication network through the communication function of the system, and generate global states in the centralized training and distributed execution phases. The input of the O2S network is the source sequence, and the output is the target sequence. The predicted global state helps to learn the global action value function without the global state being known.

[0056] The feedback mechanism is used to correct state predictions to ensure scalability and accuracy. The input of the feedback is the predicted global state of the current time step at the previous time step, and the output is the corrected global state of the current time step. In the centralized training phase, at the current time step it is possible to predict the current global state and the global state of the next time step . At the next time step it is also possible to predict , using generated when as feedback to improve the accuracy of O2S prediction. Taking the average value of the two-step prediction as the final output to correct and improve the prediction.

[0057] The hybrid network is used to aggregate the local action value functions of each UAV to form a global action value function to guide the overall behavior strategy of the multi-UAV system. The input of the hybrid network is the local action value function of each UAV and the global state output by the O2S network, and the output is the global action value function, which can evaluate the overall value of the multi-UAV system taking a certain collaborative action in a certain state. The hybrid network can reflect the collaborative effect between UAVs and guide the update of UAV strategies.

[0058] The multi-layer perceptron and gated recurrent unit are used for feature extraction and updating hidden states. The input of the multi-layer perceptron and gated recurrent unit is the predicted global state and the hidden state of the previous time step, and the output is the local action value function of all actions, which can provide a basis for the policy network to select actions in the distributed execution phase.

[0059] The policy network is used to generate the local action value function of the UAV. The input of the policy network is the action value function of all actions, and the output is the local action value function of the UAV. Actions are obtained by sampling the local action value functions of all actions according to the policy, and a conservative policy constraint is used to avoid generating out-of-distribution actions.

[0060] Specifically, the Transformer-based sequence-to-sequence structure consists of an encoder and a decoder, and uses stacked self-attention layers with residual connections to model sequence data. Each self-attention layer model processes embedded vectors corresponding to the input tokens , generates and outputs embedded vectors . A linear transformation maps the tokens to keys , queries and values . The output of the token is calculated by using the weighted sum of the values , and the weights are obtained by calculating the normalized dot product of and other key-value pairs , as shown in Equation (1): Specifically, as Figure 2 shown, the steps to obtain the global state through the O2S network are as follows: the local state and shared information are obtained by the drone, the input information is used as the source sequence, the global state is used as the target sequence, position encoding and serialization representation are performed, and the source sequence is sent into the encoder and decoder through a sliding window to generate the global state.

[0061] Considering the high-dimensional problems of multi-drones and continuous states and observations, a discrete method is used to standardize the input and output. The elements in the sequence are regarded as tokens, and a vocabulary is constructed for the source sequence and the target sequence. The vocabulary maps the tokens to digital indices starting from 0. The sequence elements are usually continuous variables, and the quantile method is used to discretize the variables to construct a vocabulary of size . The offline dataset is divided into intervals according to the vocabulary size, and thresholds are calculated for each interval . The input data is compared with the threshold , and the discrete index of the dimension is calculated using Equation (3) : During the reconstruction process, the original data is approximated by the midpoint of the interval: In addition, special tokens are defined for the start part of the sequence, the end part of the sequence, and the padding tokens .

[0062] During the training phase, a specific token is added to the original output sequence and used as the input to the decoder. The loss function is defined as shown in Equation (4): wherein is a parameter of O2S, is the induced conditional probability.

[0063] O2S simulates the relationship between the local state and the global state, as shown in Equation (5): which is essentially used as the observation function of the inverse function. O2S inherently has memory. Therefore, the relationship between the local state and the global state can be modified as shown in Equation (6): The loss function can be modified to: The sliding window size is defaulted to 2. This fixes the token length of the context window in the O2S encoder to .

[0064] Specifically, to improve the global state prediction accuracy through the feedback mechanism, during distributed execution, the O2S encoder receives the local state of the drone and the messages of other drones at time , and the content of the next time step is filled with the token . Since the local state of the next time step cannot be accessed during distributed execution, the feedback mechanism is ineffective. However, during centralized training, the O2S encoder can receive the local state and the feedback mechanism works properly.

[0065] Specifically, the steps to obtain the global action value function by learning the hybrid network are as follows. During the centralized training phase, the drone swarm learns the hybrid network , wherein and , to approximate the global action value function . The parameter is learned by minimizing the expected temporal error, and the temporal error is: wherein, is the batch size sampled from the replay buffer .

[0066] Example 2: As Figure 3 , Figure 4 shown, the quadrotor drone is a high-stability and high-expandability quadrotor industrial drone with a fully polymer composite fuselage and equipped with multi-modal sensors. In this example, the quadrotor drone is used as the drone, the target search is used as the task, and the technical solution of the present invention is implemented on this premise.

[0067] The multi-UAV reinforcement learning module 1 with implicit communication will distinguish two task execution situations for the quadrotor UAV swarm according to different execution policies, providing support for the full-process logic analysis and diagnosis. The specific situations are as follows: Situation 1: When the quadrotor UAV swarm needs to execute a target search task that can access the global state; Situation 2: When the quadrotor UAV swarm needs to execute a target search task that cannot access the global state.

[0068] As Figure 4 described, (a) is the flowchart for the quadrotor UAV swarm to execute a target search task that can access the global state, and (b) is the flowchart for the quadrotor UAV swarm to execute a target search task that cannot access the global state. When the quadrotor UAV swarm faces Situation 1, the specific steps for performing related tasks through the implicit communication module and the multi-UAV reinforcement learning method are as follows: Step S101: Build an integrated equipment platform including the UAV swarm control system, and arrange multi-modal sensors and a multi-UAV reinforcement learning network module for implicit communication in the offline scenario; Step S102: The UAV swarm control system issues a collaborative target search task instruction for the quadrotor UAV cluster to the multi-UAV system; Step S103: Input the target search task instruction of the UAV swarm control system and the data captured by the multi-modal sensors into the MARL module through a wireless receiving device, and the MARL module analyzes the data; Step S104: Process all the input information through the observed O2S network; input the processed data and the global state of the system into the trained policy network, determine the search policy of the current quadrotor UAV cluster according to the policy function, and send the search policy to the UAV swarm control system; Step S105: The UAV swarm control system receives and evaluates the search policy issued by the implicit communication module. If the actions performed by the current quadrotor UAV swarm meet the requirements of the search task instruction issued by the UAV swarm control system, it is default to allow the quadrotor UAV swarm to execute the current action; if the actions performed by the current quadrotor UAV swarm do not meet the requirements of the search task instruction issued by the UAV swarm control system, the UAV swarm control system will send an instruction to stop the search policy of the current quadrotor UAV swarm, and require the module to collect all the data again and repeat Steps S3 to S5 until the quadrotor UAV swarm finds the task target issued by the UAV swarm control system, then stop this process.

[0069] When the quadrotor UAV faces Situation 2, the specific steps for performing related tasks through the implicit communication module and the multi-UAV reinforcement learning method are as follows: Step S201: Build an integrated equipment platform including a UAV swarm control system, and deploy multi-modal sensors and a multi-UAV reinforcement learning network module for implicit communication in an offline scenario; Step S202: The UAV swarm control system issues a collaborative target search task instruction for the quadrotor UAV swarm to the multi-UAV system; Step S303: Input the target search task instruction of the UAV swarm control system and the data captured by the multi-modal sensors into the MARL module through a wireless receiving device, and the MARL module analyzes the data; Step S204: Process all the input information through the O2S network, perform information prediction through the O2S network and generate the global state of the current quadrotor UAV swarm; input the predicted global state and the input information into the trained policy network, determine the search strategy of the current quadrotor UAV swarm according to the policy function, and send the search strategy to the UAV swarm control system; Step S205: The UAV swarm control system receives the search strategy sent by the implicit communication module for evaluation and analysis. If the actions performed by the current quadrotor UAV swarm meet the requirements of the search task instruction issued by the UAV swarm control system, it is default to allow the quadrotor UAV swarm to perform the current action; if the actions performed by the current quadrotor UAV swarm do not meet the requirements of the search task instruction issued by the UAV control system, the UAV swarm control system will send an instruction to stop the search strategy of the current quadrotor UAV swarm, and require the module to collect all the data again and repeat steps S3 to S5 until the quadrotor UAV swarm finds the task target issued by the UAV swarm control system, then stop this process.

[0070] Compared with the prior art, the present invention has the following beneficial effects: (1) The multi-UAV reinforcement learning method for implicit communication in an offline scenario of the present invention integrates a sequence-to-sequence communication model for communication in offline multi-UAV reinforcement learning. During the training process of the present invention, the global state can be predicted according to the local state, and the prediction accuracy can be improved through a feedback mechanism, thereby reducing the system's dependence on the real global state.

[0071] (2) The multi-UAV reinforcement learning method for implicit communication in an offline scenario of the present invention has a flexible design, enabling the present invention to be pre-trained and smoothly integrated into different offline multi-UAV reinforcement learning algorithms with the CTDE paradigm. By adopting a sliding window and fixing the context sequence length of the window, the problem of the token length increasing with the number of UAVs is solved, thereby improving the scalability of the method.

[0072] (3)The multi-UAV reinforcement learning method for implicit communication in the offline scenario of the present invention designs an O2S structure, and uses communication to solve the problem of the difference between the local policy function and the global policy function in the traditional behavior cloning method.

[0073] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0074] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A coordinated control method for a group of unmanned aerial vehicles with implicit communication in an offline scenario, characterized in that: The following steps are involved: In the centralized training phase, the UAV obtains local state information and communication information from the offline dataset; Based on the local state and communication information, the source sequence is obtained by vocabulary-based sequence modeling; the source sequence is input into the encoder, the encoder processes and fuses the information, and the information processed by the encoder is input into the decoder, the decoder extracts features and predicts the global state; the local action value function is obtained from the offline data set, the global state and the local action value function are input into the hybrid network, the global action value function is calculated, and based on the global action value function, the action and reward are obtained from the offline data set, and the hybrid network parameters are updated according to the reward; In the distributed execution stage, the drone obtains the local state from the environment and the communication information from other drones. Based on the local state and communication information, the drone obtains the global state through the encoder and decoder, obtains the hidden state of the previous time step from the gated recurrent unit, combines the global state and the hidden state, and obtains the local action value function through the multi-layer perceptron and the gated recurrent unit. The local action value function and the global state of each drone are input into the updated hybrid network to obtain the global action value function, and the drone swarm takes action according to the global action value function.

2. According to claim 1, a coordinated control method for a drone swarm with implicit communication in an offline scenario is characterized in that: The offline data set includes: Local state information: the drone’s own observation data, including position, speed, attitude, sensor data, and environmental perception information; Communication information obtained from other UAVs: observation information, historical decisions, and mission status shared by other UAVs as input for implicit communication; Local action value function: the value of each drone taking different actions in a specific state, estimated by the policy network, for decision optimization; The global action value function of the drone swarm based on the local state information at each moment, the actions taken based on the global action value function, and the corresponding rewards.

3. According to claim 2, a coordinated control method for a drone swarm with implicit communication in an offline scenario is characterized in that: The offline data set acquisition process is as follows: In the data collection phase, the drone swarm interacts with the environment through different strategies and collects multiple trajectory information. Each trajectory information contains a series of actions, states, rewards, and the next state at the next moment, including: The drone obtains its own local state information through its sensors, navigation system and environmental perception module, and obtains observation information, historical decisions and task status shared by other drones through the wireless communication network; uses the policy network to calculate the local action value function, and aggregates the local action value function through the hybrid network to calculate the global action value function; selects the optimal joint action based on the global action value function, and records the reward value after executing the action and the global state of the next time step; stores the collected local state, communication information, global state, action and reward data in the offline data set.

4. According to claim 1, a coordinated control method for a drone swarm with implicit communication in an offline scenario is characterized in that: The method of obtaining a source sequence based on the local state and the communication information by using a vocabulary-based sequence modeling specifically includes: Setting the vocabulary size , used to store discretized values ​​of local states and communication information; Extract the local states and communication information of all drones from the offline dataset to form a data set D: in, represents the local state information of UAV i at time step t, represents the communication information received by UAV i through the communication network at time step t; According to the set vocabulary size , divide the data set D into intervals and calculate the threshold for each interval : in, For interval Discrete index of ; Arrange the local states and communication information of n drones at time step t in chronological order to obtain continuous values , for each continuous value , find the interval to which it belongs, calculate the discrete index I, and construct the source sequence based on the discrete index I of the local state and communication information : in, represents the N-th dimension discretized local state index of the n-th UAV at time step t, Represents the N-th dimension discretized communication information index of the n-th UAV at time step t.

5. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1 or 4, characterized in that: The source sequence is input into the encoder, the encoder processes and fuses the information, the information processed by the encoder is input into the decoder, the decoder extracts features and predicts the global state, specifically including: The source sequence The Transformer-based encoder is input through the sliding window. The encoder extracts the correlation between different time steps through the stacked multi-head self-attention layer. After attention calculation, the encoded fusion information is generated : in, For encoder operation; The fusion information output by the encoder As input, the Transformer-based decoder is input, and the decoder combines the fusion information of the current time step and the predicted global state at the previous time step Perform optimization calculations to generate a target sequence, where the target sequence consists of discrete indexes of the global state of multiple time steps; Through the discrete index in the target sequence, the corresponding interval threshold is found in the vocabulary, and the discrete index is converted into a continuous value using the interval midpoint approximation method to obtain the global state of time step t: in, is the i-th dimension information of the global state at time step t, and are the interval thresholds corresponding to the index respectively.

6. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1, characterized in that: The loss function of the encoder and decoder is: in, is the loss function of the encoder and decoder, is the number of time steps, M is the dimension of the global state, and Represents the i-th dimension information of the global state at time step t and t+1 respectively, For the parameter The encoder and decoder of the control predict the conditional probability of the global state, represents the time step t and The collection of all drone local states and communication information; By minimizing this loss function, the parameters of the encoder and decoder are optimized.

7. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1, characterized in that: The steps of obtaining actions and rewards from offline data sets based on the global action value function specifically include: In the offline dataset, according to the global state and the global action-value function , select the action with the highest action value : Extract execution actions based on the task completion status in the offline dataset The reward value after , indicating the feedback of the drone swarm performing the action in this state: in, is the reward function value pre-stored in the offline dataset.

8. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1, characterized in that: The updating of the hybrid network parameters according to the reward specifically includes: Reward according to time step t And the global action value function at the future time step t+1 , calculate the target value : in, is a discount factor used to control the impact of future rewards on current decisions, are possible actions for the next time step, is the global action value function in the global state The maximum estimate of all possible actions; Calculate the current global action value function With target value The error between them is: in, is the loss function of the hybrid network, is the total number of time steps in the centralized training phase; The gradient descent method is used to minimize the loss function of the hybrid network and optimize the parameters of the hybrid network.

9. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1, characterized in that: The method combines the global state and the hidden state to obtain the local action value function through a multi-layer perceptron and a gated recurrent unit, specifically including: In the distributed execution stage, the global state of the current time step and the hidden state of the previous time step are combined and input into the multi-layer perceptron for feature extraction and preliminary estimation of the local action value function; the output of the multi-layer perceptron is used as input and passed to the gated recurrent unit, and the state is modeled and updated through the gating mechanism to further optimize the estimation of the local action value function; the local action value function of the current drone is obtained through the output of the gated recurrent unit; The local action value function is the expected reward value of the action selected by each drone according to the current strategy network under a given local state and communication information, reflecting the value of the action in the current state.

10. The method for coordinated control of a drone swarm with implicit communication in an offline scenario according to claim 1, characterized in that: The hybrid network is a deep Q network, and the hybrid network is optimized by conservative strategy constraints.

Citation Information

Patent Citations

  • Path planning method for unmanned aerial vehicle crowd sensing task based on reinforcement learning

    CN113641192A

Cited By

  • Cluster control method and system for underwater unmanned equipment

    CN121433263A