Multi-unmanned aerial vehicle confrontation task execution method and device based on model reinforcement learning, and medium
By employing a model-based reinforcement learning approach with distributed execution and adaptive weight optimization, the problems of independent execution and gradient imbalance in multi-UAV adversarial missions are addressed, enabling efficient policy training and sample utilization, and improving the performance of UAV adversarial missions.
Patent Information
- Application Number
- CN202511115048.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-25
AI Technical Summary
Existing multi-agent systems in UAV adversarial missions suffer from a centralized training-centralized execution mode that restricts independence, preventing the policy network from executing independently. Furthermore, existing model methods exhibit gradient imbalance in multi-objective optimization, leading to degraded policy performance and low sample efficiency.
We employ a model-based reinforcement learning approach under a distributed execution mode. We train the world model through autoregressive supervision, combine observation difference and adaptive weights to optimize the policy network for independent execution, and improve the policy learning process by refining the extraction of enemy and friend features.
It improves the efficiency of policy training samples for multi-UAV systems, alleviates the numerical dependency problem, optimizes the policy learning process, reduces invalid exploration, and improves the rationality of action output and sample efficiency.
Smart Images

Figure CN121009950A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-agent cooperative control technology, and in particular to a method, device, and medium for performing multi-UAV adversarial missions based on model-based reinforcement learning. Background Technology
[0002] In recent years, with the rapid development of reinforcement learning theory and methods, and leveraging the achievements of deep learning, the performance of reinforcement learning methods in high-dimensional, continuous state spaces has been significantly improved, making its application in the multi-agent domain possible. A typical multi-agent reinforcement learning process can be summarized as follows: each agent learns the optimal policy through interaction with the environment and other agents, leading to collaborative or competitive behaviors among a group of autonomous agents to jointly solve complex optimization problems. Compared to single-agent systems, multi-agent systems are significantly more complex and bring a series of new challenges, including environmental non-stationarity, continuous changes in agent policies, and the accompanying dynamic changes in the environment; partial observability, where limited observations of the environment and opponents by agents challenge their ability to adapt to policies; and sample efficiency, as the interaction data required for training reinforcement learning methods increases significantly with the expansion of the state and action spaces, necessitating more trial-and-error opportunities.
[0003] In multi-agent systems for UAV (Unmanned Aerial Vehicle) warfare, adversarial missions hold significant promise for real-world applications. UAV adversarial missions typically involve cooperative or adversarial operations between multiple UAVs, such as air combat between UAV swarms, target interception, reconnaissance, and counter-reconnaissance. From both security and economic perspectives, these missions need to minimize the interaction costs between UAVs and the real-world environment while mitigating the impact of adversary policy non-stationarity (i.e., changes in adversary policy over time) to achieve optimal adversarial game outcomes. To address these issues, model-based reinforcement learning methods have been introduced into multi-agent systems for UAV warfare. This method constructs a world model to learn the dynamic transition patterns of the UAV adversarial environment, such as UAV kinematics models, sensor detection models, and adversary behavior models. Based on the simulation of the environment using the world model, a reinforcement policy network is trained using simulated pseudo-data, thereby reducing the number of interactions required between UAVs and the real-world environment. This method not only saves on the costs and risks of actual flight testing but also provides an additional model containing the environmental operating rules. This model can be used for a range of additional tasks or requirements, such as boundary prediction (e.g., UAV flight boundaries, mission execution boundaries) and safety control (e.g., obstacle avoidance, collision avoidance). However, due to the limitations of complex multi-agent interaction environments, current work on applying model-based methods to multi-agent systems and multi-machine adversarial tasks is still in its early stages, and existing work typically has the following shortcomings: (1) In multi-agent tasks, to maintain the scalability and deployment flexibility of the system, it is usually necessary to satisfy the Centralized Training with Decentralized Execution (CTDE) paradigm. That is, the algorithm can introduce centralized information interaction during the training process, but assumes that multiple UAVs cannot communicate during the execution phase. The input, output and operation of the algorithm must be completely independent. When existing model-based methods are introduced into multi-agent tasks, information is usually first input into a centralized world model during the execution phase, and then the information is distributed to each UAV agent. This is actually a centralized training-centralized execution (CTCE) mode, which greatly simplifies the training process of the policy network, does not guarantee the independent information acquisition of agents, and limits the design and application of the policy network. Therefore, it is necessary to optimize the existing mode and adjust the information flow of multiple UAVs so that the policy network can execute independently without communication.
[0004] (2) In adversarial missions, the strategies of friendly and enemy UAV agents are often non-stationary and change with the strategies of other units, making static strategies less optimal and increasing the complexity of agent learning and exploration, thus leading to a decline in the performance of adversarial strategies. Therefore, it is necessary to further decompose and learn the behavioral patterns of the enemy and friendly agents for adversarial missions to achieve better strategy performance and more efficient strategy exploration.
[0005] (3) Existing model reinforcement methods typically employ multi-objective unified optimization during the optimization process, directly summing the losses of multiple optimization objectives and updating network parameters accordingly. The main optimization objective losses include distribution approximation loss, observation reconstruction loss, and reward prediction loss, all of which play a crucial role in the subsequent trajectory prediction extension quality of the world model. However, when using direct summation for unified optimization, the gradient magnitudes of these three factors are affected by the loss calculation method and the differences in the data itself. The optimizer tends to reduce the loss term with the largest gradient, usually the distribution approximation loss, while reducing the optimization intensity for the other two terms. This results in insufficient accuracy in the model's reconstruction of observations and rewards, affecting the efficiency of policy learning. Therefore, an adaptive weight allocation mechanism is needed to balance multiple optimization objectives, accelerate the optimization process of the world model, and improve the quality of predicted trajectories. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method, device, and medium for performing multi-UAV adversarial missions based on model-based reinforcement learning.
[0007] In a first aspect, embodiments of the present invention provide a method for performing multi-UAV adversarial tasks based on model-based reinforcement learning, the method comprising the following steps: The action policy network onboard the drone, at each time step, uses local observation information provided by the environment. As input, generate the drone joint action for the current step. It is then input into the environment for interaction, and local observation information for the next time step is obtained. By analogy, a complete interaction experience trajectory is obtained, which includes local observations during the interaction process. ,action and reward sequence Available action sequences End marker Store all interaction experience trajectories in the experience pool; Training the world model includes: randomly sampling interaction experience fragments from the experience pool as training samples, and supervising the world model in an autoregressive mode; Interactive experience trajectories are randomly sampled from the experience pool as baseline trajectories, and virtual interactive data is obtained by expanding the baseline trajectories. The action policy network and evaluation network are trained based on virtual interactive data. The trained action policy network is used to output action policies for multi-UAV adversarial missions.
[0008] Secondly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-described multi-UAV adversarial task execution method based on model reinforcement learning.
[0009] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described multi-UAV adversarial task execution method based on model reinforcement learning.
[0010] Fourthly, embodiments of the present invention provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-described multi-UAV adversarial task execution method based on model reinforcement learning.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a model-based reinforcement learning method for multi-UAV adversarial task execution, realizing a model-based reinforcement learning control method for multi-UAV systems in a distributed execution mode, significantly improving the sample efficiency of policy training. Through observation difference and adaptive weighting methods, the training process of the world model is improved, alleviating the problems of multi-task optimization and the numerical dependence of the world model. For multi-UAV adversarial tasks, an additional method for refining and extracting enemy and friendly features is added to the policy network, further optimizing the policy learning process, making action outputs more rational, reducing invalid exploration, and thus further improving sample efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is an overall schematic diagram of a multi-UAV adversarial task execution method based on model-based reinforcement learning provided by an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the construction and operation process of the world model in the method provided by the present invention; Figure 3 This is a schematic diagram illustrating the construction and operation process of the policy network in the method provided by the present invention; Figure 4 This is a comparison of the actual performance verification of the method provided by this invention in the SMAC multi-agent adversarial environment with benchmark methods and existing mainstream methods; Figure 5 This is a comparison of the errors caused by different historical lengths in offline testing using the method provided by this invention; Figure 6 This is a comparison of ablation performance tests using historical information improvements in the method provided by this invention; Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.
[0016] like Figure 1 , Figure 2 and Figure 3 As shown, this embodiment of the invention provides a method for performing multi-UAV adversarial tasks based on model-based reinforcement learning. The method includes the following steps: Step S1: Collect real-world policy interaction and experience data: Action policy network carried by the drone agent. At each time step, based on the local observation information provided by the environment... As input, a specific action probability distribution is given and random sampling is performed to generate the UAV joint action for the current step. Input into the environment and obtain local observation information for the next time step. This process continues until the end of the round or the maximum number of interactions is reached, thus obtaining a complete interaction experience trajectory. The content of the interaction experience trajectory includes local observations during the interaction process. ,action and the reward sequence not provided to the agent during the interaction process. Available action sequences End marker The collection of all the above sequence data is called the experience trajectory. After the interaction is completed, it is stored in the experience pool exp buffer for subsequent model training.
[0017] Furthermore, the local observation information This includes: the drone's own status (health, location information), the status of friendly drones within its field of view (health, location), and the status of enemy drones within its field of view (health, location); the drone's coordinated actions. This includes movement actions (stop, move) and attack actions (selecting enemy units within the attack range).
[0018] Step S2, training the world model, includes: randomly sampling experience fragments from the experience pool as training samples, and supervising the world model in an autoregressive mode.
[0019] The world model consists of a core recurrent state space model (RSSM) and multiple predictor components, expressed as follows: As shown in the formula, training the cyclic state-space model in the world model requires iterating the posterior model and the prior model sequentially. Therefore, the training process needs to iterate Length times on the sampling trajectory.
[0020] Furthermore, the world model uses the Transformer attention module and the Gated Recurrent Unit (GRU) as its core components, and the Classification Variational Autoencoder (CVAE) as its main model construction mode. The main components of the model include the RSSM model (containing the transition model, prior distribution model, and posterior distribution model) as well as the encoder and decoder (including the observation encoder / decoder, reward decoder, available action decoder, and stopping decoder). The transition model uses an MLP layer to encode the input data, a Transformer Encoder to extract multi-agent features, and a GRU unit as the core to process temporal information features and obtain a compactly represented hidden state vector. The pre- and posterior distribution models are primarily composed of MLP networks, with output layers using discrete multi-class distributions and one-hot encoding for efficient representation of the state transition process. Unimix noise is used during training to enhance model robustness and generation diversity. A pass-through sampling method is used for the noisy probability distribution to avoid gradient interruptions caused by random sampling. Both the encoder and decoder primarily use MLP networks, with output layers using different distributions based on data distribution characteristics. The reward decoder directly outputs the ground truth, the stopping decoder uses a Bernoulli distribution, and the available action decoder uses multi-class outputs and one-hot encoding. The specific structure and key parameters of these networks are shown in Table 1.
[0021] Table 1 Specifically, the detailed training process and information flow of the world model are as follows: 1) Encoding local observation information: Perform backward differencing on the observation information in the interactive experience trajectory, and then encode the differencing observations. With the original splicing, as input for expanded observational information Subsequently, the augmented observation information is encoded and embedded. The observation encoder is implemented using a two-layer MLP to obtain the observation embedding vector. ; 2) The output of the loop model includes the hidden states from the previous step (including random hidden states). With determining the hidden state (Integrated actions with drone swarms as input) The output is the determined hidden state for the next step. The actual implementation of the cyclic model is a serial structure, which includes an embedding module composed of MLP, an agent-dimensional feature interaction module composed of Transformer Encoder, and a cyclic unit composed of GRU, thereby realizing the extraction of temporal features of trajectory data in the time dimension and the extraction of multi-agent features in the agent dimension, and completing the spatiotemporal modeling of environmental interaction dynamics.
[0022] 3. Prior and posterior model hidden state output: The prior and posterior models are designed similarly, both using a two-layer MLP. During output, the logits values are converted into a OneHotCategorical probability distribution, i.e., a multi-class probability distribution using OneHot encoding, thus achieving efficient and compact encoding of the input information. Here, the input to the prior model is the deterministic hidden state calculated in the previous step. The input to the posterior model is the observation embedding vector. With deterministic hidden states Therefore, the posterior model is directly visible to the observation information, while the prior model can only accept the hidden state input and lacks the current actual observation information. Thus, inference is performed based on the prior conditions, which is consistent with the situation in actual trajectory prediction where there is no observation input. Subsequently, the prior model can only be used for subsequent prediction inference. When the posterior and prior models output, pass-through sampling is used to prevent gradient interruption caused by random sampling.
[0023] 4. After completing the above process, the hidden state at each moment can be obtained by recursion along the interaction experience trajectory. By concatenating them, we can obtain the feature vector at each time step. This information is used as input for each predictor. The predictors are similarly designed, each consisting of a two-layer MLP, with outputs that are probability distributions, using different distribution types depending on the data type. Specifically, the observation decoder and reward prediction use the MLP's logits directly as output, the action prediction uses a OneHot encoded multi-class distribution, and the cutoff prediction uses a Bernoulli distribution.
[0024] 5. The loss function of the world model is as follows: Losses include: observation reconstruction losses Reward prediction loss Successive and subsequent distribution difference loss Termination predictor loss Available action predictor loss Among them, the observation reconstruction loss and reward prediction loss both adopt smooth L1 loss, the prior and subsequent distribution difference loss adopts balanced KL divergence loss, and the termination predictor loss and available action predictor loss both adopt negative logarithmic loss.
[0025] Specifically, for the prior and posterior models, a balanced Kullback-Leibler Divergence loss is used to make the probability distribution of the prior output approximate the probability distribution of the posterior model, and a balance coefficient constraint is added to reduce the deviation between the two. For the training of each detector, the observation reconstructor and reward predictor use smooth L1 loss to make their reconstructed or predicted values approximate the true trajectory. Since the available action prediction and stop prediction are probability distribution outputs, a negative logarithmic loss is used to calculate the probability distribution difference, making the probability distribution approximate the distribution in the sampled trajectory. Finally, the above losses are adjusted by adaptive coefficients and then accumulated. The optimizer performs gradient backpropagation and parameter optimization on the entire world model.
[0026] Furthermore, considering the large number of loss terms in the model, their varying calculation methods, and differing numerical values, a dynamic adaptive coefficient is proposed to better optimize the entire world model. This coefficient determines the weight of each loss term, balances their contribution to parameter optimization, and prevents differences in optimization strength caused by numerical factors. The reconstruction loss weights are defined. Reward prediction loss weight Successive and subsequent distribution difference loss After adaptation, the calculation of these three losses is as follows: Repeat the above training process Second-rate.
[0027] Furthermore, step S2 also includes: validating the trained world model, specifically including: Training loss of the world model This doesn't directly reflect the predictive ability of the world model, so we consider adding an additional validation dataset and calculating validation errors. The experience pool used in validation is similar to the online experience pool, but the difference is that the validation experience pool is a complete offline experience pool retained from other training rounds. The experience in this pool follows the same distribution or interaction rules as the online experience, but it is not numerically identical to the online experience in the current training process, and its distribution in the state space is wider. Therefore, it can be used to validate the model's predictive performance and generalization ability under the same distribution. Similar to the model's training process, the network parameters are frozen on the validation set, and the same information input and loss function are used for prediction. The validation dataset is used to calculate the validation error. The magnitude of the prediction loss reflects the model's predictive performance and its ability to adapt to unseen data.
[0028] Furthermore, the training process of the world model also includes: In this example, considering the weak numerical dependence and generalization of the world model's learning process, observation difference information is introduced to enhance the model's predictive ability. During the training process of the world model, since the physical meaning of the observation information dimension is clear, backward differencing is performed on the training experience trajectory. Use all zeros to fill the starting point. This directly introduces explicit state change quantities, reducing the difficulty of model learning and prediction, and reducing the absolute numerical dependence of the model, thus enabling better learning of environmental transition patterns.
[0029] Step S3, Generating Virtual Interaction Data: To address the issue that reinforcement learning networks require a large amount of interaction experience to gradually optimize their strategies, and considering the high interaction costs and difficulty in obtaining experience samples in adversarial tasks, to reduce the number of trial-and-error interactions between friendly and enemy UAV agents, the interaction cost is transferred to the world model. Based on existing interaction data, the world model engages in virtual interaction with the friendly agent's strategy. According to the learned state transition rules and the joint actions given by the friendly multi-UAV agent swarm, it predicts the future interaction trajectory after the action, thereby significantly expanding the existing limited experience data using the predictive power of the world model. A sampling length of [missing information] is taken from the experience pool. The baseline trajectory, and decomposed into The world model first preloads short fragments as historical data, uses the original data as the initial state, and then uses the existing policy network. action The output is used as a state transition condition to iteratively predict future trajectories, thereby expanding the data to sampled data. The system then interacts with the policy network frame by frame and outputs future trajectory predictions, i.e., future virtual interaction data sequences of multiple enemy and friendly UAV units. This expands the interactive experience and significantly reduces the need for interaction with the real environment. In this example, the actions output by the policy network in the virtual interaction are also retained. , The hidden state of a network The expanded virtual interaction data, along with the original data, will be used to train the policy network and evaluation network of the drone agent.
[0030] Step S4: Train the action policy network and evaluation network based on virtual interactive data. The trained action policy network is used to output action policies for multi-UAV adversarial missions.
[0031] The process of training the action policy network and evaluation network based on virtual interaction data includes: The system uses observation reconstruction, reward prediction, and available action prediction from the world model's predicted output as input to maintain consistency with local observation inputs during real-world interaction. This reduces distributional bias while preserving the centralized training-distributed execution mode of the multi-agent system, making it more flexible and scalable. Theoretically, this pre-built world model and data expansion method can interface with any multi-agent reinforcement network. The action network design uses a serial structure of GRU and MLP to process temporal information input, reusing parameters for each UAV agent while maintaining independent local observation inputs. For training the action network and evaluation network, the Lambda-Return cumulative reward is first calculated as the value evaluation of the current state of the UAV agent: Subsequently, the value network adopts a joint state based on the world model. As input, the output fits the cumulative return, and the regression loss is calculated using smooth-l1: The action network loss calculation uses a proximal optimization model, and the loss function is as follows: In the formula, This represents the λ-discounted return value. Represents the joint state of the world model. This represents the reward value at time t. Indicates the discount factor. Indicates the weighting factor; Represents local observation information. This represents the probability distribution of actions of the new policy network under the current prediction. This represents the probability distribution of actions of the old policy network under the current prediction. ε Let T represent the hyperparameters and T represent the time series.
[0032] The process by which the trained action policy network outputs action policies to execute multi-UAV adversarial tasks includes: For multi-agent adversarial tasks, an observation input decomposition and action output decomposition design was added to the action network. (Flowchart provided) Figure 3 After data augmentation, the world model's predicted local observation information still possesses clear physical meaning. This characteristic allows for further refinement of the input observations and enables additional modeling and action optimization of adversary and friendly features in multi-agent adversarial tasks. Considering the universality of this method in adversarial tasks, the observation information is categorized into its own features based on the meaning of the observation dimensions. (Location information, machine status), enemy characteristics (Enemy relative position, enemy unit status information), friendly characteristics (Relative position of friendly units, status information of friendly units), Unit number The four types of information mentioned above are all common basic information in adversarial tasks, therefore this method can be applied to most adversarial tasks. For the relatively fixed self-features and identification information, a one-layer fully connected (FC) network and a one-layer embedding network are used for feature mapping, respectively. The more important enemy and friendly features require special handling; the former uses a supernetwork composed of two MLP layers. The mapping is performed, and the output includes enemy embedded feature weights. Attack weight actions Attack action bias Multiple components: in The weights and biases of the last two terms are used for enemy feature extraction calculations, and are then used for subsequent attack action calculations. Similarly, friendly features are processed in a similar way, and the hypernetwork output includes the weights of the friendly embedded features. Weight of friendly interaction actions Friendly interaction action bias If there is no interaction with friendly units, the latter two items are ignored. The action network calculation process is based on the stages of input and output, revolving around the core GRU module. The input calculation is (time-independent): The core hidden state iteration process is (time-dependent): The output is calculated as follows (time-independent): In the formula, act base Indicates a basic action, act attack Indicates an attack action, act aid Indicates an interactive action.
[0033] Considering that the additional supernetworks introduced here are all implemented by MLP and the information is independent of the actual time sequence, the action network is implemented using batch computation. During training, only the core GRU module is cyclically computed, thereby greatly reducing the amount of training computation.
[0034] In this example, the training environment uses the StarCraft Multi-Agent Challenge (SMAC), a complex and challenging multi-agent adversarial environment based on the popular real-time strategy game StarCraft II, used to test and evaluate the performance of multi-agent systems. The SMAC environment simulates cooperative combat and confrontation between multiple agents, covering complex issues such as strategy planning, resource allocation, and dynamic decision-making, making it a classic platform for researching multi-agent reinforcement learning. Different map settings utilize different combinations of unit types to compete against the environment's built-in combat strategies. Because the SMAC environment is highly similar to multi-UAV adversarial environments in terms of complexity, dynamics, and task diversity, multi-agent algorithms trained and tested on SMAC (such as model-based reinforcement learning methods) can provide important theoretical and technical support for UAV adversarial tasks. By validating the effectiveness of the algorithms in the SMAC environment, these algorithms can be further transferred and optimized to address the practical challenges in multi-UAV adversarial tasks. Furthermore, considering the need for fair horizontal comparison of algorithm performance, comparing the method with similar multi-agent methods using a publicly available testing environment better illustrates the effectiveness and advancement of the method presented in this invention.
[0035] In the above experiments, the selected baseline methods include, but are not limited to, the following: multi-agent model reinforcement learning methods include the current state-of-the-art algorithm MARIE and the typical mainstream algorithm MAMBA; mainstream multi-agent model-free reinforcement learning algorithms include MAPPO, QMIX, QPLEX, and MBVD; among them, MARIE, MAPPO, and MBVD algorithms are strict CTDE modes, while the rest are CTCE modes or CTDE modes under relaxed constraints (which may include global state).
[0036] The key experimental parameters were set as follows: world model learning rate 0.0005, AdamW optimizer used, gradient clipping 5, and training epochs. Training sequence length Training batch size Experience pool capacity 250,000, minimum capacity 2,000, prediction length Historical data length The policy network uses the AdamW optimizer. The learning rate is 0.0005, and the agent's experience pool capacity is 80 virtual trajectories. The policy network entropy loss weight is 0.003, and the decay rate is 0.0005.
[0037] For a detailed comparison of the experimental and baseline results on different maps, please refer to Table 2 and... Figure 4This paper includes a comparison with the state-of-the-art (SOTA) multi-agent algorithm MARIE, which uses model reinforcement within the CTDE (Computer-to-Demand) framework. In the SMAC (Simplified Chinese Multi-Agent Mapping) environment, the experimental metric is the average win rate of the relevant methods in the last 20K steps of training with a fixed training step size. Our experimental results show that the convergence performance and sample efficiency of our invention are significantly better than the benchmark model and similar state-of-the-art mainstream multi-agent algorithms across multiple different SMAC tasks.
[0038] Table 2 As can be seen from the specific win rate comparison in Table 2, under the constraint of a limited number of interaction steps, the method with model enhancement significantly outperforms the method without model enhancement, demonstrating significantly stronger sample efficiency. Specifically, combining... Figure 4 The training curves show that in relatively simple environments, all methods learn quickly with minimal differences in convergence performance. However, the learning speed of the method in this invention is significantly faster than the current state-of-the-art (SOTA) method, MARIE, demonstrating a significant improvement in sample efficiency. As the difficulty of adversarial environments increases (3s vs 3z and beyond), the performance gap between various methods becomes more pronounced. Model-free methods struggle to generate effective policies within a limited number of steps, exhibiting significant performance differences. The ineffective exploration of model-free methods necessitates extensive interactive trial and error to gradually find the optimization direction. In contrast, model-based methods save considerable interaction costs under simulated environmental interaction rules. On more difficult maps, the difference between the method in this invention and the SOTA method, MARIE, also gradually increases, fully demonstrating the effectiveness of the proposed improvements and further enhancing the sample efficiency of model-based methods.
[0039] To further illustrate the effectiveness of the method proposed in this invention, Figure 4 , Figure 5 , Figure 6 Related ablation experiments were conducted to verify this. Figure 4 As can be seen, after removing the relevant improvements in the policy network, both policy performance and sample efficiency significantly decrease in both simple and complex environments, and the stationarity of policy learning also significantly declines, resulting in greater fluctuations in win rate. Figure 5In the experiment, offline testing was used to compare the effectiveness of the improvements to the world model. The tests showed that as the historical data read by the world model gradually increased (len=1~4), the accuracy of the model's observation reconstruction significantly improved, and the loss was significantly lower than the original default prediction mode (len=1). This indicates that the original prediction method had a large error, which may have led to a decrease in the quality of the prediction data, thus resulting in a decrease in the final policy performance. Furthermore, the impact of the error on the reward prediction term did not decrease linearly. At len=4, the model's reward prediction error was actually larger, in some cases even exceeding the loss of the default case, and fluctuating significantly, making it unsuitable for practical use. Considering both data cost and error elimination effect, len=2 was the optimal choice. Figure 6 In actual testing, it can be seen that after adding historical information for prediction, the strategy performance and sample efficiency have been significantly improved, indicating that appropriately improving the prediction accuracy during the data generation process is effective for training under the observation reconstruction mode of this invention.
[0040] In summary, the multi-UAV adversarial task execution method based on model-based reinforcement learning proposed in this invention achieves significantly better sample efficiency than other multi-agent algorithms. It also demonstrates stronger convergence performance and maintains superior sample efficiency compared to mainstream model-based multi-agent reinforcement learning methods. Furthermore, the improvements to the policy network in this invention are transferable, decoupling the data format of the world model and the policy reinforcement network. This allows for the transfer of relevant patterns and improved structures to similarly designed multi-agent algorithms. The policy network also strictly conforms to the CTDE pattern, facilitating flexible algorithm deployment and providing a novel and effective design approach for solving multi-UAV adversarial tasks.
[0041] Accordingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the multi-UAV adversarial task execution method based on model reinforcement learning as described above. Figure 7 The diagram shown is a hardware structure diagram of any device with data processing capabilities for executing a multi-UAV adversarial task based on model reinforcement learning, as provided in an embodiment of the present invention. Except for... Figure 7 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0042] Accordingly, this application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the multi-UAV adversarial task execution method based on model reinforcement learning as described above. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0043] The purpose of the above embodiments is to illustrate the design concept and features of the present invention, so that those skilled in the art can understand and implement the present invention. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or improvements made based on the principles and design concepts of the present invention should be considered as part of the scope of protection of the present invention.
Claims
1. A method for performing multi-UAV adversarial missions based on model-based reinforcement learning, characterized in that, The method includes the following steps: The action policy network onboard the drone, at each time step, uses local observation information provided by the environment. As input, generate the drone joint action for the current step. It is then input into the environment for interaction, and local observation information for the next time step is obtained. By analogy, a complete interaction experience trajectory is obtained, which includes local observations during the interaction process. ,action and reward sequence Available action sequences End marker Store all interaction experience trajectories in the experience pool; Training the world model includes: randomly sampling interaction experience fragments from the experience pool as training samples, and supervising the world model in an autoregressive mode; Interactive experience trajectories are randomly sampled from the experience pool as baseline trajectories, and virtual interactive data is obtained by expanding the baseline trajectories. The action policy network and evaluation network are trained based on virtual interactive data. The trained action policy network is used to output action policies for multi-UAV adversarial missions.
2. The multi-UAV adversarial task execution method based on model-based reinforcement learning according to claim 1, characterized in that, The training process of the world model includes: Randomly sample several interactive experience fragments of length T from the experience pool. ; Based on interactive experience fragments This allows the world model to be trained in a supervised manner using an autoregressive pattern, including: the world model using local observation information. As input, the coordinated actions of our drone swarm Given the state transition conditions, we obtain the next time step. Predicting the interaction experience trajectory and comparing it with Error calculations are performed on the real-time interaction experience trajectories to obtain gradient information, which is then used to update the parameters of the world model, enabling the world model to learn the state transition interaction rules between friendly and enemy targets.
3. A method for performing multi-UAV adversarial missions based on model-based reinforcement learning according to claim 1 or 2, characterized in that, The training process of the world model specifically includes: Several interactive trajectories of length T are sampled from the experience pool; observation sequence Input the posterior model and output the stochastic latent state for each UAV agent. After being concatenated with the action sequence, it is input into the transition model to obtain intermediate features. ; intermediate features After inputting the prior model, the deterministic hidden state is obtained. ; The random latent states output by the posterior model The deterministic hidden state output by the prior model By concatenating the data, the predicted features are obtained. , which serve as the input to each predictor; A joint loss function is constructed, and the parameters of the world model are updated based on the gradient descent method according to the joint loss function. The joint loss function includes: observation reconstruction loss, reward prediction loss, a priori distribution difference loss, termination predictor loss, and available action predictor loss. The observation reconstruction loss and reward prediction loss both adopt smooth L1 loss, the a priori distribution difference loss adopts balanced KL divergence loss, and the termination predictor loss and available action predictor loss both adopt negative logarithmic loss.
4. The multi-UAV adversarial task execution method based on model-based reinforcement learning according to claim 3, characterized in that, The weight coefficients corresponding to the observation reconstruction loss, reward prediction loss, and prior and subsequent distribution difference loss are adjusted through dynamic adaptive coefficients.
5. The method for performing multi-UAV adversarial tasks based on model-based reinforcement learning according to claim 1, characterized in that, The process of randomly sampling interactive experience trajectories from the experience pool as baseline trajectories and expanding upon these baseline trajectories to obtain virtual interactive data includes: The sampling length from the experience pool is The actual interaction trajectory is used as the starting data, and it is decomposed into The sub-fragments are preloaded into the world model, and then interact with the policy network frame by frame to output future trajectory predictions. This allows us to obtain virtual interactive data.
6. The multi-UAV adversarial task execution method based on model-based reinforcement learning according to claim 1, characterized in that, The expression for the loss function used to train and evaluate the network is as follows: In the formula, This represents the λ-discounted return value. Represents the joint state of the world model. This represents the reward value at time t. Indicates the discount factor. Indicates the weighting factor; The expression for the loss function used to train the action policy network is as follows: In the formula, Represents local observation information. This represents the probability distribution of actions of the new policy network under the current prediction. This represents the probability distribution of actions of the old policy network under the current prediction. ε Let T represent the hyperparameters and T represent the time series.
7. The method for performing multi-UAV adversarial tasks based on model-based reinforcement learning according to claim 1, characterized in that, The process of using a trained action policy network to output action policies for multi-UAV adversarial missions includes: Classify observation information into its own characteristics Enemy characteristics Friendly characteristics Body number ; Regarding its own characteristics Body number Feature mapping is performed using one fully connected layer and one embedding layer, respectively. Characteristics of the enemy Supernetwork composed of two-layer multilayer perceptron (MLP) Perform mapping and output enemy embedded feature weights. Attack weight actions Attack action bias ; Friendly characteristics Supernetwork composed of two-layer multilayer perceptron (MLP) Perform mapping and output the weights of the friendly embedded features. Weight of friendly interaction actions Friendly interaction action bias ; The input to the action policy network is calculated as follows: The hidden state iteration process is as follows: The output is calculated as follows: In the formula, act base Indicates a basic action, act attack Indicates an attack action, act aid Indicates an interactive action.
8. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the multi-UAV adversarial mission execution method based on model reinforcement learning as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the multi-UAV adversarial mission execution method based on model reinforcement learning as described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the multi-UAV adversarial mission execution method based on model reinforcement learning as described in any one of claims 1-7.
Citation Information
Cited By
Multi-agent multi-task collaborative reinforcement learning method based on space-time fusion architecture
CN121859981A
Multi-agent multi-task collaborative reinforcement learning method based on space-time fusion architecture
CN121859981B