一种基于反事实基线的无人机集群对抗博弈仿真方法

By using a simulation method for adversarial game in drone swarms based on counterfactual baselines, and optimizing agent policies through evaluation and action networks, the problem of solving Nash equilibrium in drone swarm adversarial games is solved, achieving more efficient policy learning and faster convergence of the reward function.

CN116136945BActive Publication Date: 2026-07-17SHENYANG AEROSPACE UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG AEROSPACE UNIVERSITY
Filing Date
2023-02-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively solve the problem of finding Nash equilibrium in drone swarm warfare, and the communication and cooperation among multiple agents are highly complex.

Method used

A simulation method for adversarial game of UAV swarms based on counterfactual baselines is adopted. By setting up a combat data replay buffer and counterfactual baseline policy gradient update, the agent policy is optimized by using evaluation network and action network to solve the Nash equilibrium.

Benefits of technology

It simplifies the simulation process of drone swarm adversarial games, improves the efficiency and accuracy of policy learning, and enables the drones to reach high-reward states more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116136945B_ABST
    Figure CN116136945B_ABST
Patent Text Reader

Abstract

本发明提供一种基于反事实基线的无人机集群对抗博弈仿真方法,涉及无人机及强化学习技术领域。该方法首先设定对抗博弈的智能体数和对抗博弈回合数等信息;并初始化每个智能体的动作网络和评估网络参数;然后初始化无人机集群对抗博弈环境,获取环境的初始状态空间以及每个智能体的观察值;计算评估网络输出的损失函数,把评估网络输出误差最小化;最后通过当前动作策略计算每个智能体每个步长的基线;使用无人机集群中所有智能体对应的评估网络计算当前智能体在当前环境下的优势函数,比较当前智能体动作的价值与除去当前智能体动作并保持其他智能体动作不变的反事实基线,更新智能体的动作网络,直至对抗博弈回合数为止。
Need to check novelty before this filing date? Find Prior Art