基于蒙特卡洛树搜索和强化学习的近距空战机动决策方法

By constructing a virtual air combat environment and a neural network model, and combining Monte Carlo tree search and reinforcement learning, the problems of high-dimensional state and delayed feedback in autonomous air combat maneuver decision-making of aircraft were solved, and fast, globally optimal maneuver decision-making was achieved.

CN120277980BActive Publication Date: 2026-07-17SICHUAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2023-08-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for autonomous air combat maneuver decision-making methods for aircraft in high-dimensional state spaces and with long-term delayed feedback suffer from the curse of dimensionality, slow learning, and instability, and it is difficult to find a balance between exploring unknown states and utilizing known states.

Method used

A virtual air combat environment is constructed using Monte Carlo tree search and reinforcement learning methods. By combining a policy deep neural network and a value neural network, and collecting experience sample pool data through Selfplay, the policy deep neural network and the value neural network are trained offline to optimize maneuver decisions.

Benefits of technology

In high-dimensional state spaces, it can quickly find optimal action strategies, improve decision-making speed and global optimality, adapt to environmental uncertainty, and improve the accuracy and efficiency of autonomous maneuver decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277980B_ABST
    Figure CN120277980B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于蒙特卡洛树搜索和强化学习的近距空战机动决策方法,涉及飞行器机动决策技术领域,包括构建空战虚拟环境、策略深度神经网络、价值神经网络和蒙特卡洛树搜索;基于蒙特卡洛树搜索、空战虚拟环境和飞行器历史数据,通过Selfplay计算得到经验样本池数据,通过经验样本池数据对策略深度神经网络和价值神经网络进行离线训练,在训练过程中,蒙特卡洛树搜索输出的概率需根据噪声进行修正;根据飞行器实时数据、空战虚拟环境、蒙特卡洛树搜索和训练好的策略深度神经网络和价值神经网络,计算并选择得到飞行器的近距空战机动决策。本发明提升了决策速度及决策全局最优性,可以更好地适应环境的不确定性。
Need to check novelty before this filing date? Find Prior Art