Intelligent decision-making method and device based on offline-online hybrid reinforcement learning
CN117648548BActive Publication Date: 2026-08-28NAT UNIV OF DEFENSE TECH
Patent Information
- Application Number
- CN202311639444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-12-01
AI Technical Summary
Technical Problem
[0003]智能博弈对抗问题被建模为马尔科夫决策问题时,仿真环境构建难度高,状态空间和动作空间大,任务奖励设计复杂,仿真推演的成本高且高质量的离线数据稀缺,当使用纯离线或纯在线强化学习方法进行策略学习时,面临比求解一般的控制决策问题更为复杂的学习训练过程,在无人机仿真场景下难以实现智能体的有效决策
Benefits of technology
[0027]上述基于离线-在线混合强化学习的智能决策方法和装置,通过获取空中博弈仿真对抗中对抗双方交互产生的离线数据集,将离线数据集存储在经验回放池,采用预先设置的混合采样方法从经验回放池中采集离线训练样本,能够高样本利用率和模型训练效率,提高策略的可靠性,根据离线强化学习算法和离线训练样本训练预先构建的策略网络和Q值网络,将训练好的策略网络和Q值网络作为在线训练的初始网络,接着,将在线数据集存储在经验回放池,采用混合采样方法从当前经验回放池中采集在线训练样本,根据在线强化学习算法和在线训练样本对初始网络的网络参数以及在线数据集进行迭代更新,直到满足预设的迭代终止条件时,停止迭代,得到训练好的智能决策模型,最后,根据智能决策模型辅助第一方无人机进行空中博弈仿真对抗场景下的智能决策。本发明实施例,在小样本数据和专家经验的条件下,能够提高无人机在面对复杂空中博弈环境下的自主智能决策能力和适应复杂环境的能力,从而灵活应对各种博弈挑战。
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN117648548B_ABST
Abstract
The application relates to an intelligent decision-making method and device based on offline-online hybrid reinforcement learning. The method comprises the following steps: acquiring an offline data set generated by interaction between two parties in air game simulation confrontation; collecting offline training samples from the offline data set by using a hybrid sampling method, training a pre-constructed strategy network and a Q value network according to an offline reinforcement learning algorithm and the offline training samples, and taking the trained strategy network and Q value network as initial networks for online training; storing an online data set in an experience playback pool where the offline data set is located, collecting online training samples from the current experience playback pool by using a hybrid sampling method, training the initial networks according to an online reinforcement learning algorithm and the online training samples, and obtaining a trained intelligent decision-making model; and assisting a first unmanned aerial vehicle in intelligent decision-making in an air game simulation confrontation scene according to the intelligent decision-making model. The method can improve the reliability and accuracy of intelligent decision-making of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Batch constraint offline reinforcement learning algorithm for mixed behavior space
CN115640830A
Integrated modeling method and system for combat entity behavior model
CN115906673A