一种基于强化学习的无人机自主目标跟踪控制方法、系统及设备

By constructing an autonomous tracking reinforcement learning task model and designing the TD3 reinforcement learning training method, combined with teacher-guided annealing training mechanism and multidimensional reward function, the problem of high-precision tracking of UAVs in complex environments was solved, achieving smooth and stable tracking of UAVs and avoiding the risk of loss of control.

CN122086092BActive Publication Date: 2026-07-17CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT
Filing Date
2026-04-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing reinforcement learning-based UAV control methods struggle to achieve high-precision autonomous tracking in complex environments and are prone to issues such as unsmooth trajectories and control divergence. Furthermore, existing reward function designs fail to adequately consider the relative geometric relationship between the UAV and the target.

Method used

An autonomous tracking reinforcement learning task model was constructed, the TD3 reinforcement learning training method was designed, a teacher-guided annealing training mechanism was introduced, and flight control constraints and multi-dimensional reward functions were designed. Through joint optimization of the policy network and the dual-value network, a target UAV tracking and control model was established.

Benefits of technology

It achieves high-precision autonomous tracking of drones in complex environments, avoiding uncontrolled behaviors such as overspeeding and rapid acceleration, and ensuring safety and smooth operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086092B_ABST
    Figure CN122086092B_ABST
Patent Text Reader

Abstract

本申请公开了一种基于强化学习的无人机自主目标跟踪控制方法、系统及设备,涉及计算机技术领域,包括:S1、构建自主跟踪强化学习任务模型,为无人机跟踪控制模型的训练和优化提供约束基础;S2、设计TD3强化学习训练方法,通过策略网络与双价值网络的联合优化进行自学习与动态更新;S3、引入教师引导退火训练机制,并设计飞行控制约束、强化学习稳定机制和多维奖励函数,以建立目标无人机跟踪控制模型;飞行控制约束包括控制指令约束和加速度约束;S4、利用目标无人机跟踪控制模型进行针对运动目标的无人机目标跟踪控制操作。在复杂环境下实现无人机的高精度平滑稳定自主跟踪操作。
Need to check novelty before this filing date? Find Prior Art