一种基于物理蒸馏强化学习的船舶预设时间跟踪控制方法

By employing physical distillation reinforcement learning, combined with teacher and student network architectures and an execution-evaluation network, the problem of precise tracking and energy consumption optimization in unknown dynamic environments during ship control was solved, achieving efficient and economical navigation within a preset time.

CN122411029APending Publication Date: 2026-07-17DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN MARITIME UNIVERSITY
Filing Date
2026-05-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing ship control technologies lack the ability to effectively integrate prior physical knowledge into deep reinforcement learning, making it difficult to achieve accurate tracking and optimal control under strict time constraints. Furthermore, traditional methods face risks of energy consumption spikes and model divergence in unknown dynamic environments.

Method used

A physical distillation-based reinforcement learning approach is adopted. The ship's energy dissipation rules are extracted through a teacher-student network architecture. An adaptive virtual control law is constructed by combining backstepping and a saturated preset time function. An execution-evaluation network is introduced to perform optimal feedback control, thereby achieving path tracking within a preset time.

Benefits of technology

In an unknown dynamic environment, the ship path tracking error was strictly converged, reducing energy consumption and optimizing control costs, thereby improving the system's environmental adaptability and navigation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122411029A_ABST
    Figure CN122411029A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于物理蒸馏强化学习的船舶预设时间跟踪控制方法包括建立船舶三自由度欠驱动船舶非线性数学模型;构建基于教师网络‑学生网络的物理蒸馏神经网络;构建预设时间自适应虚拟控制律;根据预设时间自适应虚拟控制律与物理蒸馏神经网络,构建欠驱动船舶的自适应预设时间前馈控制律与权值更新律;根据自适应预设时间前馈控制律与权值更新律,构建欠驱动船舶的预设时间最优反馈控制律,并通过结合自适应预设时间前馈控制律,实现船舶预设时间路径跟踪控制。解决了现有方法在面对动态特性未知时,传统深度强化学习因脱离船舶物理特性及缺乏动态环境适应能力,导致不能在严格预设时间约束下兼顾船舶的精准路径跟踪与控制代价最优的问题。
Need to check novelty before this filing date? Find Prior Art