一种微推力航天器协同追逃的多智能体强化学习控制方法及装置

By using Beta distribution and adaptive action mapping techniques, combined with a high-precision orbital dynamics model, and optimizing spacecraft strategies, the problems of low strategy exploration efficiency and large gradient estimation errors in multi-agent orbital pursuit and escape were solved, achieving efficient collaborative spacecraft pursuit.

CN121634861BActive Publication Date: 2026-07-17PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV
Filing Date
2026-01-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multi-agent orbital pursuit game algorithms are inefficient in strategy exploration in highly dynamic and constrained environments, have large gradient estimation errors, and perform poorly at the boundaries, making it difficult to achieve efficient collaborative pursuit in complex orbital environments.

Method used

A stochastic strategy for the spacecraft is constructed using a Beta distribution strategy. The Beta distribution sampled values ​​are converted into three-dimensional thrust vectors through adaptive action mapping technology. Combined with a high-precision J2 perturbation orbital dynamics model, a hierarchical reward mechanism is designed to optimize the strategy parameters and achieve a balance from uniform exploration to single-peak utilization.

Benefits of technology

It significantly improves the spacecraft's adaptability and trajectory flexibility in complex orbital environments, increases mission success rate, first capture time, and capture duration, reduces gradient estimation errors, and achieves a high-quality pursuit strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121634861B_ABST
    Figure CN121634861B_ABST
Patent Text Reader

Abstract

本发明属于航空航天领域,为了解决航天器在轨道追逃随机博弈中微推力动作空间探索效率低、策略梯度估计误差大的问题,本发明提供了一种微推力航天器协同追逃的多智能体强化学习控制方法及装置,该方法包括:采用Beta分布策略构建航天器的随机策略,通过调节Beta分布可学习参数实现从均匀探索到单峰利用的自适应策略优化;通过自适应动作映射技术将Beta分布采样值精确转换为三维推力向量,建立与航天器实际推力空间的完整对应关系。本发明能够以优秀的探索‑利用平衡性能实现最高效的协同追捕。本发明展现出动态适应特性,可以通过自动调整Beta分布的形状参数,改变策略的探索和利用强度,增强航天器集群对不同轨道构型和对抗策略的适应能力。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Pulse type track pursuit barrier cooperative game intelligent decision control method

    CN116991067A

  • Cluster spacecraft multi-target intelligent cooperative tracking method based on reinforcement learning

    CN121325612A