一种微推力航天器协同追逃的多智能体强化学习控制方法及装置
By using Beta distribution and adaptive action mapping techniques, combined with a high-precision orbital dynamics model, and optimizing spacecraft strategies, the problems of low strategy exploration efficiency and large gradient estimation errors in multi-agent orbital pursuit and escape were solved, achieving efficient collaborative spacecraft pursuit.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multi-agent orbital pursuit game algorithms are inefficient in strategy exploration in highly dynamic and constrained environments, have large gradient estimation errors, and perform poorly at the boundaries, making it difficult to achieve efficient collaborative pursuit in complex orbital environments.
A stochastic strategy for the spacecraft is constructed using a Beta distribution strategy. The Beta distribution sampled values are converted into three-dimensional thrust vectors through adaptive action mapping technology. Combined with a high-precision J2 perturbation orbital dynamics model, a hierarchical reward mechanism is designed to optimize the strategy parameters and achieve a balance from uniform exploration to single-peak utilization.
It significantly improves the spacecraft's adaptability and trajectory flexibility in complex orbital environments, increases mission success rate, first capture time, and capture duration, reduces gradient estimation errors, and achieves a high-quality pursuit strategy.
Smart Images

Figure CN121634861B_ABST
Abstract
Citation Information
Patent Citations
Pulse type track pursuit barrier cooperative game intelligent decision control method
CN116991067A
Cluster spacecraft multi-target intelligent cooperative tracking method based on reinforcement learning
CN121325612A