一种基于动作掩码与奖励塑形MAPPO的人机协同动态调度系统及方法
By using the MAPPO algorithm based on action masking and reward shaping, the problems of worker fatigue management and dynamic resource matching in workshop scheduling are solved, achieving efficient and adaptive dynamic scheduling and improving production efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HOHAI UNIV
- Filing Date
- 2026-05-13
- Publication Date
- 2026-07-17
AI Technical Summary
Existing workshop scheduling methods face challenges in dynamic environments, including the inability to effectively handle dynamic changes in orders, real-time fluctuations in resource status, and worker fatigue management, leading to low production efficiency, frequent safety accidents, and increased defect rates.
A multi-agent near-field policy optimization algorithm (MAPPO) based on action masking and reward shaping is adopted. Through state feature extraction, action mask generation, policy output and environmental interaction and reward feedback modules, dynamic matching and scheduling of worker and robot resources are realized. The fatigue state of workers is optimized by combining a fatigue evolution model. An event-driven time-progression mechanism and a hybrid reward mechanism are adopted to avoid scheduling deadlock.
It achieves efficient and adaptive dynamic scheduling while ensuring worker health, reduces scheduling costs, improves production efficiency, avoids scheduling deadlock, and has better multi-objective solution capabilities and industrial generalization value.
Smart Images

Figure CN122172756B_ABST