一种基于动作掩码与奖励塑形MAPPO的人机协同动态调度系统及方法

By using the MAPPO algorithm based on action masking and reward shaping, the problems of worker fatigue management and dynamic resource matching in workshop scheduling are solved, achieving efficient and adaptive dynamic scheduling and improving production efficiency and safety.

CN122172756BActive Publication Date: 2026-07-17HOHAI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HOHAI UNIV
Filing Date
2026-05-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing workshop scheduling methods face challenges in dynamic environments, including the inability to effectively handle dynamic changes in orders, real-time fluctuations in resource status, and worker fatigue management, leading to low production efficiency, frequent safety accidents, and increased defect rates.

Method used

A multi-agent near-field policy optimization algorithm (MAPPO) based on action masking and reward shaping is adopted. Through state feature extraction, action mask generation, policy output and environmental interaction and reward feedback modules, dynamic matching and scheduling of worker and robot resources are realized. The fatigue state of workers is optimized by combining a fatigue evolution model. An event-driven time-progression mechanism and a hybrid reward mechanism are adopted to avoid scheduling deadlock.

Benefits of technology

It achieves efficient and adaptive dynamic scheduling while ensuring worker health, reduces scheduling costs, improves production efficiency, avoids scheduling deadlock, and has better multi-objective solution capabilities and industrial generalization value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122172756B_ABST
    Figure CN122172756B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于动作掩码与奖励塑形MAPPO的人机协同动态调度系统及方法,包括:状态特征提取模块在触发调度决策事件时,提取车间内各工位的局部状态特征,动作掩码生成模块生成动态动作掩码;策略输出模块输出合规的调度动作;环境交互与奖励反馈模块根据合规的调度动作锁定对应的作业、工位、工人或机器人资源,利用疲劳演化模型更新工人疲劳状态,并根据作业基准加工时间、工人疲劳状态和所选加工模式计算作业实际加工时间;基于基础运行惩罚、系统势能差和连续时间变步长计算混合奖励,并进行优势函数估计以及更新策略和价值网络参数;本发明有效降低调度死锁风险,为人机协同车间调度提供了一种高效、安全、科学的解决方案。
Need to check novelty before this filing date? Find Prior Art