This invention discloses a non-periodic dynamic detection and maintenance method for multi-state systems based on
Gaussian demand and customized PPO (Progressive Point of Action), comprising the following steps: A)
System and
demand modeling: including a multi-state
system model, time-varying
demand modeling, and definition of detection and
maintenance actions; B) Continuous-time MDP modeling: constructing the dynamic detection and maintenance
decision problem as a continuous-time Markov
decision process, including
state space, action space, state transition probabilities, reward function, and objective function; C) Customized PPO
algorithm framework: designing a deep
reinforcement learning framework based on PPO, adapting to the
hybrid action space, and efficiently solving the MDP model. The
demand modeling of this invention is accurate: by using a
Gaussian process to model time-varying demand, it can simultaneously capture the expected trend, random fluctuations, and
time correlation of demand. Compared with traditional constant, linear, or simplified Markov demand modeling, it is more in line with actual industrial scenarios and effectively reduces the risk of supply-demand mismatch.