The invention discloses a cloth folding optimization method and equipment based on deep
reinforcement learning. The method comprises the following steps: acquiring a cloth image, a key point and a folding
actuator state in a
mass point-spring physical
simulation environment, and outputting a folding action by utilizing a model comprising a visual
encoder, a strategy network and a double-action
value network; a task progress is constructed based on a key point distance, sub-rewards such as
angular point alignment, flatness, wrinkles and symmetry are dynamically weighted, and a stable folding strategy model is trained in combination with SumTree-based priority experience playback and
advantage weighting strategy updating. The equipment comprises a folding
actuator, an
image acquisition unit, a memory and a processor, and the processor executes the strategy model and generates a control instruction according to an acquired image and an
actuator state to drive the folding actuator to complete cloth folding. According to the scheme, the sample
utilization rate and stability are improved, and a cloth folding strategy with higher folding quality and better generalization performance is obtained.