The invention discloses a fixed base mechanical arm point-to-
point trajectory optimization method based on a near-end strategy optimization
algorithm, solves the problem that a high-dimensional continuous action space exists when deep
reinforcement learning is directly used for a fixed base mechanical arm
tail end position to reach a task, and belongs to the field of
robot operation control and mechanical arm autonomous
motion optimization. The method comprises the steps that a mechanical arm
kinematics model is established, and constraint conditions are defined; modeling a
tail end point-to-point arrival task as a
sequential decision process, constructing a
state vector as the input of a strategy network, and taking a
tail end displacement increment
direction vector as the action output of the strategy network; mapping the tail end displacement increment
direction vector into a joint increment through differential
inverse kinematics, and performing constraint
processing according to a constraint condition to generate an
executable joint control instruction; and training the strategy network by adopting a near-end strategy optimization
algorithm, obtaining a single-step reward through a multi-target reward function, and evaluating the performance output by the action of the strategy network according to the single-step reward until convergence.