The invention discloses a navigation method, device and equipment based on a GRPO
algorithm and a medium, and relates to the technical field of
reinforcement learning, and the method comprises the steps: calculating the average similarity between a current strategy and a plurality of previous iteration strategies based on KL
divergence; updating a step length factor through an average reward change rate and an average similarity determined based on a plurality of iterated rewards; determining a
gradient estimation correction item based on the
gradient estimation of the sampling trajectory, determining target
gradient estimation according to the gradient
estimation correction item and the original gradient
estimation, and updating the current strategy through the target gradient
estimation and the updated step length factor; when the current strategy is updated, the
importance weight is
cut, the target function of the GRPO
algorithm is corrected according to the
cut weight, the GRPO
algorithm is trained based on the corrected function and the updated strategy, so that the
intelligent agent learns the optimal strategy based on the trained GRPO algorithm, and the outlet of the labyrinth is determined according to the optimal strategy. Therefore, the stability of the algorithm is improved.