The invention relates to a
robot control method and
system based on a cyclic depth deterministic strategy gradient (RDDPG)
algorithm, and aims to improve the decision and control capability of a
robot in a complex dynamic environment. The method comprises the following steps: firstly, modeling a
robot control system into a
partially observable Markov decision process, and defining an environment state, an action space, a
state transition probability, a reward function and an
observation function; secondly, constructing a cyclic
encoder by adopting a cyclic neural network, taking motion
time sequence data and environment
observation time sequence data of the robot as input, and outputting meta-parameters for identifying environment differences; thirdly, designing an evaluation
value network, an
evaluation strategy network, a target
value network and a target strategy network, generating a control action and evaluating the value of the control action; finally,
empirical data of interaction between the robot and the environment is stored through playback memory, network parameters are updated through
time sequence difference learning, and target network parameters are updated through a
moving average method. According to the invention, the cyclic
encoder is introduced, the
time sequence characteristics of environmental information are fully utilized, the adaptability, learning ability and decision accuracy of the robot in a complex dynamic environment are enhanced, and the method can be widely applied to the fields of industrial
automation, intelligent logistics, service robots and the like.