The application discloses a
NAO robot object grasping training method based on direct preference optimization, and comprises the following steps: (1) collecting and saving human expert data,
imitation data and trajectory data of interactive grasping actions in a
robot simulator using a
reinforcement learning strategy; (2) scoring state-action trajectory sequence data pairs in the trajectory data by a human annotator to obtain a trajectory
data set based on human preferences; (3) designing a reference strategy network using the human
preference data set, designing a parameterized strategy network, constructing a maximum likelihood target in combination with the reference strategy network, and performing
gradient descent on the target function to obtain an optimal strategy; (4) training an object recognition model, deploying the object recognition model and the optimal strategy to a real
NAO robot, and evaluating the grasping actions of the
robot; and (5) iteratively training and evaluating the process until the
robot can successfully grasp the target object.