This invention relates to the technical field of
robot training. Specifically, it relates to a
robot adaptive training method, apparatus, and medium based on
reinforcement learning. The method includes generating task sub-
objective information using a high-level policy network, inputting the task sub-
objective information into a low-level execution network, generating
action control commands based on the task sub-
objective information, collecting interaction feedback information during execution, calculating a
reward value based on the interaction feedback information, associating the
reward value with scene complexity parameters, performing dynamic reward shaping operations to generate an adjusted reward
signal, generating a policy model optimized by meta-
learning based on the adjusted reward
signal, loading the policy model into a
simulation environment to generate a target policy model optimized by
simulation training, loading the target policy model into the
robot, and controlling the robot to perform task operations in an actual interactive
scenario. This invention achieves the effect of adaptive
task learning for robots in multi-interaction scenarios.