The application provides a policy transfer method and device for
reinforcement learning and a storage medium, and belongs to the technical field of
machine learning. The method comprises the following steps: obtaining first sample data at multiple
simulation time points based on a
simulation scene of an autonomous driving scene; determining first policy information and a first prediction model based on multiple groups of the first sample data; in each
iteration process, obtaining second sample data at a target time point in the autonomous driving scene based on the first policy information of the current
iteration process, updating the first prediction model of the current
iteration process based on the second sample data, updating the first policy information based on the updated first prediction model, and performing the next iteration process based on the updated first prediction model and the first policy information. Through the method, the prediction model and the policy information in the
simulation scene can be transferred to the autonomous driving scene, so that the action to be performed by the autonomous vehicle can be accurately predicted in the autonomous driving scene, and a better effect can be achieved after the action is performed.