According to the model training method and device, the task execution method and device, the
electronic equipment and the storage medium, in the method, the prompt content, the reply content output by the target model for the prompt content and the process data for generating the reply content can be firstly obtained, and then the prompt content, the reply content and the process data are input into the
reward system; therefore, the
reward value of each reasoning step in the process data is obtained, and finally, the target model is iteratively trained based on the
reward value of each reasoning step. The
reward system of the method no longer generates the
reward value for the token of the sample, but generates the reward value for each reasoning step in the process data, so that the target model can pay attention to the integrity and logicality of the reply content in the training process, and then the performance and stability of the target model in a complex task and the robustness of the model are improved.