The invention provides a model training method and device for a text open task, on one hand, a reward model is trained through a training text pair marked with a relative quality
advantage relation, original complex absolute quality scores of the reward model are converted into simpler and more stable relative
advantage judgment, the model training cost and subjective influence are reduced, and the training efficiency is improved. The accuracy and confidence of the reward
signal are improved, one first reply in the output replies of the existing model is set as a reference text reply, and it is ensured that the output quality of the trained text model is superior to the output quality of the existing model; on the other hand, in the training process of the to-be-trained model, a second reply output by the to-be-trained model is monitored, and the target reward model is adjusted according to the monitoring result; and / or, the
reward value output by the target reward model for the second reply is monitored, and the reference text reply is adjusted according to the monitoring result, so that the optimization direction of the text model is ensured, and the
training effect of the text model is improved.