The present disclosure provides a method for evaluating the capability of a large
language model, aligning a large
language model, related devices and
computer program products, relating to the technical field of
artificial intelligence such as large
language model alignment, model capability evaluation and
deep learning. A specific embodiment of the method for evaluating the capability of a large language model comprises:
processing a sample question by using a large language model to be evaluated to obtain at least two answers to be evaluated; determining a correct answer set from the at least two answers to be evaluated by using a sample answer corresponding to the sample question; in response to the correct answer set including at least two correct answers, generating a first capability evaluation value based on a similarity comparison result between the correct answers, and generating a second capability evaluation value based on a quantity relationship between the correct answers in the correct answer set and the answers to be evaluated; and generating a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value. Thus, the model capability of the large language model can be evaluated more comprehensively, with higher quality and efficiency.