The invention provides a thinking chain
selection method based on internal confidence of a large
language model, and belongs to the field of
artificial intelligence. Relates to a big
language model reasoning optimization technology, in particular to a training-free thinking chain
selection method based on big
language model inherent confidence, and aims to improve reasoning efficiency and evaluate reasoning process rationality and answer accuracy at the same time. Based on a probability framework and taking internal confidence as a scoring
signal, setting selection of an optimal answer question as maximum likelihood
estimation for solving a joint probability distribution space of a thinking chain and a final answer; setting X to represent a possible prompt set, R to represent a thinking chain set and Y to represent a possible final answer set, and for a given input prompt x belongs to X, an optimal thinking chain r belongs to R and an answer y belongs to Y corresponding to the optimal thinking chain; identifying a pair (r, y) with the highest joint
conditional probability; wherein the probability is a joint
conditional probability, and refers to the probability that the model generates a thinking chain r and a final answer y under a given prompt x.