Method and apparatus for calibrating a language model
The temperature scaling method addresses the overconfidence issue in aligned language models by calibrating their predictive distributions, improving their reliability and consistency, and enabling safer applications.
Patent Information
- Application Number
- PCT/CN2023/120048
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-09-20
- Publication Date
- 2025-06-12
AI Technical Summary
Aligned language models tend to be overconfident in their predictions, making it difficult to discern between truthful and hallucinated answers, which hinders their application in safety-critical domains.
A simple, effective, and sample-efficient temperature scaling method is proposed to calibrate aligned language models by refining their predictive distributions based on a temperature scaling coefficient, aligning them with the predictive distributions of pre-trained models.
The proposed method effectively calibrates aligned language models, reducing overconfidence and improving the consistency between predictive confidence and accuracy, thereby enhancing their reliability in practical applications.
Smart Images

Figure CN2023120048_12062025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR CALIBRATING A LANGUAGE MODELFIELD
[0001] Aspects of the present disclosure relate generally to artificial intelligence (AI) , and more particularly, to method and apparatus for calibrating a language model (LM) .BACKGROUND
[0002] Uncertainty calibration, as an important metric for building reliable deep learning systems, measures the consistency of the posterior probability (or predictive confidence) that the model gives about the output with the true correctness likelihood. For example, when a well-calibrated model gives some predictions with 80%confidence, then those predictions should achieve exactly 80%accuracy, i.e., the model knows what it knows.
[0003] Having calibrated predictive confidence is helpful for the application of LMs.
[0004] Such property allows human users to better detect and correct undesired behaviors such as hallucinations in the LM by receiving reliable signals given by the model about the correctness of the answers and the answers themselves, establishing trust in the LM-based applications.
[0005] Aligning pre-trained LMs with human feedback has achieved great success in real-world application scenarios. However, the aligned LMs are known to be more overconfident in their answers compared to the pre-trained LMs, adding the difficulty of discerning between truthful and hallucinated answers of the models, thus hindering the application of aligned LMs in safety-critical domains.SUMMARY
[0006] In order to address the above-mentioned problem, the disclosure propose a simple effective, and sample-efficient temperature scaling method which effectively calibrate aligned LMs in practical applications using the predictive distribution of the pre-trained LMs.
[0007] According to an embodiment, there provides a computer implemented method for calibrating a LM, comprising: receiving, by a first LM, a set of language questions and generating a first set of predictive answer confidence distributions corresponding respectively to the set of language questions; receiving, by a second LM, the set of language questions and generating a second set of predictive answer confidence distributions corresponding respectively to the set of language questions, wherein the generating the second set of predictive answer confidence distributions comprising calibrating the second set of predictive answer confidence distributions based on a temperature scaling coefficient; determining a distance between the first set of predictive answer confidence distributions and the second set of predictive answer confidence distributions; and updating the temperature scaling coefficient based on the distance.
[0008] According to an embodiment, there provides a computer implemented method for performing a language task, comprising: receiving a language question by a LM which is obtained by using the calibrating method according to aspects of the disclosure; outputting an answer by the LM in response to the language question.
[0009] According to an embodiment, there provides a computer system, which comprises one or more processors and one or more storage devices storing computer- executable instructions that, when executed, cause the one or more processors to perform the operations of the method as mentioned above as well as to perform the operations of the method according to aspects of the disclosure.
[0010] According to an embodiment, there provides one or more computer readable storage media storing computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method as mentioned above as well as to perform the operations of the method according to aspects of the disclosure.
[0011] According to an embodiment, there provides a computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method as mentioned above as well as to perform the operations of the method according to aspects of the disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
[0013] Fig. 1 is a schematic block diagram illustrating a pre-trained LM and an aligned LM according to aspects of the disclosure.
[0014] Fig. 2 illustrates an exemplary language question according to aspects of the disclosure.
[0015] Fig. 3 illustrates an exemplary reliability diagram according to aspects of the disclosure.
[0016] Fig. 4 illustrates exemplary experimental results of accuracy and ECE according to aspects of the disclosure.
[0017] Fig. 5 illustrates exemplary changes of experimental results of accuracy, ECE, confidence under different settings according to aspects of the disclosure.
[0018] Fig. 6 illustrates an exemplary process for calibrating an aligned LM according to aspects of the disclosure.
[0019] Fig. 7 illustrates exemplary calibration results of different methods according to aspects of the disclosure.
[0020] Fig. 8 illustrates an exemplary process for calibrating a LM according to aspects of the disclosure.
[0021] Fig. 9 illustrates an exemplary process for performing a language task according to aspects of the disclosure,
[0022] Fig. 10 illustrates an exemplary computing system according to aspects of the disclosure.DETAILED DESCRIPTION
[0023] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
[0024] Various embodiments will be described in detail with reference to the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. References made to particular examples and embodiments are for illustrative purposes, and are not intended to limit the scope of the disclosure.
[0025] Fig. 1 is a schematic block diagram illustrating a pre-trained LM and an aligned LM according to aspects of the disclosure.
[0026] A pre-trained LM 110 may be denoted as Given a sequence of text tokens x with length L, the pre-trained LM models the conditional probability p (xl|x<l) of the token xl at position l over a vocabulary space V, given all previous tokens x<l. In practical applications or scenarios, the LM 110 is expected to generate high quality response y based on human instructions x on some downstream tasks Dtask, such as question answering or code completion. For the pre-trained LM 110, such task may be accomplished by in-context learning (ICL) , where the LM needs to learn how to perform the task by following few-shot demonstrations under the same task. In specific, denoting the instruction-response pair of a downstream task as (x, y) ~ Dtask, ICL gives the response y of an instruction x by where SK is a concatenation of K independent instruction-response pairs sampled from the downstream task, where the K independent instruction-response pairs are used as the demonstrations. For example, as illustrated in Fig. 2, the in-context examples 220 allow the LM 110 to accomplish the ICL.
[0027] Although ICL enables efficient adaptation of pre-trained LMs to downstream tasks, recent works have shown that further aligning LMs with high-quality human preference data can significantly boost the performance and generalization of LMs in downstream applications. As illustrated in Fig. 1, an alignment process 130 may be performed to the pre-trained LM 110 to obtain aligned LM 120. The alignment process 130 may include supervised fine-tuning (SFT) and / or learning from pairwise feedback (LPF) . The alignment process 130 may be implemented with existing methods. In an embodiment, the alignment process 130 includes the SFT stage and the LPF stage. For the SFT stage, the pre-trained LM is fine-tuned on curated responses y with length Ly of the corresponding instructions x through the language modeling objective, i.e., maximizing The LM obtained through SFT based on the pre-trained LM may be denoted as For the LPF stage, pairwise preference data (x, yw, yl) are collected based on an implicit human preference reward function r: where r (x, yw) >r (x, yl) . Then the SFT model is further optimized to align with the pairwise data, which can be achieved by Reinforcement Learning from Human Feedback (RLHF) policy or preference-based cross-entropy loss. The resulted LPF model may be denoted as which is an example of the aligned LM 120. As an example, Llama is a known pre-trained LM and Vicuna is the aligned LM for the pre-trained LM Llama. As another example, Llama-2 is a known pre-trained LM and Llama-2-Chat is the aligned LM for the pre-trained LM Llama-2. It is appreciated that the pre-trained LM and corresponding aligned LM are not limited to the exampled ones, and the solution of the disclosure is applicable to any suitable pre-trained LM and corresponding aligned LM.
[0028] Fig. 2 illustrates an exemplary language question according to aspects of the disclosure.
[0029] A considerable portion of downstream tasks or applications can be adapted to the multiple-choice setting, where logit-based uncertainty quantification (UQ) is tractable. Consider a task whose sample consists of an instruction (or question) x and a set of candidate responses yc∈Y. To assess the language answer pθ (yc∣x) through language modeling, the sample (x, Y) may be formatted into a multiple-choice question. As illustrated in Fig. 2, a question body 200 for the multiple-choice question may be created using a mapping The question body 200 may include the task description 210, the in-context examples 220, the question example 230. The question example 230 may include the question description 2310 and all candidate responses yc 2320. A choice letter 2330 is assigned for each candidate yc, and a trigger token 2340 and a target generation position 2350 may also be optionally included in the question body 200. The task description 210 and the in-context examples 220 is optional and may not be included in the question body 200. If the in-context examples 220 are included in the question body 200, the ICL is employed where the in-context examples 220 provides demonstrations for the LM. If the in-context examples 220 are not included in the question body 200, the zero-shot leaning (ZSL) is employed where no in-context examples 220 provides demonstrations for the LM. Then the answer pθ (yc∣x) with the probability of the choice letter may be estimated over all given tokens by the LM 110 or 120.
[0030] Denoting the LM’s prediction as and the ground truth as uncertainty calibration examines the consistency between the LM’s correctness and confidence in population level. A prefect calibrated model holds and thus have zero expected calibration error (ECE) , i.e. To evaluate calibration with ECE in practice with N finite samples, the model’s confidence may be grouped into M bins. In an embodiment, 10 equal-sized bins may be employed to estimate ECE, denote as ECE10. Denoting Bm as the indices of samples in the m-th bin, then the ECE can be estimated by the weighted average of the difference between confidence and accuracy in each bin:
[0031] Fig. 3 illustrates an exemplary reliability diagram according to aspects of the disclosure.
[0032] The dash line in Fig. 3 illustrates a perfect calibration (PC) of a LM, that is, the LM’s confidence (CON) and accuracy (ACC) are consistent with zero ECE. The darker line illustrates the reliability of the pre-trained LM which is Llama in this example, the brighter line illustrates the reliability of the aligned LM which is Vicuna in this example. As shown in Fig. 3, although the overall accuracy (58.63) of the aligned LM is higher than that (57.76) of the pre-trained LM, the aligned LM are overconfident as can be observed from the comparison between the line of PC and the line of the aligned LM, that is to say, aligned LMs are mis-calibrated due to overconfidence on multiple choice questions with logit-based evaluation. On the other hand, the large pre-trained LM is well-calibrated as can be observed from the comparison between the line of PC and the line of the pre-trained LM. The mis-calibration of the aligned LM over confidence may hinder the application of the aligned LMs in many scenarios as the aligned LM is too confident about its prediction.
[0033] To find a way to alleviate the mis-calibration of the aligned LM, comprehensive experiments on several pairs of pre-trained and aligned LMs are performed, focusing on the differences in calibration and other related metrics between the both under multiple-choice setting. An exemplary experiments setting is shown in the following.
[0034] Model: Llama family ranging from model size 7B to 70B is used as the pre-trained LMs, Vicuna and Llama-2-Chat are used as the aligned LM for the pre-trained LMs Llama and Llama-2, respectively.
[0035] Prompt Format: all data are adapted to the multiple-choice format. As shown in Fig. 2, an input prompt consists of an optional task description 210, optional in-context examples 220, and a test example 230. Two types of format for the choices: “A” and “ (A) ” may be used, where the LM need to directly output the choice letter after “Answer: ” for the former case, while the LM will use the trigger token “ (” as a sort of hint and output the choice letter after it for the latter case.
[0036] Evaluation Protocol: both zero-shot learning (ZSL) and few-shot in-context learning (ICL) evaluation are performed for each dataset. For zero-shot setting, the in-context examples part 220 of the prompt are simply omitted. For few-shot setting, the in-context examples part 220 of the prompt are maintained.
[0037] Metrics: the evaluation is based on the LM’s output logits of a specific target generation position. As illustrated in Fig. 2, for example, the choice letter with the highest probability among all choices may be taken as the prediction of the LM and its probability over the whole token space may be used as LM’s predictive confidence. For each task, the accuracy and ECE10 are used as main metrics. The LM’s average predictive confidence, sum of all choice letter probabilities, and probability of the trigger token (if available) may be tracked to better understand the LM’s behavior.
[0038] Fig. 4 illustrates exemplary experimental out-of-the-box results of accuracy and ECE averaged across all seven tasks and two prompt formats for all tested LMs according to aspects of the disclosure.
[0039] The out-of-the-box results refer to the results obtained by using the LM’s confidence without post-processing. The model size (MS) includes 7B, 13B, 33B or 70, and the LMs include two pre-trained LMs (Llama-1, Llama-2) and two corresponding aligned LMs (Vicuna-v1.3, Llama-2-Chat) . The pre-trained LMs overall have low ECE in the ICL setting, from which the largest pre-trained model, Llama-2 70B, achieves the best performance in both accuracy and ECE among all pre-trained LMs and aligned LMs. All aligned LMs have higher ECE than their corresponding pre-trained models, regardless of size. Unlike the pre-trained LMs, the aligned LMs have less change in accuracy and ECE between ZSL and ICL settings.
[0040] To better understand how calibration varies between pre-trained and aligned LMs and between ZSL and ICL settings, it is monitored how LMs’ predictive confidence behaved with these variables. Fig. 5 illustrates changes of experimental results of accuracy (ACC) , ECE, averaged confidence (AVG CON) under different settings according to aspects of the disclosure. The top half of Fig. 5 corresponds to the choice format “A” and the bottom half of Fig. 5 corresponds to the choice format “ (A) ” . As shown in the top half of Fig. 5, for pre-trained LMs under the choice format “A” , the sum of choice letters’ probabilities (SoCL) is low in the ZSL setting, indicating that pre-trained LMs are uncertain about the format of their responses, i.e., whether they should directly output the choice letters at first for multiple choice questions, which may be referred to as the format uncertainty. Such uncertainty leads to low predictive confidence, resulting in high ECE for pre-trained LMs in ZSL setting. Looking at the ICL settings, the pre-trained LMs eliminate the format uncertainty by following in-context examples. Notably, the predictive confidence adjusted by ICL is well-calibrated out-of-the-box, demonstrating that pre-trained LMs are calibrated in-context learners. On the other hand, under the choice format “ (A) ” , the format uncertainty is decomposed into the given trigger token “ (” , so the difference between the calibration of the model on ZSL and ICL is smaller in this setup. Nonetheless, the predictive confidence of the models is still slightly affected by ICL, perhaps due to the fact that the in-context examples can provide information beyond the response format for LMs.
[0041] In contrast to the pre-trained LMs, the aligned LMs are overconfident out-of-the-box in both ZSL and ICL settings and prefer to directly output the choice letter to answer multiple-choice questions, i.e., the format uncertainty is relatively low compared to pre-trained LMs. Using ICL for the aligned LMs can also eliminate the format uncertainty, but its effect on accuracy and calibration is minimal, unlike pre-trained LMs. These results suggest that the alignment process destroys the well-calibrated predictive distributions of the pre-trained LMs and cannot be repaired by ICL at inference time.
[0042] To tackle the mis-calibration caused by the alignment process, an effective approach is to conduct post-hoc calibration on the predictive confidence of aligned LMs. Unlike traditional deep learning models, a language model will face diverse tasks at the same time in real-world deployments, so it needs a calibration approach that is computationally friendly, sample-efficient, and accuracy-maintaining.
[0043] Fig. 6 illustrates an exemplary process for calibrating predictive confidence of the aligned LM according to aspects of the disclosure.
[0044] A validation set of language data is input into the pre-trained LM 110 and the aligned LM 120, where (xi, yi) denotes a pair of language question and answer. For each language question xi, each of the pre-trained LM 110 and the aligned LM 120 generates predictive answer confidence distribution that is, the aligned LM 120 generates predictive answer confidence distribution and the pre-trained LM 110 generates predictive answer confidence distribution Taking the question illustrated in Fig. 2 as an example, there are four candidate choices with four choice letters, the predictive answer confidence distribution denotes the respective possibilities of the four candidate choices, and thus may be a vector having four elements indicating possibilities in this example.
[0045] The aligned LM120 includes a post-hoc calibration module 1210. The post-hoc calibration module 1210 may perform temperature scaling to refine the predictive distribution of the aligned LM120. In an embodiment, the post-hoc calibration module 1210 may scales the raw logits li at the positions of choice letters, particularly the raw logits li are scaled to li / T by the temperature scaling coefficient T. Taking the question illustrated in Fig. 2 as an example, the raw logits li are four raw numbers corresponding to the four choice letters. Then the scaled logits li / T are processed by the probability generation module 1220 to generate the predictive answer confidence distribution which includes respective possibilities of the four candidate choices. In an embodiment, the probability generation module 1220 may be implemented with a softmax function, that is, = softmax (li / T) .
[0046] A distance determination module 130 may determine a distance between the predictive answer confidence distribution of the aligned LM 120 and the predictive answer confidence distribution of the pre-trained LM 110. In an embodiment, the distance may be a Kullback-Leibler (KL) divergence: In an embodiment, the accumulated distance over the validation set may be obtained at the distance determination module 130:
[0047] A parameter updating module 140 may update the temperature scaling coefficient T in a direction of minimizing the KL divergence: The update of the coefficient T may be performed by using a gradient descent method, which is a known method for updating parameters.
[0048] The updated temperature scaling coefficient T output from the module 140 is used to update the temperature scaling coefficient T of the post-hoc calibration module 1210. Then in the next iteration, the process described above is repeated with the new temperature scaling coefficient T. And the process is iteratively performed to minimize the KL divergence. In this way, the temperature scaling approach is used to enforce the consistency of the predictive distribution between the aligned LM and pre-trained LM, where the degree to which the predictive distribution of the aligned LM changes relative to the pre-trained LM is learned by simply using one parameter T.
[0049] Fig. 7 illustrates calibration results of the proposed calibration method compared to other base line methods according to aspects of the disclosure.
[0050] All the calibration methods are tested on the aligned LM Llama-2-Chat 70B whose corresponding pre-trained LM Llama-2 70B has the best calibration performance among all tested LMs. The base line methods include out-of-the-box (OOTB) method, the few-shot temperature scaling (TS) method, Kernel density estimation (KDE) method, the TS with constant T=2.5, and enhanced TS (ETS) method according to the embodiment of Fig. 6. The five methods are respectively denoted with different shadings as shown in Fig. 7. The horizontal axis denotes different tasks or applications, where seven tasks are tested and illustrated. As shown in Fig. 7, both TS and KDE cannot calibrate the LM well for all tasks with few-shot examples. In some tasks (e.g., LogiQA, IMDB) their roles can be complementary, while in others (e.g., OpenbookQA) both are bad than out-of-the-box calibration. Based on the overconfident a priori of the aligned LMs, using one temperature uniformly (T == 2.5) for all tasks is a strong baseline. However, the optimal temperature under different tasks may be very different, resulting in a sub-optimal solution. Among these methods, the proposed method ETS is the only one that outperforms out-of-the-box calibration on all tasks and calibrates the language model most effectively in the vast majority of scenarios. This suggests that learning the degree to which the predictive distribution of the aligned LM changes relative to the pre-trained LM by using one parameter for each task is an efficient and powerful post-hoc calibration solution.
[0051] Fig. 8 illustrates an exemplary process for calibrating a language model (LM) according to aspects of the disclosure.
[0052] At step 810, a first LM receives a set of language questions and generates a first set of predictive answer confidence distributions corresponding respectively to the set of language questions. For example, the set of language questions may be the questions in the validation set.
[0053] At step 820, a second LM receives the set of language questions and generates a second set of predictive answer confidence distributions corresponding respectively to the set of language questions, wherein the generating the second set of predictive answer confidence distributions comprises calibrating the second set of predictive answer confidence distributions based on a temperature scaling coefficient.
[0054] At step 830, a distance between the first set of predictive answer confidence distributions and the second set of predictive answer confidence distributions is determined. According an embodiment, the distance is Kullback-Leibler (KL) divergence.
[0055] At step 840, the temperature scaling coefficient is updated based on the distance.
[0056] According an embodiment, the steps 810 to 840 are iteratively performed to minimize the distance.
[0057] According an embodiment, the first LM is a pretrained LM, and the second LM is an aligned LM obtained by fine-tuning the pretrained LM.
[0058] According an embodiment, the first LM and the second LM are configured with ICL.
[0059] According an embodiment, the fine-tuning the pretrained LM comprises fine-tuning the pretrained LM through a SFT method and / or a LPF method.
[0060] According an embodiment, the calibrating the second set of predictive answer confidence distributions comprises: scaling raw logits corresponding to each of the second set of predictive answer confidence distributions based on the temperature scaling coefficient; and generating the predictive answer confidence distribution based on the scaled raw logits.
[0061] According an embodiment, the generating the predictive answer confidence distribution based on the scaled raw logits comprises: generating the predictive answer confidence distribution by using a softmax function to the scaled raw logits.
[0062] According an embodiment, the set of language questions are multiple-choice questions, and each of the first and second sets of predictive answer confidence distributions is a distribution of probabilities of candidate choices of a corresponding language question.
[0063] According an embodiment, wherein the set of language questions are multiple-choice questions, the raw logits corresponding to each of the second set of predictive answer confidence distributions are raw logits at positions of choice letters of candidate choices of a corresponding language question.
[0064] Fig. 9 illustrates an exemplary process for performing a language task according to aspects of the disclosure.
[0065] At step 910, a LM receives a language question, where the LM is obtained by using the calibrating method of Fig. 8 as well as by using the calibrating method according to aspects of the disclosure. For example, the LM is the aligned LM which has been calibrated using the calibrating method of Fig. 8 as well as by using the calibrating method according to aspects of the disclosure.
[0066] At step 920, the LM outputs an answer in response to the language question.
[0067] Fig. 10 illustrates an exemplary computing system according to aspects of the disclosure. The computing system 1000 may comprise at least one processor 1010. The computing system 1000 may further comprise at least one storage device 1020. The storage device 1020 may store computer-executable instructions that, when executed, cause the processor 1010 to perform any operations according to the embodiments of the present disclosure as described in connection with Figs. 1-9.
[0068] The embodiments of the present disclosure may be embodied in a computer-readable medium such as non-transitory computer-readable medium. The non-transitory computer-readable medium may comprise instructions that, when executed, cause one or more processors to perform any operations according to the embodiments of the present disclosure as described in connection with Figs. 1-9.
[0069] The embodiments of the present disclosure may be embodied in a computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform any operations according to the embodiments of the present disclosure as described in connection with Figs. 1-9.
[0070] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts.
[0071] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
[0072] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims.
Claims
1.A computer implemented method for calibrating a language model (LM) , comprising:receiving, by a first LM, a set of language questions and generating a first set of predictive answer confidence distributions corresponding respectively to the set of language questions;receiving, by a second LM, the set of language questions and generating a second set of predictive answer confidence distributions corresponding respectively to the set of language questions, wherein the generating the second set of predictive answer confidence distributions comprising calibrating the second set of predictive answer confidence distributions based on a temperature scaling coefficient;determining a distance between the first set of predictive answer confidence distributions and the second set of predictive answer confidence distributions; andupdating the temperature scaling coefficient based on the distance.2.The method of claim 1, wherein the receiving and generating by the first LM, the receiving and generating by the second LM, the determining the distance and the updating the temperature scaling coefficient are iteratively performed to minimize the distance.3.The method of claim 1, wherein the first LM is a pretrained LM, and the second LM is an aligned LM obtained by fine-tuning the pretrained LM.4.The method of claim 3, wherein the first LM and the second LM are configured with in-context leaning (ICL) .5.The method of claim 3, wherein the fine-tuning the pretrained LM comprises fine-tuning the pretrained LM through a supervised fine-tuning (SFT) method and / or a learning from pairwise feedback (LPF) method.6.The method of claim 1, wherein the calibrating the second set of predictive answer confidence distributions comprises:scaling raw logits corresponding to each of the second set of predictive answer confidence distributions based on the temperature scaling coefficient; andgenerating the predictive answer confidence distribution based on the scaled raw logits.7.The method of claim 6, wherein the generating the predictive answer confidence distribution based on the scaled raw logits comprises:generating the predictive answer confidence distribution by using a softmax function to the scaled raw logits.8.The method of claim 1, wherein the distance is Kullback-Leibler (KL) divergence.9.The method of claim 1, wherein the set of language questions are multiple-choice questions, and each of the first and second sets of predictive answer confidence distributions is a distribution of probabilities of candidate choices of a corresponding language question.10.The method of claim 6, wherein the set of language questions are multiple-choice questions, the raw logits corresponding to each of the second set of predictive answer confidence distributions are raw logits at positions of choice letters of candidate choices of a corresponding language question.11.A computer implemented method for performing a language task, comprising:receiving a language question by a LM which is obtained by using the calibrating method of one of claims 1-10;outputting an answer by the LM in response to the language question.12.A computer system, comprising:one or more processors; andone or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of one of claims 1-11.13.One or more computer readable storage media storing computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method of one of claims 1-11.