Bayesian sparse low-rank fitting large language model confidence calibration method and device

CN122735784APending Publication Date: 2026-09-11JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610914077.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

此类方法虽然能够引入参数不确定性,但仍然可能带来额外可训练参数、后验参数维护成本和推理采样开销,并且未充分利用低秩适配本身由多个秩一分量组合而成的结构特征

Benefits of technology

[0020] First, this disclosure shifts the uncertainty modeling of large language models from a dense model parameter space or a complete low-rank fitting parameter space to a rank-dimensional structure space, requiring only a small number of posterior parameters to be maintained for each rank dimension, thus reducing the additional parameter overhead of Bayesian uncertainty modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735784A_ABST
    Figure CN122735784A_ABST
Patent Text Reader

Abstract

The present disclosure provides a Bayesian sparse low-rank adaptation confidence calibration method and device for large language models. The method obtains a pre-trained large language model and training samples, freezes the original parameters and constructs a low-rank adaptation module; sets low-rank adaptation parameter matrices and and introduces a random diagonal rank mask matrix, so that the low-rank parameter is updated to ; the diagonal elements of are taken as latent variables, the prior distribution and the learnable variational posterior distribution are set, and and the variational parameter are jointly optimized; when reasoning, multiple random diagonal rank mask matrices are sampled, multiple prediction probability distributions are calculated and aggregated to obtain the calibrated prediction probability distribution, and the prediction result, confidence and uncertainty estimation result are output. The present disclosure transfers the uncertainty modeling from the parameter space to the rank dimension structure space, while maintaining the low-rank fine-tuning lightweight advantage, realizes Bayesian regularization and ensemble reasoning, and effectively alleviates the over-confidence problem after fine-tuning of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of natural language processing and large language model fine-tuning, and in particular to a method and apparatus for calibrating the confidence of large language models using Bayesian sparse low-rank adaptation. Background Technology

[0002] In recent years, large language models have demonstrated strong capabilities in scenarios such as text understanding, commonsense reasoning, question answering, code generation, and solving complex tasks. To adapt pre-trained large language models to specific downstream tasks, task adaptation is typically achieved through instruction-based fine-tuning, supervised fine-tuning, or parameter-efficient fine-tuning methods. However, during deterministic fine-tuning, models often develop memory for the training data distribution and overfit, causing them to output high confidence even when making incorrect predictions, resulting in a mismatch between confidence and true accuracy. This overconfidence phenomenon reduces the interpretability and reliability of prediction probabilities and limits the deployment of large language models in high-reliability scenarios such as healthcare, finance, education, autonomous driving assistance, intelligent customer service, and scientific research support.

[0003] Confidence calibration aims to ensure that the predicted probabilities output by the model more accurately reflect the true reliability of the prediction results. Existing uncertainty estimation methods include Bayesian neural networks, Monte Carlo random deactivation, deep ensembles, and posterior approximation methods. These methods can capture model parameter uncertainty or prediction uncertainty to some extent, but when directly applied to large language models with billions or even more parameters, they often require maintaining a huge parameter distribution, training multiple models, or performing costly posterior inference, resulting in significant computational and memory overhead.

[0004] Low-rank fitting is a commonly used, highly efficient parameter fine-tuning method. It significantly reduces the number of training parameters and storage overhead by freezing the original weights of the pre-trained model and introducing low-rank parameters only in the target linear layer for updates. For a pre-trained weight matrix... Standard low-rank adaptation typically represents task-related parameter updates as... ,in and For low-rank fitting parameter matrix, This is the preset rank value. However, standard low-rank adaptation usually... As a fixed global hyperparameter, it is difficult to adaptively adjust according to the actual intrinsic dimensional requirements of different tasks, layers, or projection matrices. When the fixed rank is greater than the effective rank actually required by the task, redundant low-rank components may introduce unnecessary model capacity, thereby exacerbating overfitting and overconfidence; when the fixed rank is too small, it may impair task adaptability.

[0005] Existing Bayesian low-rank fitting methods mostly rely on the low-rank fitting parameter matrix. or Starting from the continuous weight space, models are created for its mean, variance, or posterior approximation. While such methods introduce parameter uncertainty, they may still incur additional trainable parameters, posterior parameter maintenance costs, and inference sampling overhead, and fail to fully utilize the structural feature of low-rank adaptation itself, which consists of multiple rank-one components. Therefore, a novel confidence calibration method for large language models is urgently needed, capable of Bayesian modeling of the effective rank and rank component activation states in the low-rank structure while maintaining the lightweight advantage of low-rank adaptation, thereby simultaneously achieving adaptive model capacity control, overconfidence mitigation, and reliable uncertainty estimation. Summary of the Invention

[0006] To overcome the shortcomings of the existing technology, this disclosure provides a confidence calibration method for large language models using Bayesian sparse low-rank fitting. This method does not perform complex posterior modeling of the complete large language model parameters or all low-rank fitting weights. Instead, it transfers uncertainty modeling to the rank-dimensional structure of the low-rank fitting, controlling the activation states of different rank-1 components through learnable random diagonal rank masks. Bayesian regularization is implemented during the training phase, and multi-mask sampling aggregation is achieved during the inference phase, thereby improving the confidence calibration effect of large language models.

[0007] According to a first aspect of this disclosure, a method for calibrating the confidence of a large language model using Bayesian sparse low-rank fitting is provided, comprising the following steps:

[0008] S1. Model preprocessing: Obtain training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and construct a low-rank adaptation module in at least one target linear layer of the pre-trained large language model.

[0009] S2. Construct a sparse low-rank adaptation structure: For the target linear layer, set a low-rank adaptation parameter matrix. and ,matrix and During training, it is optimized as a trainable deterministic matrix, and in the matrix... and Set a random diagonal rank mask matrix between them , To control the participation of low-rank components corresponding to different rank dimensions in parameter updates, the low-rank parameter update amount of the target linear layer is defined as:

[0010] S3. Latent variable modeling and joint optimization: During training, the matrix... The diagonal elements are used as latent variables, and a prior distribution and a learnable variational posterior distribution are set for these latent variables. The optimization matrix is ​​then jointly optimized by combining the task loss function and the regularization term of the variational posterior distribution relative to the prior distribution. and matrix Variational posterior parameters;

[0011] S4. Multi-mask sampling and probability aggregation: During the inference phase, multiple random diagonal-rank mask matrices are sampled from the variational posterior distribution obtained after the training process converges. The prediction probability distribution corresponding to each random diagonal-rank mask matrix is ​​calculated. All prediction probability distributions are aggregated to obtain the calibrated prediction probability distribution.

[0012] S5. Output Results: Based on the calibrated prediction probability distribution, output the target prediction results, confidence level, and uncertainty estimation results. According to a second aspect of this disclosure, a Bayesian sparse low-rank fitting large language model confidence calibration device is provided, comprising:

[0013] The model preprocessing module is used to obtain training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and construct a low-rank adaptation module in at least one target linear layer of the pre-trained large language model.

[0014] A sparse low-rank adaptation structure module is constructed to build the sparse low-rank adaptation structure: for the target linear layer, a low-rank adaptation parameter matrix is ​​set. and ,matrix and During training, it is optimized as a trainable deterministic matrix, and in the matrix... and Set a random diagonal rank mask matrix between them , To control the participation of low-rank components corresponding to different rank dimensions in parameter updates, the low-rank parameter update amount of the target linear layer is defined as:

[0015] ;

[0016] The latent variable modeling and joint optimization module is used to perform matrix optimization during training. The diagonal elements are used as latent variables, and a prior distribution and a learnable variational posterior distribution are set for these latent variables. The optimization matrix is ​​then jointly optimized by combining the task loss function and the regularization term of the variational posterior distribution relative to the prior distribution. and matrix Variational posterior parameters;

[0017] The multi-mask sampling and probability aggregation module is used to sample multiple random diagonal-rank mask matrices from the variational posterior distribution obtained after the training process converges during the inference phase, calculate the prediction probability distribution corresponding to each random diagonal-rank mask matrix, and aggregate all prediction probability distributions to obtain the calibrated prediction probability distribution.

[0018] The results output module is used to output the target prediction results, confidence level and uncertainty estimation results based on the calibrated prediction probability distribution.

[0019] The technical effects of this disclosure are as follows:

[0020] First, this disclosure shifts the uncertainty modeling of large language models from a dense model parameter space or a complete low-rank fitting parameter space to a rank-dimensional structure space, requiring only a small number of posterior parameters to be maintained for each rank dimension, thus reducing the additional parameter overhead of Bayesian uncertainty modeling.

[0021] Second, this disclosure uses a random diagonal rank mask matrix to adaptively control the low-rank adaptation capacity, which can suppress overfitting and overconfidence caused by redundant rank-1 components, while retaining the necessary task adaptation capability.

[0022] Third, this disclosure aggregates multiple prediction probability distributions through multi-mask sampling during the inference stage, forming a lightweight ensemble effect without the need to train multiple complete large language models, thereby improving the reliability of prediction probabilities and the degree of confidence calibration.

[0023] Fourth, this disclosure can be integrated as a parameter fine-tuning module into various large language models, Transformer structures, or other deep neural networks containing linear mapping layers, and is suitable for multiple-choice question answering, commonsense reasoning, text classification, natural language reasoning, intent recognition, and other natural language processing tasks that require confidence calibration. Attached Figure Description

[0025] Figure 1 The flowchart of the steps of the Bayesian sparse low-rank fitting method for calibrating the confidence of large language models is provided in this disclosure.

[0026] Figure 2 This is a schematic diagram of the structure of the random diagonal rank mask matrix in the low-rank adaptation module of the large language model confidence calibration method for Bayesian sparse low-rank adaptation provided in this disclosure.

[0027] Figure 3 This is a schematic diagram of the structure of a Bayesian sparse low-rank fitting confidence calibration device for large language models provided in this disclosure. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the described embodiments are only a part of the embodiments of this disclosure, and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are all within the protection scope of this disclosure.

[0029] Example 1

[0030] like Figure 1 As shown, this disclosure provides a method for calibrating the confidence of a large language model using Bayesian sparse low-rank fitting, comprising the following steps:

[0031] S1. Model Preprocessing: Obtain the training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and select at least one target linear layer within the pre-trained large language model. A low-rank adaptation module is constructed within the selected target linear layer. The pre-trained large language model can be an autoregressive language model based on a Transformer structure, an encoder-decoder model, or other large-scale neural network models capable of text generation or text discrimination. The target linear layer includes at least the query projection layer, key projection layer, value projection layer, and output projection layer in the attention module of the pre-trained large language model, the linear mapping layer and output mapping layer in the feedforward network, or any combination of the above linear layers.

[0032] S2. Construct a sparse low-rank adaptation structure: For each selected target linear layer, set a low-rank adaptation parameter matrix. and and in the matrix Set a random diagonal rank mask matrix between them This is the frozen original model weight matrix. , Input features, , Module output characteristics Standard low-rank adaptation represents task-related parameter updates as... ,in , , To preset the maximum low-rank adaptation rank, ;

[0033] In this embodiment, the update amount of the low-rank parameter of the target linear layer is defined as:

[0034] ;

[0035] in, , , .when At that time, the first The low-rank components corresponding to each rank dimension are activated; when At that time, the first The low-rank components corresponding to the rank dimension are masked. Therefore, the low-rank parameter update can be further written as:

[0036] ;

[0037] in, For matrix The column vectors, For matrix The Row vectors. Matrix The low-rank fitting parameter matrix and the random diagonal-rank mask matrix are optimized as trainable deterministic matrices during training. This enables independent control over each rank dimension component.

[0038] S3. Latent variable modeling and joint optimization: using a random diagonal-rank mask matrix The diagonal elements are used as latent variables, and a fixed prior distribution and a learnable variational posterior distribution are set for the latent variables. In this embodiment, for the ... One latent variable Set the Bernoulli prior distribution:

[0039] ;

[0040] in, The preset prior activation probability is used to represent the activation probability of the first time when there is no observation task data. The prior tendency of each rank component to be activated. To approximate the unanalyzable true posterior distribution, this disclosure employs a factorized Bernoulli variational posterior distribution:

[0041] ;

[0042] in, For learnable variational parameters, This is the Sigmoid activation function. ,matrix and It can be learned jointly through stochastic gradient descent, adaptive moment estimation, or other gradient optimization algorithms.

[0043] Regarding the training objective, this disclosure minimizes an objective function comprising a task negative log-likelihood term and a variational posterior KL divergence term relative to the prior. In this embodiment, a joint objective function is constructed, consisting of a task negative log-likelihood term and a Bernoulli distribution KL divergence term. The KL divergence term is used to constrain the deviation of the variational posterior distribution from the prior distribution. The joint objective function expression is:

[0044] ;

[0045] in, This represents the number of samples during the training phase. The first digit obtained from the variational posterior distribution A rank mask sample, These are the weighting coefficients of the KL divergence regularization term. The input text / question to be predicted, For the candidate tag set, For candidate tag set One of the candidate tags in the data. For learnable variational parameters, This is the frozen original model weight matrix. , and For low-rank fitting parameter matrix, It follows the Bernoulli prior distribution. For Nully variational posterior distribution. It can be set to a fixed value, or an annealing strategy can be used, or it can be adaptively adjusted according to the training process. This represents the conventional value range in this field. By minimizing the above joint objective function, the low-rank adaptation parameter matrix is ​​simultaneously and jointly optimized. and random diagonal-rank mask matrix The variational posterior parameters.

[0046] When both the prior distribution and the variational posterior distribution are Bernoulli distributions, the first... The KL divergence of rank dimension can be written as:

[0047] ;

[0048] in, To preset the prior activation probability, Indicates the first rank mask variable The corresponding learnable variational parameters, This represents the Sigmoid function, used to convert learnable variational parameters. Mapped to Within the interval. To achieve the desired results for discrete latent variables. Differentiability optimization is achieved by employing Gumbel-Sigmoid reparameterization or continuous relaxation sampling during the training phase in this embodiment. Specifically, let... Sample from a uniform distribution Uniform(0,1). Then the rank mask variable after continuous relaxation can be expressed as:

[0049] ;

[0050] in, For temperature parameters, Let represent a random variable obtained by independent sampling from the standard uniform distribution Uniform(0,1). This represents the Logistic noise term obtained by transforming the uniform random variable. The temperature parameter can be a fixed value or it can vary during training. Through this continuous relaxation, the model can adjust the variational parameters during backpropagation. Perform gradient updates.

[0051] S4. Multi-mask sampling and probability aggregation: After the model completes training, it enters the inference phase, where multiple random diagonal-rank mask matrices are sampled from the variational posterior distribution of the trained model. For the input sample Calculate the prediction probability distribution corresponding to each rank mask matrix, and average the multiple prediction probability distributions to obtain the calibrated prediction probability distribution:

[0052] ;

[0053] in, The number of samples during the inference phase. It is a positive integer. This is the frozen original model weight matrix. , and For low-rank fitting parameter matrix, The input text / question to be predicted, For the candidate tag set, For candidate tag set One of the candidate tags in the data. Given input At that time, candidate tags The calibrated predicted probability, The randomized diagonal-rank mask matrix can be configured based on computational budget, deployment latency, and calibration requirements. This step does not require training multiple full large language models; instead, it achieves similar ensemble prediction-calibration results by sampling multiple rank masks to form multiple lightweight low-rank substructures.

[0054] S5. Output Results: Based on the calibrated prediction probability distribution, output the target prediction result and the corresponding confidence level or uncertainty estimate. In classification or multiple-choice question-answering scenarios, the candidate label with the highest prediction probability can be selected as the final prediction result, and the maximum prediction probability can be used as the confidence level.

[0055] ;

[0056] in, The input text / question to be predicted, For the candidate tag set, For candidate tag set One of the candidate tags in the data. For the final predicted label, To predict confidence levels, Given input At that time, candidate tags In other embodiments, the calibrated predicted probability can also output uncertainty estimation results based on indicators such as prediction entropy, variance between multiple sampled predictions, mutual information, negative log-likelihood, or calibration error.

[0057] Example 2

[0058] like Figure 3 As shown in the embodiments of this disclosure, a confidence calibration device for a large language model using Bayesian sparse low-rank adaptation is also provided, comprising:

[0059] The model preprocessing module is used to obtain training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and construct a low-rank adaptation module in at least one target linear layer of the pre-trained large language model.

[0060] A sparse low-rank adaptation structure module is constructed to build the sparse low-rank adaptation structure: for the target linear layer, a low-rank adaptation parameter matrix is ​​set. and ,matrix and During training, it is optimized as a trainable deterministic matrix, and in the matrix... and Set a random diagonal rank mask matrix between them , To control the participation of low-rank components corresponding to different rank dimensions in parameter updates, the low-rank parameter update amount of the target linear layer is defined as:

[0061] ;

[0062] The latent variable modeling and joint optimization module is used to perform matrix optimization during training. The diagonal elements are used as latent variables, and a prior distribution and a learnable variational posterior distribution are set for these latent variables. The optimization matrix is ​​then jointly optimized by combining the task loss function and the regularization term of the variational posterior distribution relative to the prior distribution. and matrix Variational posterior parameters;

[0063] The multi-mask sampling and probability aggregation module is used to sample multiple random diagonal-rank mask matrices from the variational posterior distribution obtained after the training process converges during the inference phase, calculate the prediction probability distribution corresponding to each random diagonal-rank mask matrix, and aggregate all prediction probability distributions to obtain the calibrated prediction probability distribution.

[0064] The results output module is used to output the target prediction results, confidence level and uncertainty estimation results based on the calibrated prediction probability distribution.

[0065] The method disclosed herein can be applied to low-rank adaptation fine-tuning of commonsense reasoning and multiple-choice question-answering tasks in large language models. Evaluation metrics include accuracy. Expected calibration error and negative log-likelihood ,in A higher value indicates a higher prediction accuracy (↑). and The lower the value, the better the prediction probability calibration and reliability (↓).

[0066] In practical applications, a pre-trained large language model can be used as the base model, and comparisons can be made with representative methods such as standard low-rank fitting, Bayesian low-rank fitting, post-trained Bayesian low-rank fitting, and contextualized low-rank fitting on multiple commonsense reasoning datasets. Experimental results show that the proposed method can reduce performance while maintaining competitive accuracy. and in The near-optimal results achieved demonstrate its ability to mitigate the overconfidence problem caused by deterministic fine-tuning. In some embodiments of this disclosure, experiments were conducted on six datasets: ARC-C, OBQA, BoolQ, AAO, WB, and Combined. The results obtained using the method of this disclosure were compared with those obtained using existing training methods for uncertainty estimation. Tables 1 and 2 show... The number of mask samplings during the inference phase is shown in Table 1. The selected value is 10, in Table 2 The selected value is 5. The results in Table 1 are obtained by fine-tuning with Llama-3.1-8B as the backbone, and the results in Table 2 are obtained by fine-tuning with Llama2-7B as the backbone. The following baseline methods were selected for comparison in this disclosure: (1) Maximum Likelihood Estimation (MLE); (2) Bayesian Low-Rank Adaptive Algorithm BloB based on backpropagation; (3) Contextual Low-Rank Adaptive C-LoRA for Uncertainty Estimation of Large Language Models; (4) Training-free Bayesian TFB for Low-Rank Adapters of Large Language Models; (5) Training-free Bayesian TFB-MLE based on Maximum Likelihood Estimation Adapter; (6) Training-free Bayesian TFB-MAP based on Maximum A posteriori Estimation Adapter; (7) Training-free Bayesian TFB-BloB based on Backpropagation Bayesian Low-Rank Adapter.

[0067] Table 1 presents some exemplary experimental results. The listed values ​​are only used to illustrate the technical effects of the methods disclosed herein and do not constitute a limitation on the scope of protection of this disclosure.

[0068] Table 1

[0069] ACC↑ MLE 81.08 87.90 89.58 ACC↑ BLoB 80.86 87.66 88.69 ACC↑ TFB 80.18 88.20 88.84 ACC↑ C-LoRA 81.70 86.93 87.77 ACC↑ This disclosure 81.74 88.24 89.43 ECE↓ MLE 16.35 9.77 8.69 ECE↓ BLoB 5.87 3.35 2.46 ECE↓ TFB 6.19 4.51 3.80 ECE↓ C-LoRA 10.75 6.50 4.36 ECE↓ This disclosure 5.60 3.12 1.82 NLL↓ MLE 1.20 0.61 0.52 NLL↓ BLoB 0.59 0.38 0.27 NLL↓ TFB 0.62 0.36 0.29 NLL↓ C-LoRA 0.68 0.40 0.30 NLL↓ This disclosure 0.56 0.36 0.27

[0070] As shown in Table 1, in the embodiments, the method of this disclosure achieves low calibration errors on multiple tasks and maintains competitive performance in terms of accuracy. This result demonstrates that by learning the rank mask posterior and performing multi-mask sampling aggregation during the inference phase, the reliability of prediction confidence for large language models can be improved without significantly increasing the number of fine-tuning parameters.

[0071] To verify the applicability of the disclosed method under multi-task hybrid distribution, multiple commonsense reasoning datasets were fused to construct a fused dataset, and the accuracy of different low-rank adaptation uncertainty estimation methods was compared on the fused dataset. Expected calibration error and negative log-likelihood The fused dataset includes three settings: AAO, WB, and Combined. AAO is obtained by combining ARC-Easy, ARC-Challenge, and OpenBookQA; WB is obtained by combining Winogrande-Medium and BoolQ; and Combined is obtained by merging all six commonsense reasoning datasets. Experimental results are shown in Table 2. In the table, Indicates accuracy; the higher the value, the better (↑). This indicates the desired calibration error; the smaller the value, the better (↓). This represents the negative log-likelihood, and the smaller the value, the better (↓).

[0072] Table 2

[0073] ACC↑ BLoB 81.58 81.31 79.16 ACC↑ C-LoRA 79.56 78.36 81.00 ACC↑ TFB-MLE 82.10 79.89 79.77 ACC↑ TFB-MAP 81.80 80.97 80.81 ACC↑ TFB-BLoB 80.56 78.11 76.87 ACC↑ This disclosure 81.48 80.26 80.16 ECE↓ BLoB 2.58 1.42 1.14 ECE↓ C-LoRA 2.78 3.30 2.39 ECE↓ TFB-MLE 4.58 2.29 2.79 ECE↓ TFB-MAP 3.20 4.84 2.14 ECE↓ TFB-BLoB 3.98 3.01 1.47 ECE↓ This disclosure 1.81 1.35 1.82 NLL↓ BLoB 0.52 0.39 0.45 NLL↓ C-LoRA 0.57 0.44 0.44 NLL↓ TFB-MLE 0.55 0.45 0.48 NLL↓ TFB-MAP 0.53 0.43 0.46 NLL↓ TFB-BLoB 0.60 0.48 0.53 NLL↓ This disclosure 0.52 0.39 0.45

[0074] As shown in Table 2, under the fused dataset setting, the proposed method achieves low expected calibration error on both AAO and WB while maintaining competitive accuracy, and approaches the optimal result on the negative log-likelihood index. This result demonstrates that the proposed method is applicable not only to single commonsense reasoning tasks but also to complex data distributions formed by mixing multiple tasks. By learning the rank mask posterior and performing multi-mask sampling aggregation during the inference stage, the proposed method can improve the reliability and confidence calibration effect of prediction probabilities in multi-task fusion scenarios. The experimental results listed in Table 2 are only used to illustrate the technical effects of the proposed method and do not constitute a limitation on the scope of protection of this disclosure.

[0075] Because this method primarily introduces a small number of variational posterior parameters in each rank dimension, rather than building a high-dimensional continuous posterior for the complete large language model parameters or all low-rank adaptation weights, the number of additional parameters is relatively small. Compared to methods that require maintaining multiple model replicas or performing full Gaussian posterior modeling on the low-rank parameter matrix, this method is more suitable for deployment under resource-constrained conditions.

[0076] Optional implementation methods and equivalent substitutions.

[0077] Without departing from the core ideas of this disclosure, the prior distribution is not limited to the Bernoulli distribution with fixed activation probabilities, but may also employ hierarchical priors, layer-related priors, module-related priors, or prior distributions determined by task information; the variational posterior distribution is not limited to the fully factorized Bernoulli distribution, but may also employ discrete distributions with related structures, mixed distributions, or other learnable posterior approximations.

[0078] The continuous relaxation sampling is not limited to the Gumbel-Sigmoid method, but can also use the Concrete distribution, the pass-through estimator, or other sampling methods that can support approximate backpropagation of discrete variables.

[0079] The low-rank adaptation module is not limited to the ordinary LoRA form, but can also be combined with other parameter-efficient fine-tuning structures, such as adaptive LoRA, hierarchical LoRA, low-rank plus bias adaptation, prefix tuning, hint tuning, or other fine-tuning modules with decomposable structures.

[0080] The inference aggregation method is not limited to simple arithmetic average, but can also use weighted average, temperature-scaled average, aggregation based on uncertainty weights, or other probability fusion methods.

[0081] The method is applicable not only to large language models, but also to multimodal large models containing low-rank adaptation modules, visual language models, speech language models, or other Transformer-type deep learning models.

[0082] The embodiments described above are merely preferred embodiments of this disclosure and are not intended to limit the scope of protection of this disclosure. Various modifications, substitutions, and improvements made by those skilled in the art to the technical solutions of this disclosure without departing from its spirit should fall within the scope of protection defined by the claims of this disclosure.

Claims

1. A confidence calibration method for large language models using Bayesian sparse low-rank fitting, characterized in that, Includes the following steps: S1. Model preprocessing: Obtain training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and construct a low-rank adaptation module in at least one target linear layer of the pre-trained large language model. S2. Construct a sparse low-rank adaptation structure: For the target linear layer, set a low-rank adaptation parameter matrix. and ,matrix and During training, it is optimized as a trainable deterministic matrix, and in the matrix... and Set a random diagonal rank mask matrix between them , To control the participation of low-rank components corresponding to different rank dimensions in parameter updates, the low-rank parameter update amount of the target linear layer is defined as: ; S3. Latent variable modeling and joint optimization: During training, the matrix... The diagonal elements are used as latent variables, and a prior distribution and a learnable variational posterior distribution are set for these latent variables. The optimization matrix is ​​then jointly optimized by combining the task loss function and the regularization term of the variational posterior distribution relative to the prior distribution. and matrix Variational posterior parameters; S4. Multi-mask sampling and probability aggregation: During the inference phase, multiple random diagonal-rank mask matrices are sampled from the variational posterior distribution obtained after the training process converges. The prediction probability distribution corresponding to each random diagonal-rank mask matrix is ​​calculated. All prediction probability distributions are aggregated to obtain the calibrated prediction probability distribution. S5. Output Results: Based on the calibrated prediction probability distribution, output the target prediction results, confidence level, and uncertainty estimation results.

2. The method for calibrating the confidence level of a large language model using Bayesian sparse low-rank fitting according to claim 1, characterized in that, The target linear layer includes at least one of the following: an attention projection layer, a feedforward linear layer, and an output mapping layer; The random diagonal-rank mask matrix satisfy: ; Where, vector , To preset the maximum low-rank adaptation rank, Indicates the first The activation state of the low-rank component corresponding to each rank dimension; The low-rank parameter update amount is further expressed as: ; in, For matrix The column vectors, For matrix The row vectors; when At that time, the first Each low-rank component participates in the model parameter update, when At that time, the first One low-rank component is masked.

3. The method for calibrating the confidence level of a large language model using Bayesian sparse low-rank fitting according to claim 1 or 2, characterized in that, The prior distribution of the latent variables in S3 is a Bernoulli prior distribution, and the variational posterior distribution is a factorized Bernoulli distribution; Among them, for the first Latent variables corresponding to each rank dimension Its prior distribution is: ; in, Preset prior activation probability; Its variational posterior distribution is: ; in, For learnable variational parameters, For the Sigmoid function; The joint optimization includes: sampling latent variables from the variational posterior distribution, using the sampled random diagonal-rank mask matrix to complete the model forward propagation, and minimizing the objective function composed of the task negative log-likelihood term and the Bernoulli distribution KL divergence term; wherein, the Bernoulli distribution KL divergence term is used to constrain the degree of deviation of the variational posterior distribution from the Bernoulli prior distribution. During the training phase, the latent variables are sampled using the Gumbel-Sigmoid reparameterization method or the continuous relaxation sampling method to achieve differentiability optimization of the discrete rank mask variables.

4. A confidence calibration device for large language models using Bayesian sparse low-rank fitting, characterized in that, include: The model preprocessing module is used to obtain training samples of the pre-trained large language model and the task to be calibrated, freeze the original model parameters in the pre-trained large language model, and construct a low-rank adaptation module in at least one target linear layer of the pre-trained large language model. Construct a sparse low-rank adaptor structure module to set the low-rank adaptor parameter matrix for the target linear layer. and ,matrix and During training, it is optimized as a trainable deterministic matrix, and in the matrix... and Set a random diagonal rank mask matrix between them , To control the participation of low-rank components corresponding to different rank dimensions in parameter updates, the low-rank parameter update amount of the target linear layer is defined as: ; The latent variable modeling and joint optimization module is used to perform matrix optimization during training. The diagonal elements are used as latent variables, and a prior distribution and a learnable variational posterior distribution are set for these latent variables. The optimization matrix is ​​then jointly optimized by combining the task loss function and the regularization term of the variational posterior distribution relative to the prior distribution. and matrix Variational posterior parameters; The multi-mask sampling and probability aggregation module is used to sample multiple random diagonal-rank mask matrices from the variational posterior distribution obtained after the training process converges during the inference phase, calculate the prediction probability distribution corresponding to each random diagonal-rank mask matrix, and aggregate all prediction probability distributions to obtain the calibrated prediction probability distribution. The results output module is used to output the target prediction results, confidence level and uncertainty estimation results based on the calibrated prediction probability distribution.