Fine adjustment method and system for rejection perception intelligent question and answer model based on gradient driving

By constructing a rejection-aware dataset and utilizing gradient-driven and meta-learning optimization methods, the problems of illusion and over-rejection in open-domain question-answering scenarios for large language models are solved, achieving high efficiency and reliability in model responses.

CN120929569APending Publication Date: 2025-11-11WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511059084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing large-scale language models suffer from illusion problems and overrejection in open-domain question-answering scenarios, making it difficult to balance reducing illusions with avoiding overrejection.

Method used

A rejection-aware dataset is constructed, and sample selection is performed using gradient-driven methods. The confidence threshold and temperature are optimized through second-order gradient guidance and meta-learning, and adaptive weight fine-tuning is carried out to form a distillation dataset and fine-tune the model.

Benefits of technology

It significantly reduces the hallucination rate, suppresses excessive rejection, improves the quality and usability of model responses, and achieves a balanced optimization of model rejection and response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929569A_ABST
    Figure CN120929569A_ABST
Patent Text Reader

Abstract

The invention discloses a gradient driving-based refusal perception intelligent question-answer model fine tuning method, which comprises the following steps of: constructing a refusal perception data set which comprises known correct samples and unknown samples needing refusal; constructing a second-order gradient guided influence rejection formula, and performing sample selection by utilizing gradient driving to form a distillation data set; distributing a weight to each sample in the distillation data set, optimizing a confidence coefficient threshold value and temperature by utilizing meta-learning, and carrying out self-adaptive weight fine tuning; and completing deployment and fine tuning of the large model, and outputting a new model weight. The problem that it is difficult to balance'reduce illusion 'and'avoid excessive rejection' in the prior art is solved, various questions can be answered accurately and efficiently, and powerful support is provided for improvement of generalization ability of intelligent question answering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of large language model (LLM) fine-tuning and security and reliability technology, specifically involving a gradient-driven rejection-aware intelligent question answering model fine-tuning method and system, which is used to reduce the model's "illusion" wrong answers while avoiding "over-rejection", and to achieve a balanced optimization of the model's rejection and answering capabilities. Background Technology

[0002] Large language models (LLMs), such as GPT and LLaMA, have made significant progress and demonstrated outstanding capabilities in a variety of downstream tasks. Despite this success, the "hallucination" problem is becoming increasingly prominent in scenarios such as open-domain question answering, especially in the generation of semi-hallucinations—when faced with unfamiliar or ambiguous queries, the model generates incorrect or fictitious information, affecting the reliability of downstream applications.

[0003] Ideally, a responsible LLM should refuse to answer questions beyond its knowledge scope to minimize hallucinations. Recent research has developed Rejection-Aware Instruction Tuning (RAIT), which constructs a rejection-aware dataset and teaches the model to appropriately reject responses using supervised fine-tuning (SFT). Typically, the rejection-aware dataset divides training samples into ik (correct) and idk (incorrect) groups based on the correctness of the responses. Samples with incorrect responses (idk) are considered to have unknown knowledge, and their answers are replaced with rejection responses such as "I don't know," while correct responses (ik) remain unchanged. Although RAIT has been successful in reducing hallucinations and effectively lowering the hallucination rate, existing RAIT methods have two major shortcomings:

[0004] Over-rejection: After learning to "reject the unknown", the model tends to reject questions that it could answer correctly, resulting in a decrease in the effective response rate;

[0005] Low sample utilization: Relying solely on model output confidence or random sampling to construct rejection samples makes it difficult to accurately select the most valuable data for fine-tuning. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, this invention provides a gradient-driven method and system for fine-tuning a rejection-aware intelligent question-answering model. This method addresses the difficulty in balancing "reducing illusions" and "avoiding excessive rejection" in the prior art, enabling accurate and efficient answers to various questions and providing strong support for improving the generalization ability of intelligent question answering.

[0007] According to one aspect of the present invention, a method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven methods is provided, comprising:

[0008] Construct a rejection-aware dataset, including known correct samples and unknown samples to be rejected;

[0009] We construct a rejection influence formula guided by second-order gradient, and use gradient-driven sample selection to form a distillation dataset.

[0010] Each sample in the distillation dataset is assigned a weight, and meta-learning is used to optimize the confidence threshold and temperature for adaptive weight fine-tuning.

[0011] Complete the deployment and fine-tuning of the large model, and output the new model weights.

[0012] As a further technical solution, a rejection-aware dataset is constructed, including:

[0013] Query the internal state of the LLM, and divide the original training samples into known correct samples and unknown rejection samples based on the query results. The known correct samples are those in which the model answers correctly, and the original answers are retained. The unknown rejection samples are those in which the model answers incorrectly, and their answers are replaced with rejection responses.

[0014] As a further technical solution, the method also includes:

[0015] A batch of samples with ambiguous boundaries are generated using GANs or diffusion models and then merged into a set of unknown samples that need to be rejected.

[0016] As a further technical solution, a second-order gradient-guided rejection influence formula is constructed, and gradient-driven sample selection is performed, including:

[0017] Construct the rejection effect formula of the second gradient, and calculate the contribution of each unknown sample to be rejected to minimizing the inaccuracy;

[0018] Gradient contrast loss is introduced for candidate known correct samples / unknown rejection samples to learn, so that the gradients within known correct samples are close to each other, and the gradients between known correct samples and unknown rejection samples are separated to the maximum extent, thereby enhancing the distinguishability of static conflicts.

[0019] For the unknown samples that need to be rejected after the influence ranking, fine-tuning is performed to form a distillation dataset for target training.

[0020] As a further technical solution, adaptive weight fine-tuning is performed, including:

[0021] Each sample is assigned a weight, and the confidence threshold and temperature parameter are optimized using meta-learning. The THS index on the validation set is updated in reverse to automatically find the optimal value. During training, the loss of unknown samples that need to be rejected is weighted to mitigate over-rejection.

[0022] As a further technical solution, the method also includes:

[0023] Before performing meta-learning hyperparameter tuning, the stability effect is calculated for each sample in the distillation dataset.

[0024] According to one aspect of the present invention, a gradient-driven rejection-aware intelligent question-answering model fine-tuning system is provided, comprising:

[0025] The first main module is used to construct the rejection-aware dataset, which includes known correct samples and unknown samples that need to be rejected;

[0026] The second main module is used to construct the rejection effect formula guided by the second-order gradient, and to use gradient-driven sample selection to form a distillation dataset.

[0027] The third main module is used to assign a weight to each sample in the distillation dataset, and to use meta-learning to optimize the confidence threshold and temperature for adaptive weight fine-tuning.

[0028] The fourth main module is used to complete the deployment and fine-tuning of the large model and output new model weights.

[0029] According to one aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the gradient-driven rejection-aware intelligent question-answering model fine-tuning method.

[0030] According to one aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the gradient-driven rejection-aware intelligent question-answering model fine-tuning method.

[0031] According to one aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the gradient-driven rejection-aware intelligent question-answering model fine-tuning method.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] 1. This invention significantly reduces the hallucination rate while effectively suppressing excessive rejection, thus improving the overall quality and usability of responses. The introduction of second-order gradients and adversarial examples makes the fine-tuned data more representative and informative, reducing redundant training overhead. The meta-learning automatic tuning and multimodal fusion mechanism make this method adaptable to various question-answering tasks and data distributions.

[0034] 2. This invention provides a novel fine-tuning framework that, while efficiently utilizing training samples, achieves a balance between "reducing illusions" and "avoiding excessive rejection" through more refined sample selection and weighting strategies. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart illustrating the gradient-driven rejection-aware intelligent question-answering model fine-tuning method provided in an embodiment of the present invention.

[0037] Figure 2 This is a schematic diagram of a gradient-driven rejection-aware intelligent question-answering model fine-tuning system provided in an embodiment of the present invention.

[0038] Figure 3 A schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0039] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0040] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be decomposed, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0042] This invention combines gradient-driven and LLM techniques to develop a rejection-aware intelligent question-answering model fine-tuning system. While efficiently utilizing training samples, it achieves a balance between "reducing illusions" and "avoiding excessive rejection" through more refined sample selection and weighting strategies.

[0043] Please see Figure 1 The gradient-driven rejection-aware intelligent question-answering model fine-tuning method described in this embodiment of the invention includes the following steps:

[0044] Step 1: Construct a rejection-aware dataset, which involves dividing the original training samples into two classes by querying the internal state of the LLM.

[0045] Step 2: Use gradient-driven sample selection. Through the rejection influence formula guided by the second-order gradient, efficiently select the samples that are most critical to reducing hallucinations.

[0046] Step 3: Use meta-learning to optimize the threshold and temperature for adaptive weight fine-tuning to alleviate over-rejection and ensure that the model does not lose its ability to answer known questions;

[0047] Step 4: Complete the deployment and fine-tuning of the large model, output the new model weights, and realize the end-to-end question-answering system;

[0048] Step 1, which involves constructing a rejection-aware dataset from the acquired raw dataset, includes:

[0049] By querying the internal state of the LLM (such as the confidence level of the generated answer), the original training samples are divided into two categories: D_ik (known correct samples): samples in which the model answers correctly, and the original answer is retained. D_idk (unknown and to be rejected samples): samples in which the model answers incorrectly, and their answers are replaced with rejection responses (such as "I don't know"). In addition to the existing confidence-based idk selection, "blurred" text samples are synthesized using GAN or a diffusion model and labeled as pseudo-idk. These adversarial samples are merged with the original idk to expand the coverage of D_idk, providing structured rejection-aware data for subsequent training, distinguishing the known and unknown knowledge boundaries of the model, and improving the model's robustness to diverse unknown scenarios.

[0050] Specifically, for each sample x in the original dataset D_src, use Generate N answers and calculate the accuracy C(x) to divide the confidence level; if C(x) ≥ TC, then it is assigned to the known set D_ik; otherwise, it is assigned to the unknown set D_idk; then use GAN or diffusion model to generate a batch of boundary-ambiguous samples D_adv and merge them into D_idk.

[0051] Step 2, which utilizes gradient-driven sample selection, includes:

[0052] The second-order gradient rejection influence formula is constructed to calculate the contribution of each idk sample to minimizing inaccuracy. Gradient contrast loss is introduced to the candidate ik / idk samples for learning, so that the gradients within ik are close to each other and the gradients between ik and idk are separated to the maximum, thereby enhancing the distinguishability of static conflicts. Then, the idk samples after influence ranking are fine-tuned to form a distillation dataset for target training.

[0053] Specifically, sample selection is driven by second-order gradients; for each x ∈ D_idk, the first-order influence is calculated using the following formula:

[0054]

[0055] Where x represents the training sample, For gradient operators, For loss function, Represents the original (unperturbed) samples and labels. Representative and original sample The true labels of the pairing, These are the model parameters.

[0056] For the same batch of samples, calculate the Hessian vector product:

[0057]

[0058] Where v is the average gradient of idk.

[0059] Then calculate using the following formula:

[0060]

[0061] in It is a custom weight parameter.

[0062] Take the first N_ik samples in descending order to form D_sel_ik; take the first N_idk samples in descending order to form And perform gradient contrastive pre-training, that is, add gradient contrastive loss to {D_sel_ik∪D_sel_idk}:

[0063]

[0064] in , Let represent the gradient direction vectors of the i-th sample on ik and idk respectively, such that the gradient directions of ik / idk are maximized.

[0065] The adaptive weight fine-tuning in step 3 includes:

[0066] Each sample is assigned a weight, and the confidence threshold and temperature parameter are optimized using meta-learning. The THS index on the validation set is updated in reverse to automatically find the optimal value. During training, the loss of the idk samples is weighted to mitigate excessive rejection.

[0067] Specifically, adaptive weight fine-tuning first involves adjusting the weights for each x∈ Perform stability impact calculations:

[0068]

[0069] Then, meta-learning parameter tuning is performed, that is, on a small-scale validation set, the confidence threshold TC and temperature parameter τ are updated in reverse with the TruthfulHelpfulness Score (THS) as the meta-objective.

[0070] Reweight normalization:

[0071]

[0072] in, This represents a temporary variable used during the summation and iteration process.

[0073] Finally, a weighted supervised fine-tuning (SFT) loss is applied:

[0074]

[0075] right Fine-tune the model to obtain the final model θ*.

[0076] The implementation of the various embodiments of the present invention is based on programmed processing through a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of the present invention are encapsulated into various modules. Based on this reality, and building upon the above embodiments, the embodiments of the present invention provide a gradient-driven rejection-aware intelligent question-answering model fine-tuning system, which is used to execute a gradient-driven rejection-aware intelligent question-answering model fine-tuning method from the above method embodiments.

[0077] See Figure 2 The system includes: a first main module for constructing a rejection-aware dataset, including known correct samples and unknown samples to be rejected; a second main module for constructing a rejection influence formula guided by second-order gradients, using gradient-driven sample selection to form a distillation dataset; a third main module for assigning a weight to each sample in the distillation dataset, using meta-learning to optimize the confidence threshold and temperature, and performing adaptive weight fine-tuning; and a fourth main module for completing the deployment and fine-tuning of the large model, outputting new model weights.

[0078] This invention provides a gradient-driven rejection-aware intelligent question-answering model fine-tuning system, addressing the difficulty in balancing "reducing illusions" and "avoiding over-rejection" in existing technologies. Figure 2 Several modules within the framework provide a novel fine-tuning framework that, while efficiently utilizing training samples, achieves a balance between "reducing illusions" and "avoiding excessive rejection" through more refined sample selection and weighting strategies.

[0079] It should be noted that the system embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The only difference is that corresponding functional modules are set. The principle is basically the same as that of the above system embodiments provided by the present invention. As long as those skilled in the art can improve the modules in the above system embodiments by referring to the specific technical solutions in other method embodiments and combining technical features to obtain corresponding technical means and technical solutions composed of these technical means, on the basis of the above system embodiments, and on the premise of ensuring the practicality of the technical solutions, they can obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.

[0080] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3As shown, the electronic device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can call logical instructions in the memory to execute a method for intelligent drone-assisted disaster-stricken road damage detection and repair.

[0081] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0082] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the intelligent drone-assisted method for detecting and repairing road damage after disasters provided by the above methods.

[0083] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent unmanned aerial vehicle-assisted method for detecting and repairing road damage after disasters provided by the methods described above.

[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0085] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0086] In summary, this invention belongs to the fields of large-scale language models, artificial intelligence, and machine learning. The provided gradient-driven rejection-aware intelligent question-answering model fine-tuning method includes the following steps: using a pre-trained large-scale language model to infer the original training set, calculating the answer accuracy of each sample, and dividing the samples into a known sample set (ik) and an unknown sample set (idk) according to a set threshold; expanding the boundary-ambiguous samples of the unknown sample set using adversarial generation technology and merging them with the original unknown samples; calculating the rejection influence of each unknown sample based on the first-order gradient inner product and the second-order Hessian-vector product, sorting the known samples by accuracy, and selecting the optimal subsets; introducing gradient contrastive learning loss into the candidate sample set to maximize the separation of gradient directions between known and unknown samples; automatically optimizing the confidence threshold TC and temperature parameter τ on a small-scale validation set with the Truthful Helpfulness Score as the meta-objective; calculating the stable influence of unknown samples and normalizing it to weight ω; and finally performing weighted supervised fine-tuning on the selected known and unknown samples. The system described in this invention includes a data acquisition module, a gradient calculation module, a sample selection module, a meta-learning module, a weighted fine-tuning module, and a model deployment module. Compared with existing technologies, this invention can significantly reduce the model hallucination error rate while effectively suppressing excessive rejection, thereby improving the model's response rate and reliability.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A gradient-driven method for fine-tuning a rejection-aware intelligent question-answering model, characterized in that, include: Construct a rejection-aware dataset, including known correct samples and unknown samples to be rejected; We construct a rejection influence formula guided by second-order gradient, and use gradient-driven sample selection to form a distillation dataset. Each sample in the distillation dataset is assigned a weight, and meta-learning is used to optimize the confidence threshold and temperature for adaptive weight fine-tuning. Complete the deployment and fine-tuning of the large model, and output the new model weights.

2. The method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven approach according to claim 1, characterized in that, Construct a rejection-aware dataset, including: Query the internal state of the LLM, and divide the original training samples into known correct samples and unknown rejection samples based on the query results. The known correct samples are those in which the model answers correctly, and the original answers are retained. The unknown rejection samples are those in which the model answers incorrectly, and their answers are replaced with rejection responses.

3. The method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven approach according to claim 2, characterized in that, The method further includes: A batch of samples with ambiguous boundaries are generated using GANs or diffusion models and then merged into a set of unknown samples that need to be rejected.

4. The method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven approach according to claim 1, characterized in that, Construct a second-order gradient-guided rejection effect formula and utilize gradient-driven sample selection, including: Construct the rejection effect formula of the second gradient, and calculate the contribution of each unknown sample to be rejected to minimizing the inaccuracy; Gradient contrast loss is introduced for candidate known correct samples / unknown rejection samples to learn, so that the gradients within known correct samples are close to each other, and the gradients between known correct samples and unknown rejection samples are separated to the maximum extent, thereby enhancing the distinguishability of static conflicts. For the unknown samples that need to be rejected after the influence ranking, fine-tuning is performed to form a distillation dataset for target training.

5. The method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven approach according to claim 1, characterized in that, Adaptive weight fine-tuning includes: Each sample is assigned a weight, and the confidence threshold and temperature parameter are optimized using meta-learning. The THS index on the validation set is updated in reverse to automatically find the optimal value. During training, the loss of unknown samples that need to be rejected is weighted to mitigate over-rejection.

6. The method for fine-tuning a rejection-aware intelligent question-answering model based on gradient-driven approach according to claim 5, characterized in that, The method further includes: Before performing meta-learning hyperparameter tuning, the stability effect is calculated for each sample in the distillation dataset.

7. A gradient-driven rejection-aware intelligent question-answering model fine-tuning system, characterized in that, include: The first main module is used to construct the rejection-aware dataset, which includes known correct samples and unknown samples that need to be rejected; The second main module is used to construct the rejection effect formula guided by the second-order gradient, and to use gradient-driven sample selection to form a distillation dataset. The third main module is used to assign a weight to each sample in the distillation dataset, and to use meta-learning to optimize the confidence threshold and temperature for adaptive weight fine-tuning. The fourth main module is used to complete the deployment and fine-tuning of the large model and output new model weights.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a gradient-driven rejection-aware intelligent question-answering model fine-tuning method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a gradient-driven rejection-aware intelligent question-answering model fine-tuning method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a gradient-driven rejection-aware intelligent question-answering model fine-tuning method as described in any one of claims 1 to 6.