Large language model fine tuning method and related equipment

By acquiring questions and their confidence levels to filter target questions, constructing answer generation instructions, and fine-tuning a large language model multiple times, the model illusion problem was solved, and its answering ability in domains with high accuracy requirements was improved.

CN120873121APending Publication Date: 2025-10-31ZHEJIANG E COMMERCE BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510902335.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Large language models tend to fabricate answers rather than admit ignorance when they lack internal knowledge, leading to the illusion problem and affecting their usability in domains with high accuracy requirements.

Method used

By acquiring questions and their confidence levels, target questions that meet the confidence level conditions are selected, answer generation instructions are constructed, and the large language model is fine-tuned multiple times to output the exact answer or admit ignorance. An unsupervised method is used to express self-awareness by utilizing the internal signals of the model.

Benefits of technology

It improves the usability of large language models in domains with high accuracy requirements, avoids labeling bias through unsupervised methods, and enhances the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873121A_ABST
    Figure CN120873121A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large language model fine tuning method and related equipment, and relates to the field of model training. In the embodiment of the specification, answers corresponding to a plurality of questions and the confidence degree of the answer corresponding to each question are obtained through a to-be-trained large language model, at least one target question in the questions is further screened out according to a preset confidence degree condition, and the target question serves as a sample for fine tuning of the to-be-trained large language model. And constructing an answer generation instruction corresponding to the at least one target question, and finely tuning the to-be-trained large language model for multiple times according to the at least one answer generation instruction, so that the to-be-trained large language model outputs an exact answer or is acknowledged based on the answer generation instruction, and obtaining the large language model until a fine tuning condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual relates to the field of model training, and in particular to a method for fine-tuning large language models and related equipment. Background Technology

[0002] Large Language Models (LLMs) have achieved state-of-the-art results on a wide range of Natural Language Processing (NLP) tasks, and applications based on LLMs are beginning to be widely used in various fields. However, LLMs suffer from the illusion problem, where the model may generate content that contradicts its internal knowledge and external facts. In applications, fabricated incorrect answers can lead to significant losses, and the illusion problem impairs the model's usability in domains where high accuracy in answering is required.

[0003] A key reason for this illusion is that when a model lacks the necessary internal parameter knowledge to answer a question, it tends to fabricate an answer rather than acknowledge its ignorance. This means the model cannot accurately represent its own knowledge boundaries and respond correctly to the input question based on its own knowledge. Summary of the Invention

[0004] This specification provides a method and related equipment for fine-tuning a large language model, which can solve the above-mentioned problems. The technical solution is as follows: Firstly, embodiments of this specification provide a method for fine-tuning a large language model, the method comprising: Multiple questions are obtained and input into a large language model to be trained to obtain the answers to the questions and the confidence scores of the answers. Based on the answers and confidence levels corresponding to the multiple questions, obtain at least one target question whose confidence level meets the confidence level condition; Construct at least one answer generation instruction corresponding to each of the target questions; The large language model to be trained is fine-tuned multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

[0005] Secondly, embodiments of this specification provide a large language model fine-tuning device, the device comprising: The confidence calculation module is used to obtain multiple questions, input the questions into the large language model to be trained, and obtain the answer to the question and the confidence of the answer; The target acquisition module is used to acquire at least one target question whose confidence level meets the confidence level condition based on the answers and confidence levels corresponding to the multiple questions respectively; The instruction construction module is used to construct at least one answer generation instruction corresponding to each of the target questions; The parameter fine-tuning module is used to fine-tune the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

[0006] Thirdly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.

[0007] Fourthly, embodiments of this specification provide a computer program product that stores multiple instructions adapted for loading by a processor and executing the above-described method steps.

[0008] Fifthly, embodiments of this specification provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.

[0009] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following: In the embodiments of this specification, multiple answers to questions and the confidence level of each answer are obtained through the large language model to be trained. Further, at least one target question is selected from the multiple questions based on preset confidence conditions. This target question serves as a sample for fine-tuning the large language model to be trained. In other words, this specification uses an unsupervised method to obtain samples. During training, the labeled questions used as samples are not required. Instead, the large language model's self-awareness in natural language form is fine-tuned by probing its internal signals. Since there is no labeling bias in a specific fine-tuning sample set, the fine-tuned large language model exhibits better generalization.

[0010] Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model to be trained is then fine-tuned multiple times based on this instruction, so that the model outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met, resulting in the final large language model. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary representation and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the architecture of a large language model fine-tuning method provided in the embodiments of this specification; Figure 2 This is a flowchart illustrating a method for fine-tuning a large language model provided in the embodiments of this specification; Figure 3 This is a schematic diagram illustrating a process of inputting a question into a large language model to be trained and obtaining the answer and confidence score, as provided in the embodiments of this specification. Figure 4 This is a flowchart illustrating a process for fine-tuning a large language model to be trained, as provided in the embodiments of this specification. Figure 5 This is a flowchart illustrating a method for fine-tuning a large language model provided in the embodiments of this specification; Figure 6 This is a flowchart illustrating an example of constructing an answer generation instruction provided in an embodiment of this specification; Figure 7 This is a flowchart illustrating a method for fine-tuning a large language model provided in the embodiments of this specification; Figure 8 This is a schematic diagram of the structure of a large language model fine-tuning device provided in the embodiments of this specification; Figure 9 This is a schematic diagram of the structure of an electronic device provided in the embodiments of this specification. Detailed Implementation

[0013] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0014] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0015] The present specification will now be described in detail with reference to specific embodiments.

[0016] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the features, information, and data involved in this specification were all obtained under full authorization.

[0017] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a large language model fine-tuning method provided in the embodiments of this specification. Figure 1 It includes at least a server 101 that performs large language model fine-tuning methods, and also includes multiple electronic devices for uploading questions as samples or uploading questions to be answered.

[0018] The multiple electronic devices include at least electronic device 1021, electronic device 1022, and electronic device 1023. It is understood that... Figure 1 The number of servers and electronic devices shown is for illustrative purposes only, and the embodiments in this specification do not impose any limitations on them.

[0019] The aforementioned server 101 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, where each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services to the outside world independently. Providing services independently can be understood as not requiring the assistance of other servers.

[0020] For example, a server can be multiple physical servers, each with independent hardware. Alternatively, a server can be multiple virtual servers deployed within the same hardware resource pool. Virtual server deployment methods include, but are not limited to, VMware, VirtualBox, and Virtual PC.

[0021] It is understood that server 101 also possesses other service capabilities and functions to complete the tasks described in the following embodiments. For example, server 101 also provides portal services, resource management services, and CI / CD services, etc.

[0022] Electronic devices include, but are not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.

[0023] In the embodiments of this specification, electronic devices such as electronic devices 1021, 1022, and 1023 may also be equipped with display devices. The display devices can be various devices capable of displaying functions, such as cathode ray tube displays (CR), light-emitting diode displays (LED), electronic ink screens, liquid crystal displays (LCD), plasma display panels (PDP), etc.

[0024] Users can use the display device on electronic device 1021 to send multiple questions and fine-tuning requests to the server 101 for fine-tuning the large language model to be trained. The server 101 then fine-tunes the large language model based on these requests and the multiple questions to obtain the large language model. Users can also use the display device on electronic device 1021 to send unanswered questions to the server 101, causing the server 101 to input the unanswered questions into the trained large language model, obtaining either a definitive answer to the unanswered question or an answer from the large language model acknowledging its lack of knowledge about the unanswered question.

[0025] Multiple electronic devices and multiple servers can communicate through communication links established by communication protocols. For example, the network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0026] In one embodiment, such as Figure 2 The diagram shown is a flowchart illustrating a large language model fine-tuning method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a large language model fine-tuning device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application.

[0027] Specifically, the fine-tuning method for this large language model includes: S102. Obtain multiple questions and input them into the large language model to be trained to obtain the answers to the questions and the confidence scores of the answers.

[0028] Obtaining multiple questions refers to collecting multiple questions that need to be answered from a single data source (such as user questions, questionnaires, online forums, etc.). These questions can be from any field, such as technical questions, academic questions, or everyday life questions. Typically, questions can be collected through: user interaction, such as through chatbots or Q&A platforms where users input their questions; datasets, using existing public datasets that contain a large number of question-and-answer pairs; and automated tools, such as web crawlers, to scrape frequently asked questions or user queries from the web.

[0029] like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating a process of inputting a question into a large language model to be trained to obtain the answer and confidence score, as provided in an embodiment of this specification. In this embodiment, the multiple questions obtained include at least question 2011, question 2012, question 2013, and question 2014.

[0030] The type of large language model to be trained can be any of the following: GPT series (e.g., GPT3, GPT4, etc.), chatGPT series, BERT (Bidirectional Encoder Representations from Transformers) series, LaMDA (Language Model for Dialogue Applications), T5, etc. The embodiments in this specification do not impose any restrictions on this.

[0031] Multiple questions are input into the large language model to be trained. The model parses the input question text and generates relevant output text as answers; this input process is the model's inference phase. In this specification, the large language model also outputs the confidence score for each answer. The confidence score refers to the degree of certainty the large language model is about the answer it generates. Confidence score is usually a numerical value representing the model's confidence in a given answer. For example, suppose the large language model gives the answer "42" with a confidence score of 90%, meaning the model believes there is a 90% probability that the answer to the question is correct.

[0032] In some models, confidence is typically represented in the following ways: probability values, where the large language model being trained can assign a probability value to each possible answer, representing the likelihood of that answer; the higher the probability value, the stronger the confidence of the answer; activation values ​​of the model's output layer, for example, the output of the last layer of the large language model being trained might be a probability distribution, and the confidence can be determined by the maximum value of this distribution; or a pre-defined scoring system, where the large language model being trained scores the quality of answers based on historical data and correlates the scores with the confidence.

[0033] like Figure 3 As shown, questions 2011, 2012, 2013, and 2014 are input into the large language model 202 to be trained, respectively, to obtain the answer 2031 and the confidence score 2041 for question 2011, the answer 2032 and the confidence score 2042 for question 2012, the answer 2033 and the confidence score 2043 for question 2013, and the answer 2031 and the confidence score 2041 for question 2014.

[0034] S104. Based on the answers and confidence levels corresponding to multiple questions, obtain at least one target question whose confidence level meets the confidence level condition.

[0035] For each question, along with its corresponding answer and confidence level, determine whether the answer's confidence level meets a preset confidence level condition. If the answer's confidence level meets the preset condition, designate the question corresponding to that answer as the target question. The target question serves as a sample for subsequent training of the large language model.

[0036] The confidence level condition serves as a selection criterion for samples; only questions meeting this condition can be used as samples to train the large language model. Furthermore, this specification employs an unsupervised method to obtain samples. During training, the labeled questions used as samples are not required; instead, the large language model's self-awareness in natural language form is fine-tuned by probing its internal signals. Because there is no labeling bias specific to the fine-tuning sample set, the fine-tuned large language model exhibits better generalization ability.

[0037] In one embodiment, the confidence level conditions include a confidence level less than a first confidence threshold and / or a confidence level greater than a second confidence threshold. The first confidence threshold is a smaller threshold, and the second confidence threshold is a larger threshold. For example, the first confidence threshold is 0.4, and the second confidence threshold is 0.99. When the confidence level of an answer is determined to be greater than the second confidence threshold, the question corresponding to that answer is determined to be the target question. When the confidence level of an answer is determined to be less than the first confidence threshold, the question corresponding to that answer is determined to be the target question.

[0038] When the confidence level of an answer to a question is less than the first confidence threshold, the trained large language model should acknowledge its ignorance in natural language and output an answer text such as "I don't know." When the confidence level of an answer to a question is greater than the second confidence threshold, the trained large language model should output an exact answer. Questions with confidence levels greater than the first and less than the second are considered to contain too much noise and are therefore removed.

[0039] S106. Construct at least one answer generation instruction corresponding to each target question.

[0040] For each target question, an answer generation instruction is constructed corresponding to that target question. The difference between using an answer generation instruction and directly inputting the target question into the large language model to be trained is that the answer generation instruction instructs the large language model to output either an exact answer or an admission of ignorance, whereas directly inputting the target question into the large language model instructs it to output an answer, even if that answer is fabricated.

[0041] Constructing an answer generation instruction corresponding to the target question can involve obtaining a preset instruction template and replacing the corresponding position in the instruction template with the target question.

[0042] S108. Based on at least one answer generation instruction, fine-tune the large language model to be trained multiple times, so that the large language model to be trained outputs the exact answer or admits ignorance based on the answer generation instruction, until the fine-tuning condition is met to obtain the large language model.

[0043] Each answer generation instruction is input multiple times into the large language model to be trained, fine-tuning the target parameters of the model so that it outputs an exact answer to the target question or admits ignorance based on the instruction. The large language model is then fine-tuned multiple times based on a preset loss function until the fine-tuning conditions are met, resulting in the final large language model. These conditions can be any one or more of the following: the number of training iterations reaches a preset number, the loss function converges, or the training effect of the pre-trained large language model reaches the expected level.

[0044] like Figure 4 As shown, Figure 4 This is a flowchart illustrating a process for fine-tuning a large language model to be trained, as provided in an embodiment of this specification. Taking question 2011 as an example, an answer generation instruction 205 is constructed for question 2011 as the target question. The answer generation instruction 205 is input into the large language model to be trained 202 to obtain an exact answer to question 2011, or the large language model to be trained 202 acknowledges its ignorance of question 2011 in natural language form. The large language model to be trained 202 is further fine-tuned until the fine-tuning conditions are met to obtain the large language model.

[0045] The trained large language model is used to output an exact answer to a question, or to admit ignorance. The question can be any question input into the large language model.

[0046] In one embodiment, based on at least one answer generation instruction, the low-rank parameters in the large language model to be trained are fine-tuned multiple times to make the large language model to be trained output the exact answer or admit ignorance based on the answer generation instruction, until the fine-tuning condition is met to obtain the large language model.

[0047] Low-rank adaptation is a fine-tuning method that adjusts the performance of a large language model by adding low-rank LOA layers (i.e., adding trainable LOA parameters) to its self-attention layers. During fine-tuning, other parameters of the large language model need to be frozen, such as all weight layers, including weight matrices, biases, and other parameters closely related to the model structure.

[0048] Furthermore, a low-rank matrix is ​​added to the self-attention layer. During multiple fine-tuning processes, the usual backpropagation algorithm is used, but gradient updates are performed only on the low-rank parameters in the Lora layer. The low-rank matrix is ​​typically a very small matrix with a rank smaller than the original matrix, thus reducing parameter complexity and computational cost.

[0049] In this embodiment, a LoRa layer is added to the self-attention layer parameters of the large language model to be trained. By repeatedly fine-tuning the low-rank parameters of the LoRa layer in the large language model to be trained with low rank, other parameters in the large language model to be trained are frozen, preventing the fine-tuning from destroying the internal parameter knowledge of the large language model to be trained. This method can improve the fine-tuning efficiency, save computational resources, and enable the trained large language model to adapt to new tasks.

[0050] In the embodiments of this specification, multiple answers to questions and the confidence level of each answer are obtained through the large language model to be trained. Further, at least one target question is selected from the multiple questions based on preset confidence conditions. This target question serves as a sample for fine-tuning the large language model to be trained. In other words, this specification uses an unsupervised method to obtain samples. During training, the labeled questions used as samples are not required. Instead, the large language model's self-awareness in natural language form is fine-tuned by probing its internal signals. Since there is no labeling bias in a specific fine-tuning sample set, the fine-tuned large language model exhibits better generalization.

[0051] Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model to be trained is then fine-tuned multiple times based on this instruction, so that the model outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met, resulting in the final large language model. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary representation and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy.

[0052] In one embodiment, such as Figure 5 The diagram shown is a flowchart illustrating a large language model fine-tuning method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a large language model fine-tuning device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application.

[0053] Specifically, the fine-tuning method for this large language model includes: S202. Obtain multiple questions and input them into the large language model to be trained to obtain the answers to the questions and the confidence scores of the answers.

[0054] See S102 above, which will not be repeated here.

[0055] S204. Based on the answers and confidence levels corresponding to multiple questions, obtain at least one target question whose confidence level meets the confidence level condition.

[0056] See S104 above, which will not be repeated here.

[0057] S206. Construct multiple answer generation instructions of different types corresponding to the target question, and obtain multiple answer generation instructions corresponding to at least one target question.

[0058] In this context, different types of answer generation instructions instruct the large language model to output different content of the desired fine-tuning answer for the target question. The desired fine-tuning answer may include either a definitive answer to the target question or an admission of ignorance. In other words, multiple different types of answer generation instructions are constructed for the target question. The commonality among these instructions is that they all instruct the large language model to output a desired fine-tuning answer for the target question. The differences lie in the different logical reasoning processes instructing the large language model to perform on the target question, and in the different content of the desired fine-tuning answers for each instruction. This could even lead to multiple desired fine-tuning answers for the target question including both a definitive answer and an admission of ignorance.

[0059] In this embodiment, by constructing multiple different types of answer generation instructions for a target question, and then fine-tuning the large language model to be trained based on these instructions, the large language model can learn different ways of thinking, problem-solving strategies, and expressions. Furthermore, having multiple types of answer generation instructions helps the large language model understand the multi-layered meanings and complexity of the question, enabling it to generalize better when dealing with new questions, rather than being limited to a single fixed answer.

[0060] In one embodiment, multiple different types of instruction construction templates are used to construct multiple different types of answer generation instructions corresponding to the target question, thereby obtaining multiple answer generation instructions corresponding to at least one target question; wherein, the multiple instruction construction templates include at least one of the following instruction construction templates: a priori self-knowledge instruction construction template, direct self-knowledge instruction construction template, and a posteriori self-knowledge construction template.

[0061] like Figure 6 As shown, Figure 6 This is a flowchart illustrating an example of constructing an answer generation instruction provided in this specification. For question 301, an answer generation instruction 3021 is constructed based on a priori self-knowledge instructions, an answer generation instruction 3022 is constructed based on a direct self-knowledge instructions, and an answer generation instruction 3023 is constructed based on a posteriori self-knowledge instructions.

[0062] Instruction templates are designed to ensure the quality, accuracy, and diversity of answers. A priori self-awareness instruction templates are based on pre-existing knowledge or assumptions, instructing the large language model to deduce and generate relevant answers based on prior knowledge. Direct self-awareness instruction templates, characterized by instructing the large language model to generate answers based on direct information or actual circumstances provided by the user, typically rely on the current context or explicit user requirements. Posterior self-awareness instruction templates generate answers based on posterior analysis or feedback, instructing the large language model to deduce possible answers or suggestions by analyzing past events, results, or data.

[0063] For example, for the target question "improve work efficiency," the answer generation instruction built based on the prior knowledge instruction template is "What are some successful ways to improve work efficiency?" This answer generation instruction will instruct the large language model to generate an answer based on already verified time management techniques or to acknowledge ignorance.

[0064] Furthermore, for the target question, the answer generation instruction based on the template constructed from direct self-awareness is "I have severe procrastination problems, how can I improve my work efficiency?" This answer generation instruction will instruct the large language model to combine the preceding information "I have severe procrastination problems" to output a precise answer to the target question.

[0065] Furthermore, for the target question, the answer generation instruction, based on the posterior self-awareness instruction template, is "Why is my work efficiency low, and are there ways to improve it?" This answer generation instruction will instruct the large language model to analyze information such as past work records and analysis results to obtain a precise answer to the target question or acknowledge its ignorance.

[0066] S208. Input the multiple answer generation instructions corresponding to the target question into the large language model to be trained, and fine-tune the large language model to be trained multiple times based on the consistency loss function so that the multiple answers to be fine-tuned output by the large language model to be trained based on the multiple answer generation instructions have consistency until the fine-tuning condition is met to obtain the large language model.

[0067] The consistency loss function measures the consistency among multiple answers to be fine-tuned when the large language model being trained outputs different types of answer generation instructions.

[0068] In other words, for the same question, even if different answer generation instructions are provided to the large language model to be trained, will the output of the large language model remain consistent? For example, if a question has answer generation instructions based on templates built from direct self-knowledge and answer generation instructions based on templates built from posterior self-knowledge, then the consistency loss function will help fine-tune the large language model to be trained, so that the core information of the two outputs of the answer to be fine-tuned remains consistent, without significant discrepancies.

[0069] During training, the large language model under training adjusts itself based on feedback from the consistency loss function. After each adjustment, the large language model attempts to generate more consistent answers for fine-tuning. Through multiple iterations of optimization, the large language model gradually learns how to maintain consistent output when faced with different types of answer generation instructions, until the fine-tuning meets the fine-tuning conditions, at which point the fine-tuning process ends.

[0070] This specification employs an unsupervised method to acquire samples. Training does not require the labeled questions used as samples; instead, it fine-tunes the large language model's ability to express its self-awareness in natural language by probing its internal signals. Because there is no labeling bias specific to the fine-tuning sample set, the fine-tuned large language model exhibits better generalization. Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model under training is then fine-tuned multiple times based on this instruction, ensuring that it outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary expression and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy.

[0071] In one embodiment, such as Figure 7 The diagram shown is a flowchart illustrating a large language model fine-tuning method provided in an embodiment of this specification. This method can be implemented using a computer program and can run on a large language model fine-tuning device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application.

[0072] Specifically, the fine-tuning method for this large language model includes: S302, Obtain multiple questions.

[0073] Obtaining multiple questions refers to collecting multiple questions that need to be answered from a single data source (such as user questions, questionnaires, online forums, etc.). These questions can be from any field, such as technical questions, academic questions, or everyday life questions. Typically, questions can be collected through: user interaction, such as through chatbots or Q&A platforms where users input their questions; datasets, using existing public datasets that contain a large number of question-and-answer pairs; and automated tools, such as web crawlers, to scrape frequently asked questions or user queries from the web.

[0074] S304. Input the question into the large language model to be trained to obtain the answer to the question.

[0075] Multiple questions are input into the large language model to be trained. The large language model parses the input question text and generates relevant output text as the answer. This input process is the inference stage of the model.

[0076] S306. Calculate multiple confidence levels corresponding to the answer based on multiple confidence level calculation methods.

[0077] The confidence level of an answer refers to the degree of certainty that the large language model being trained has regarding the answer it generates. In this embodiment, multiple confidence level calculation methods are preset to calculate the reliability of the answer as the correct answer to the question. For example, for the answer to the question, three confidence levels can be obtained through multiple confidence level calculation methods, namely 0.3, 0.4, and 0.6.

[0078] In one embodiment, the question and constraint instructions are input into the large language model to be trained to obtain the answer in phrase form for the question; according to multiple confidence calculation methods, the multiple phrases included in the answer in phrase form are calculated to obtain the multiple confidence scores corresponding to the answer.

[0079] Constraints instruct the large language model being trained to output phrase-based answers. These constraints can include word limits, specific formatting requirements, and output range limitations. The answers generated by the large language model are not complete sentences or paragraphs, but rather short phrases or keywords. Phrases typically provide concise, key information that quickly offers relevant solutions. For example, for the question "How to improve work efficiency," constraints might generate phrase-based answers such as "Set clear goals" or "Use the Pomodoro Technique." Furthermore, based on multiple confidence calculation methods, the confidence scores of the multiple phrases included in the phrase-form answer are calculated separately to obtain the multiple confidence scores corresponding to the answer.

[0080] In one embodiment, the multiple confidence calculation methods include at least one or more of the following calculation methods: calculating the minimum probability among the probabilities of decoding the problem for each of the multiple phrases, calculating the geometric mean of the probabilities of decoding the problem for each of the multiple phrases, and calculating the probability of decoding the problem for the target phrase among the multiple phrases.

[0081] The algorithm calculates the minimum probability among multiple phrases that individually decode the question. Specifically, for each generated phrase, it calculates the probability that it decodes the question (i.e., the confidence score of the large language model to be trained for each phrase). Among the multiple phrases, the one with the lowest probability is selected as the overall confidence score. This approach helps identify the least likely phrase, thus avoiding the influence of incorrect answers on the overall result.

[0082] The geometric mean method considers the probability of each phrase decoding problem, aiming to synthesize the confidence of multiple phrases. Unlike the arithmetic mean, the geometric mean places more emphasis on extreme values ​​(very low or very high probabilities), thus it can better reflect the overall confidence situation and avoid the excessive influence of certain individual extreme values ​​(such as particularly untrustworthy phrases) on the final result.

[0083] This approach considers the probability of a target phrase decoding a question across multiple phrases, focusing on the confidence level of a specific "target phrase"—a key part of the target answer. The large language model being trained calculates the probability of the target phrase decoding the question to assess its credibility. This helps filter out the most relevant and reliable phrases, especially when the question involves specific key concepts or information.

[0084] By using these different confidence calculation methods, we can more comprehensively assess the confidence level of the answer and ensure that the obtained confidence level is more accurate and reliable, especially in the case of multi-phrase generation.

[0085] S308. Based on the answers to the questions and multiple confidence levels, identify the questions with confidence levels that meet the confidence level conditions among the multiple confidence levels corresponding to the questions as target questions, and obtain at least one target question.

[0086] Based on the answers to the question and multiple confidence levels, determine whether any of the confidence levels meet the confidence criteria. If no confidence level meets the confidence criteria, then the question is identified as the target question.

[0087] For example, confidence conditions include a confidence level less than a first confidence threshold and / or a confidence level greater than a second confidence threshold. The first confidence threshold is 0.4 and the second confidence threshold is 0.99. For the answer to the question, through multiple confidence calculation methods, three confidence levels can be obtained, which are 0.3, 0.4 and 0.6 respectively. Then, the question is determined to be the target question.

[0088] In this embodiment, multiple confidence scores are obtained for the answer by setting multiple confidence score calculation methods. When determining whether a question is a target question based on the confidence score, it is determined whether there is a confidence score that meets the confidence score condition among the multiple confidence scores of the answer, thereby determining whether the question is a target question. This can reduce the possibility of removing the question as noise and enrich the types and number of target questions used as training samples.

[0089] S310. Construct at least one answer generation instruction corresponding to each target question.

[0090] See S106 above; it will not be repeated here.

[0091] S312. Based on at least one answer generation instruction, fine-tune the large language model to be trained multiple times, so that the large language model to be trained outputs the exact answer or admits ignorance based on the answer generation instruction, until the fine-tuning condition is met to obtain the large language model.

[0092] See S108 above; it will not be repeated here.

[0093] This specification employs an unsupervised method to acquire samples. Training does not require the labeled questions used as samples; instead, it fine-tunes the large language model's ability to express its self-awareness in natural language by probing its internal signals. Because there is no labeling bias specific to the fine-tuning sample set, the fine-tuned large language model exhibits better generalization. Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model under training is then fine-tuned multiple times based on this instruction, ensuring that it outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary expression and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy.

[0094] The following are embodiments of the apparatus described in this specification, which can be used to execute the embodiments of the methods described in this specification. For details not disclosed in the apparatus embodiments of this specification, please refer to the embodiments of the methods described in this specification.

[0095] Please see Figure 8This diagram illustrates the structure of a large language model fine-tuning device provided in an exemplary embodiment of this specification. The large language model fine-tuning device can be implemented as all or part of a device through software, hardware, or a combination of both. The device includes a confidence calculation module 401, a target acquisition module 402, an instruction construction module 403, and a parameter fine-tuning unit 404.

[0096] The confidence calculation module 401 is used to acquire multiple questions, input the questions into the large language model to be trained, and obtain the answer to the question and the confidence of the answer; The target acquisition module 402 is used to acquire at least one target question whose confidence level meets the confidence level condition based on the answers and confidence levels corresponding to the multiple questions respectively; Instruction construction module 403 is used to construct at least one answer generation instruction corresponding to each of the target questions; The parameter fine-tuning module 404 is used to fine-tune the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

[0097] In one embodiment, the instruction construction module 403 includes: An instruction construction unit is used to construct multiple different types of answer generation instructions corresponding to the target question, thereby obtaining at least one set of multiple answer generation instructions corresponding to the target question; wherein, the different types of answer generation instructions instruct the large language model to be trained to output different specific contents of the answers to be fine-tuned for the target question, and the answers to be fine-tuned include the exact answer to the target question or an admission of ignorance.

[0098] In one embodiment, the parameter fine-tuning module 404 includes: The consistency fine-tuning unit is used to input multiple answer generation instructions corresponding to the target question into the large language model to be trained, and to fine-tune the large language model to be trained multiple times based on the consistency loss function, so that the multiple answers to be fine-tuned output by the large language model to be trained based on the multiple answer generation instructions have consistency, until the fine-tuning condition is met to obtain the large language model.

[0099] In one embodiment, the instruction building unit includes: The instruction construction subunit is used to construct multiple different types of answer generation instructions corresponding to the target question based on multiple different types of instruction construction templates, thereby obtaining at least one of the multiple answer generation instructions corresponding to the target question; wherein, the multiple instruction construction templates include at least one of the following instruction construction templates: a priori self-knowledge instruction construction template, direct self-knowledge instruction construction template, and a posteriori self-knowledge construction template.

[0100] In one embodiment, the confidence calculation module 401 includes: The first acquisition unit is used to acquire multiple questions; The second acquisition unit is used to input the question into the large language model to be trained and obtain the answer to the question; The third acquisition unit is used to calculate multiple confidence levels corresponding to the answer based on multiple confidence level calculation methods; Target acquisition module 402 includes: The target acquisition subunit is used to determine, based on the answer to the question and multiple confidence levels, a question with a confidence level that meets the confidence level conditions among the multiple confidence levels corresponding to the question as the target question, thereby obtaining at least one target question.

[0101] In one embodiment, the confidence calculation module 401 includes: The phrase answer acquisition unit is used to input the question and constraint instructions into the large language model to be trained, and obtain the answer in phrase form for the question; The confidence calculation subunit is used to calculate the confidence scores of multiple phrases included in the phrase-form answer according to multiple confidence calculation methods, so as to obtain multiple confidence scores corresponding to the answer.

[0102] In one embodiment, the plurality of confidence calculation methods include at least one or more of the following calculation methods: calculating the minimum probability among the plurality of phrases respectively decoding the question, calculating the geometric mean of the probability among the plurality of phrases respectively decoding the question, and the probability of the target phrase among the plurality of phrases decoding the question.

[0103] In one embodiment, the confidence level condition includes a confidence level less than a first confidence level threshold and / or a confidence level greater than a second confidence level threshold.

[0104] In one embodiment, the parameter fine-tuning module 404 includes: The low-rank fine-tuning unit is used to fine-tune the low-rank parameters in the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs the exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model.

[0105] This specification employs an unsupervised method to acquire samples. Training does not require the labeled questions used as samples; instead, it fine-tunes the large language model's ability to express its self-awareness in natural language by probing its internal signals. Because there is no labeling bias specific to the fine-tuning sample set, the fine-tuned large language model exhibits better generalization. Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model under training is then fine-tuned multiple times based on this instruction, ensuring that it outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary expression and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy.

[0106] It should be noted that the large language model fine-tuning device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the large language model fine-tuning method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the large language model fine-tuning device and the large language model fine-tuning method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0107] The example numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the examples.

[0108] This specification also provides a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figure 1 - Figure 7 The large language model fine-tuning method of the embodiment shown can be found in the following document for details: Figure 1 - Figure 7 The specific details of the illustrated embodiments will not be elaborated here.

[0109] This specification also provides a computer program product that stores at least one instruction, which is loaded and executed by a processor as described above. Figure 1 - Figure 7 The large language model fine-tuning method of the embodiment shown can be found in the following document for details: Figure 1 - Figure 7 The specific details of the illustrated embodiments will not be elaborated here.

[0110] Please see Figure 9 This document provides a schematic diagram of the structure of an electronic device as an embodiment of the present specification. Figure 9 As shown, the electronic device 500 may include: at least one processor 501, at least one network interface 504, user interface 503, memory 505, and at least one communication bus 502.

[0111] The communication bus 502 is used to enable communication between these components.

[0112] The user interface 503 may include a display screen and a camera. Optionally, the user interface 503 may also include a standard wired interface and a wireless interface.

[0113] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0114] The processor 501 may include one or more processing cores. The processor 501 connects to various parts of the server 500 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 505, and by calling data stored in the memory 505. Optionally, the processor 501 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 501 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 501 and may be implemented as a separate chip.

[0115] The memory 505 may include random access memory (RAM) or read-only memory. Optionally, the memory 505 may include a non-transitory computer-readable storage medium. The memory 505 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 505 may also be at least one storage device located remotely from the aforementioned processor 501. Figure 9 As shown, the memory 505, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a large language model training application.

[0116] exist Figure 9In the illustrated electronic device 500, the user interface 503 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 501 can be used to call the large language model training application stored in the memory 505 and specifically perform the following operations: Multiple questions are obtained and input into a large language model to be trained to obtain the answers to the questions and the confidence scores of the answers. Based on the answers and confidence levels corresponding to the multiple questions, obtain at least one target question whose confidence level meets the confidence level condition; Construct at least one answer generation instruction corresponding to each of the target questions; The large language model to be trained is fine-tuned multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

[0117] In one embodiment, processor 501 executes the instruction to generate answers corresponding to at least one of the target questions, specifically: Construct multiple different types of answer generation instructions corresponding to the target question to obtain at least one set of multiple answer generation instructions corresponding to the target question; wherein, the different types of answer generation instructions instruct the large language model to be trained to output different specific contents of the answers to be fine-tuned for the target question, and the answers to be fine-tuned include the exact answer to the target question or an admission of ignorance.

[0118] In one embodiment, processor 501 executes the process of fine-tuning the large language model to be trained multiple times based on at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model. Specifically, the following is executed: Multiple answer generation instructions corresponding to the target question are input into the large language model to be trained. The large language model to be trained is fine-tuned multiple times based on the consistency loss function so that the multiple answers to be fine-tuned output by the large language model to be trained based on the multiple answer generation instructions have consistency until the fine-tuning condition is met to obtain the large language model.

[0119] In one embodiment, processor 501 executes the instruction to generate multiple different types of answers corresponding to the target question, thereby obtaining at least one instruction to generate multiple answers corresponding to the target question, specifically: Based on multiple different types of instruction construction templates, multiple different types of answer generation instructions corresponding to the target question are constructed to obtain at least one multiple answer generation instruction corresponding to the target question; wherein, the multiple instruction construction templates include at least one of the following instruction construction templates: a priori self-knowledge instruction construction template, direct self-knowledge instruction construction template, and a posteriori self-knowledge construction template.

[0120] In one embodiment, processor 501 executes the process of acquiring multiple questions, inputting the questions into a large language model to be trained, and obtaining the answers to the questions and the confidence scores of the answers. Specifically, the process involves: Get multiple questions; The question is input into the large language model to be trained, and the answer to the question is obtained. Based on multiple confidence calculation methods, multiple confidence levels corresponding to the answer are calculated; Processor 501 executes the step of obtaining at least one target question whose confidence level meets the confidence level condition based on the answers and confidence levels corresponding to the multiple questions, specifically: Based on the answers to the question and multiple confidence levels, the question with a confidence level that meets the confidence level conditions is identified as the target question, thus obtaining at least one target question.

[0121] In one embodiment, processor 501 executes the process of acquiring multiple questions, inputting the questions into a large language model to be trained, and obtaining the answers to the questions and the confidence scores of the answers. Specifically, the process involves: Multiple questions are obtained, and the questions and constraint instructions are input into the large language model to be trained to obtain the answers in phrase form for the questions; Based on multiple confidence calculation methods, the confidence scores of the phrases included in the answer in the phrase form are calculated to obtain the multiple confidence scores corresponding to the answer.

[0122] In one embodiment, the plurality of confidence calculation methods include at least one or more of the following calculation methods: calculating the minimum probability among the plurality of phrases respectively decoding the question, calculating the geometric mean of the probability among the plurality of phrases respectively decoding the question, and the probability of the target phrase among the plurality of phrases decoding the question.

[0123] In one embodiment, the confidence level condition includes a confidence level less than a first confidence level threshold and / or a confidence level greater than a second confidence level threshold.

[0124] In one embodiment, processor 501 executes the process of fine-tuning the large language model to be trained multiple times based on at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model. Specifically, the following is executed: Based on at least one of the answer generation instructions, the low-rank parameters in the large language model to be trained are fine-tuned multiple times to make the large language model to be trained output an exact answer or admit ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model.

[0125] This specification employs an unsupervised method to acquire samples. Training does not require the labeled questions used as samples; instead, it fine-tunes the large language model's ability to express its self-awareness in natural language by probing its internal signals. Because there is no labeling bias specific to the fine-tuning sample set, the fine-tuned large language model exhibits better generalization. Furthermore, at least one answer generation instruction corresponding to each target question is constructed. The large language model under training is then fine-tuned multiple times based on this instruction, ensuring that it outputs an exact answer or admits ignorance based on the instruction, until the fine-tuning conditions are met. The fine-tuning method provided in this specification allows the large language model to learn to express its knowledge boundaries using natural language, ensuring consistency between its knowledge boundary expression and internal signals. This enables it to output a clear answer regarding whether it knows the relevant knowledge, thus resolving the illusion problem of large language models and preventing the use of fabricated answers from compromising its usability in domains requiring high accuracy.

[0126] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented. Each of the above methods can be executed by a computer program instructing related hardware. The program corresponding to each method can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The storage medium of the electronic device 700 can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0127] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0128] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.

Claims

1. A method for fine-tuning a large language model, the method comprising: Multiple questions are obtained and input into a large language model to be trained to obtain the answers to the questions and the confidence scores of the answers. Based on the answers and confidence levels corresponding to the multiple questions, obtain at least one target question whose confidence level meets the confidence level condition; Construct at least one answer generation instruction corresponding to each of the target questions; The large language model to be trained is fine-tuned multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

2. The large language model fine-tuning method according to claim 1, wherein constructing at least one answer generation instruction corresponding to each of the target questions includes: Construct multiple different types of answer generation instructions corresponding to the target question to obtain at least one set of multiple answer generation instructions corresponding to the target question; wherein, the different types of answer generation instructions instruct the large language model to be trained to output different specific contents of the answers to be fine-tuned for the target question, and the answers to be fine-tuned include the exact answer to the target question or an admission of ignorance.

3. The large language model fine-tuning method according to claim 2, wherein the step of fine-tuning the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model, includes: Multiple answer generation instructions corresponding to the target question are input into the large language model to be trained. The large language model to be trained is fine-tuned multiple times based on the consistency loss function so that the multiple answers to be fine-tuned output by the large language model to be trained based on the multiple answer generation instructions have consistency until the fine-tuning condition is met to obtain the large language model.

4. The large language model fine-tuning method according to claim 2, wherein constructing multiple different types of answer generation instructions corresponding to the target question to obtain at least one set of multiple answer generation instructions corresponding to the target question respectively includes: Based on multiple different types of instruction construction templates, multiple different types of answer generation instructions corresponding to the target question are constructed to obtain at least one multiple answer generation instruction corresponding to the target question; wherein, the multiple instruction construction templates include at least one of the following instruction construction templates: a priori self-knowledge instruction construction template, direct self-knowledge instruction construction template, and a posteriori self-knowledge construction template.

5. The large language model fine-tuning method according to claim 1, wherein obtaining multiple questions, inputting the questions into the large language model to be trained, and obtaining the answers to the questions and the confidence levels of the answers, includes: Get multiple questions; The question is input into the large language model to be trained, and the answer to the question is obtained. Based on multiple confidence calculation methods, multiple confidence levels corresponding to the answer are calculated; The step of obtaining at least one target question whose confidence level meets the confidence condition based on the answers and confidence levels corresponding to the multiple questions includes: Based on the answers to the question and multiple confidence levels, the question with a confidence level that meets the confidence level conditions is identified as the target question, thus obtaining at least one target question.

6. The method for fine-tuning a large language model according to claim 1, wherein obtaining multiple questions, inputting the questions into the large language model to be trained, and obtaining the answers to the questions and the confidence levels of the answers, includes: Multiple questions are obtained, and the questions and constraint instructions are input into the large language model to be trained to obtain the answers in phrase form for the questions; Based on multiple confidence calculation methods, the confidence scores of the phrases included in the answer in the phrase form are calculated to obtain the multiple confidence scores corresponding to the answer.

7. The large language model fine-tuning method according to claim 6, wherein the plurality of confidence calculation methods include at least one or more of the following calculation methods: calculating the minimum probability among the probabilities of the plurality of phrases respectively decoding the question, calculating the geometric mean of the probabilities of the plurality of phrases respectively decoding the question, and the probability of the target phrase among the plurality of phrases decoding the question.

8. The large language model fine-tuning method according to claim 1, wherein the confidence condition includes a confidence level less than a first confidence threshold and / or a confidence level greater than a second confidence threshold.

9. The large language model fine-tuning method according to claim 1, wherein the step of fine-tuning the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model, comprises: Based on at least one of the answer generation instructions, the low-rank parameters in the large language model to be trained are fine-tuned multiple times to make the large language model to be trained output an exact answer or admit ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model.

10. A large language model fine-tuning device, the device comprising: The confidence calculation module is used to obtain multiple questions, input the questions into the large language model to be trained, and obtain the answer to the question and the confidence of the answer; The target acquisition module is used to acquire at least one target question whose confidence level meets the confidence level condition based on the answers and confidence levels corresponding to the multiple questions respectively; The instruction construction module is used to construct at least one answer generation instruction corresponding to each of the target questions; The parameter fine-tuning module is used to fine-tune the large language model to be trained multiple times according to at least one of the answer generation instructions, so that the large language model to be trained outputs an exact answer or admits ignorance based on the answer generation instructions, until the fine-tuning conditions are met to obtain the large language model; wherein, the large language model is used to output an exact answer or admit ignorance for the question to be answered.

11. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 9.

12. A computer program product storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 9.

13. A printing device, characterized in that, It includes a first sensor, a printhead, a processor, and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1 to 9.