Method and device for training model, storage medium and electronic equipment

By dividing knowledge boundaries and building high-quality data sets, and training the model using direct preference optimization methods, so that it generates accurate answers within the knowledge boundaries and rejects answers when it exceeds the boundaries, the problem of model generation illusions in the existing technology is solved, and the robustness and credibility of the model in high-risk areas is improved.

CN120523957AActive Publication Date: 2025-08-22ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511006751.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-22
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Although the existing retrieval enhancement fine-tuning method improves the robustness of the model for noise retrieval, it causes the model to generate hallucinations when it lacks reliable knowledge, weakens the reliability of the model, and requires enhanced the robustness and credibility of the model in high-risk areas.

Method used

By dividing knowledge boundaries, identifying whether the query is within the knowledge scope of the model, building a high-quality data set, and using direct preference optimization methods to train the model, so that it generates accurate answers within the knowledge boundaries and rejects answers when it exceeds the boundaries, enhancing the model's rejection ability.

Benefits of technology

It effectively balances the accuracy and rejection ability of the model, enhances the robustness and credibility of the model in high-risk areas, and solves the problem of hallucination in existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523957A_ABST
    Figure CN120523957A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method and device for training a model, a storage medium and electronic equipment. The method comprises the steps that first output of a target model for a query sample under the condition that the target model does not depend on an external knowledge source and second output obtained by searching the query sample through the external knowledge source are obtained; determining a knowledge boundary division result corresponding to the query sample according to the first output, the second output and label information corresponding to the query sample; according to the third output of the target model for the query sample, the label information and the knowledge boundary division result, constructing a preference data set; and training the target model based on the preference data set through a direct preference optimization method to obtain a trained target model, so that the trained target model can output the answer rejection response information under the condition that an input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer technology, and in particular to a method, device, storage medium and electronic device for training a model. Background Art

[0002] With the advancement of natural language processing technology, large language models (LLMs), when combined with retrieval systems, can significantly improve their performance in natural language processing tasks by integrating external knowledge sources, resulting in more accurate and context-rich responses. However, while existing retrieval augmented fine-tuning (RAFT) methods—a training method that improves the model's robustness to suboptimal retrieval results by fine-tuning using noisy retrieval context—improve the model's robustness to noisy retrieval, they also cause the model to generate answers even when lacking reliable knowledge, creating hallucinations and weakening the model's reliability. Therefore, a method is needed to enhance the model's ability to refuse to answer uncertain or noisy retrieval results, thereby improving its credibility in high-stakes domains. Summary of the Invention

[0003] The purpose of the embodiments of this specification is to provide a method, device, storage medium and electronic device for training a model.

[0004] The embodiments of this specification provide a method for training a model. By dividing the knowledge boundary, it is possible to accurately identify whether a query is within the knowledge scope of the model. A high-quality dataset for direct preference optimization is constructed based on the knowledge boundary division results corresponding to the query samples. By using this high-quality dataset to train the model based on the direct preference optimization method, the model can learn to generate accurate answers within the knowledge boundary. At the same time, the model can be given the ability to refuse to answer "I don't know" when the query exceeds the knowledge boundary. This can effectively balance the accuracy and refusal to answer of the model, enhance the robustness and credibility of the model in high-risk areas, and effectively solve the hallucination problem existing in existing retrieval enhancement fine-tuning methods. The method includes: Obtaining a first output of the target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source; Determining a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference dataset includes the query sample, a preferred response and a non-preferred response corresponding to the query sample, and one of the preferred response and the non-preferred response includes a rejection response information; The target model is trained based on the preference data set by a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

[0005] Furthermore, determining a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample includes: A knowledge boundary division result corresponding to the query sample is determined according to whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information.

[0006] Furthermore, determining the knowledge boundary division result corresponding to the query sample according to whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, includes: According to whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, assigning the query sample to one of four knowledge quadrants, wherein each knowledge quadrant corresponds to a different knowledge boundary division result; The step of constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result includes: A preference data set is constructed according to the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs.

[0007] Furthermore, constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs includes: Determining answer indication information corresponding to the third output according to whether the third output of the target model for the query sample is consistent with the label information, wherein the answer indication information is used to indicate whether the third output is a correct answer corresponding to the query sample; Based on the answer indication information and the preference data construction rule corresponding to the knowledge quadrant to which the query sample belongs, a preference data set is constructed according to the third output and the refusal response information.

[0008] Furthermore, the method further comprises: Obtaining performance data of the target model in each of the four knowledge quadrants with respect to one or more preset indicators; If the performance data of the target model in at least one knowledge quadrant does not meet a preset condition, the query samples assigned to the at least one knowledge quadrant are cleaned.

[0009] Furthermore, the loss function used in the training process of the target model includes direct preference optimization loss and at least one of supervised fine-tuning loss and knowledge boundary classification loss.

[0010] Furthermore, the method further comprises: If the similarity between the second output and the label information is less than a first preset threshold, obtaining a fourth output obtained by performing semantic matching retrieval on the query sample by the external knowledge source: If the similarity between the fourth output and the label information is greater than or equal to the first preset threshold, and the retrieval confidence corresponding to the fourth output is greater than or equal to the second preset threshold, it is determined that the second output is consistent with the label information; otherwise, it is determined that the second output is inconsistent with the label information.

[0011] Furthermore, the method further comprises: If the similarity between the first output and the label information is less than a third preset threshold, obtaining a fifth output of the target model for at least one similar query that meets a preset similarity condition with the query sample without relying on an external knowledge source; If the similarity between the fifth output and the label information is greater than or equal to the third preset threshold, and the model confidence corresponding to the fifth output is greater than or equal to the fourth preset threshold, it is determined that the first output is consistent with the label information; otherwise, it is determined that the first output is inconsistent with the label information.

[0012] Furthermore, the method further comprises: If the similarity between the first output and the label information is greater than or equal to a fifth preset threshold, obtaining reasoning step information of the target model with respect to the first output; According to the logical rationality of the reasoning step information for the query sample, it is determined whether the first output is consistent with the label information corresponding to the query sample.

[0013] Furthermore, the method further comprises: Input the target query into the trained target model, and obtain the answer response information corresponding to the target query output by the trained target model, wherein if the target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary, the answer response information includes a rejection response information. The embodiments of this specification also provide a device for training a model, including: An obtaining module, configured to obtain a first output of a target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source; a knowledge boundary division module, configured to determine a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is within the parameter knowledge boundary of the target model and whether it is within the retrieval knowledge boundary of the external knowledge source; a data set construction module, configured to construct a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference data set includes the query sample, a preferred response and a non-preferred response corresponding to the query sample, and one of the preferred response and the non-preferred response includes a rejection response information; A model training module is used to train the target model based on the preference data set through a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

[0014] An embodiment of this specification further provides a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.

[0015] An embodiment of this specification further provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.

[0016] The embodiments of this specification also provide a computer program product having at least one instruction stored thereon, wherein the at least one instruction implements the steps of the above method when executed by a processor.

[0017] According to the solution of the embodiments of this specification, by dividing the knowledge boundary, it is possible to accurately identify whether the query is within the knowledge scope of the model, and a high-quality dataset for direct preference optimization is constructed based on the knowledge boundary division results corresponding to the query samples. By using this high-quality dataset to train the model based on the direct preference optimization method, the model can learn to generate accurate answers within the knowledge boundary, and at the same time, it can give the model the ability to refuse to answer "I don't know" when the query exceeds the knowledge boundary. This can effectively balance the accuracy and refusal to answer of the model, enhance the robustness and credibility of the model in high-risk areas, and effectively solve the hallucination problem existing in the existing retrieval enhancement fine-tuning method. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of a method for training a model provided in an embodiment of this specification; Figure 2 A schematic diagram of the structure of a device for training a model provided in an embodiment of this specification; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0020] See Figure 1 , is a flow chart of a method for training a model provided in an embodiment of this specification. In an embodiment of this specification, the method for training a model is applied to the device for training a model described in an embodiment of this specification (hereinafter referred to as "model training device") or an electronic device equipped with a model training device. Figure 1 The process shown is described in detail, and the method for training the model may specifically include the following steps: S102: Obtain a first output of the target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source.

[0021] In some embodiments, the target model is a large language model. A large language model is a language model with a large number of parameters that is pre-trained on a large-scale corpus dataset and trained through instruction fine-tuning and reinforcement learning with human feedback. It possesses powerful text generation capabilities. In some embodiments, the target model includes a RAG (Retrieval-Augmented Generation) system. This system enhances the model's contextual information by integrating external knowledge sources (such as document repositories or databases; this example embodiment does not specify the specific type of external knowledge source) to generate more accurate and contextualized text responses.

[0022] In some embodiments, a query sample refers to training data containing query information used to train a target model. Query information refers to instructions or questions input by a user to the target model, such as "What medicine should I take for a cold?" The target model then processes the query information and generates a corresponding question-and-answer response. In some embodiments, the query information may include only textual content, or it may include, but is not limited to, images, audio, video, icons, and other forms of content in addition to textual content. This exemplary embodiment does not specifically limit this.

[0023] In some embodiments, when a query sample is input into the target model without relying on an external knowledge source, the target model does not perform an external search on the query sample input into the target model using the external knowledge source. Instead, the target model directly generates and outputs a corresponding question-and-answer response (i.e., the first output). In some embodiments, when a query sample is not input into the target model, an external knowledge source is used to perform a search operation on the query sample, resulting in a corresponding search result (i.e., the second output).

[0024] S104. Determine a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is located within the parameter knowledge boundary of the target model and whether it is located within the retrieval knowledge boundary of the external knowledge source.

[0025] In some embodiments, the parameter knowledge boundary of the model refers to the situation where, if a query is within the knowledge boundary, the model can output the corresponding correct answer to the query when the model does not rely on an external knowledge source. That is, the parameter knowledge boundary of the model is used to characterize whether the query can be correctly answered by the parameter knowledge of the model. The retrieval knowledge boundary of the external knowledge source refers to the situation where, if a query is within the knowledge boundary, the external knowledge source can retrieve the correct answer corresponding to the query. That is, the retrieval knowledge boundary is used to characterize whether the query can be correctly answered by the retrieval knowledge of the external knowledge source.

[0026] In some embodiments, the query sample is labeled training data, and the query sample has corresponding label information, which refers to the question-answer response label corresponding to the query information (instruction or question) contained in the query sample. In some embodiments, whether the query sample can be correctly answered by the parameter knowledge of the target model can be determined based on the first output and the label information corresponding to the query sample. Whether the query sample can be correctly answered by the retrieval knowledge of the external knowledge source can be determined based on the second output and the label information corresponding to the query sample. Based on whether the query sample can be correctly answered by the parameter knowledge of the target model and the retrieval knowledge of the external knowledge source, the knowledge boundary division result corresponding to the query sample is determined, wherein the knowledge boundary division result includes four results, specifically, the query sample is within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; the query sample is within the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source; the query sample is outside the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; and the query sample is outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source. In some embodiments, if the query sample can be correctly answered by the parameter knowledge of the target model and can also be correctly answered by the retrieval knowledge of the external knowledge source, the corresponding knowledge boundary division result is that the query sample is located within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source. If the query sample can be correctly answered by the parameter knowledge of the target model and cannot be correctly answered by the retrieval knowledge of the external knowledge source, the corresponding knowledge boundary division result is that the query sample is located within the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source. If the query sample cannot be correctly answered by the parameter knowledge of the target model but can be correctly answered by the retrieval knowledge of the external knowledge source, the corresponding knowledge boundary division result is that the query sample is located outside the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source. If the query sample cannot be correctly answered by the parameter knowledge of the target model and cannot be correctly answered by the retrieval knowledge of the external knowledge source, the corresponding knowledge boundary division result is that the query sample is located outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source.

[0027] S106. Construct a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference data set includes the query sample, the preference response and the non-preference response corresponding to the query sample, and one of the preference response and the non-preference response includes a rejection response information.

[0028] In some embodiments, a query sample is input into the target model. At this time, the target model will use an external knowledge source to perform an external search on the query sample input into the target model to obtain the corresponding retrieval results. Then, the target model will generate and output the corresponding question-and-answer response (i.e., the third output) based on the query sample and the retrieval results.

[0029] In some embodiments, a customized preference data set is constructed based on whether the third output is consistent with the label information corresponding to the query sample and the knowledge boundary division result corresponding to the query sample. The preference data set includes the query sample, the preference response and the non-preference response corresponding to the query sample. Among them, whether the third output is consistent with the label information can be determined based on the degree of similarity between the third output and the label information. For example, if the similarity between the third output and the label information is greater than or equal to a preset threshold, it can be determined that the third output is consistent with the label information. If the similarity between the third output and the label information is less than the preset threshold, it can be determined that the third output is inconsistent with the label information.

[0030] In some embodiments, if the query sample is within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source, the rejection response information (the rejection response information is used to characterize the target model's refusal to answer the query sample, for example, the rejection response information can be a "reject" text message, and this example embodiment does not specifically limit the specific content of the rejection response information) is used as a non-preferred response. If the third output is consistent with the label information, the third output is used as a preferred response. If not, the preferred response is defaulted. If the query sample is within the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source, the rejection response information is used as a non-preferred response. If the third output is consistent with the label information, the third output is used as a preferred response. If not, the preferred response is defaulted. If the query sample is outside the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source, the rejection response information is regarded as a non-preferred response. If the third output is consistent with the label information, the third output is regarded as a preferred response. If they are inconsistent, the third output is also regarded as a non-preferred response. In this case, the preferred response is defaulted. If the query sample is outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source, the rejection response information is regarded as a non-preferred response. If the third output is consistent with the label information, the third output is regarded as a preferred response. If they are inconsistent, the third output is also regarded as a non-preferred response. In this case, the preferred response is defaulted. If the query sample is outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source, the rejection response information is regarded as a preferred response. In this case, regardless of whether the third output is consistent with the label information, the third output is regarded as a non-preferred response.

[0031] S108, training the target model based on the preference data set through a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

[0032] In some embodiments, Direct Preference Optimization (DPO) is a language model training method that replaces traditional reinforcement learning alignment. It optimizes model output by directly using human preference data without explicitly building a reward model. Its core idea is to transform preference learning into an optimization problem of probability distribution. By maximizing the log-likelihood difference between "preferred response" and "non-preferred response" in the preference data, the model parameters are directly adjusted to achieve preference alignment.

[0033] In some embodiments, the constructed preference data set is used to train the target model through a direct preference optimization method to obtain a trained target model, so that the target model can better generate accurate answers within the knowledge boundary by optimizing the output preferences of the target model, while choosing to refuse to answer outside the knowledge boundary. In some embodiments, the loss function used in the training process of the target model includes DPO loss (direct preference optimization loss), which is used to learn the difference between preference pairs so that the model can distinguish between preference and non-preference outputs. In some embodiments, the target query to be answered is input into the trained target model. If the target query exceeds the parameter knowledge boundary of the target model and the retrieval knowledge boundary of the external knowledge source, the target model will output a rejection response message. Otherwise (the target query is within the parameter knowledge boundary and within the retrieval knowledge boundary, or the target query is within the parameter knowledge boundary or within the retrieval knowledge boundary), the target model will generate and output the corresponding correct answer.

[0034] According to the solution of the embodiments of this specification, by dividing the knowledge boundary, it is possible to accurately identify whether the query is within the knowledge scope of the model, and a high-quality dataset for direct preference optimization is constructed based on the knowledge boundary division results corresponding to the query sample. By using this high-quality dataset to train the model based on the direct preference optimization method, the model can learn to generate accurate answers within the knowledge boundary, and at the same time, it can give the model the ability to refuse to answer "I don't know" when the query exceeds the knowledge boundary. This can effectively balance the accuracy and refusal to answer of the model, enhance the robustness and credibility of the model in high-risk fields (such as medical, legal, etc.), and can effectively solve the hallucination problem existing in existing retrieval enhancement fine-tuning methods.

[0035] In some embodiments, determining the knowledge boundary division result corresponding to the query sample based on the first output, the second output and the label information corresponding to the query sample includes: determining the knowledge boundary division result corresponding to the query sample based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information. In some embodiments, the knowledge boundary division result corresponding to the query sample can be determined based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information corresponding to the query sample, wherein if the first output is consistent with the label information and the second output is consistent with the label information, the corresponding knowledge boundary division result is that the query sample is located within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; if the first output is consistent with the label information and the second output is inconsistent with the label information, the corresponding knowledge boundary division result is that the query sample is located within the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source; if the first output is inconsistent with the label information and the second output is consistent with the label information, the corresponding knowledge boundary division result is that the query sample is located outside the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; if the first output is inconsistent with the label information and the second output is inconsistent with the label information, the corresponding knowledge boundary division result is that the query sample is located outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source. In some embodiments, whether the first output or the second output is consistent with the label information can be determined based on the degree of similarity between the first output or the second output and the label information. For example, if the similarity between the first output or the second output and the label information is greater than or equal to a preset threshold, it can be determined that the first output or the second output is consistent with the label information. If the similarity between the first output or the second output and the label information is less than the preset threshold, it can be determined that the first output or the second output is inconsistent with the label information.

[0036] In some embodiments, determining the knowledge boundary division result corresponding to the query sample based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, includes: allocating the query sample to one of the four knowledge quadrants based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, wherein each knowledge quadrant corresponds to a different knowledge boundary division result; wherein, constructing a preference data set based on the third output of the target model for the query sample, the label information and the knowledge boundary division result, includes: constructing a preference data set based on the third output of the target model for the query sample, the label information and the knowledge quadrant to which the query sample belongs. In some embodiments, the query samples are divided into four preset knowledge quadrants, and each knowledge quadrant corresponds to a different knowledge boundary division result. For example, the query samples in the first knowledge quadrant are located within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source, the query samples in the second knowledge quadrant are located within the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source, the query samples in the third knowledge quadrant are located outside the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source, and the query samples in the fourth knowledge quadrant are located outside the parameter knowledge boundary of the target model and outside the retrieval knowledge boundary of the external knowledge source. It should be noted that the correspondence between the above-mentioned knowledge quadrants and the knowledge boundary division results is only for example and not for limitation. Those skilled in the art should understand that any correspondence can be included in the scope of protection of this specification, and this example embodiment does not make any special limitations on this. In some embodiments, based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, the query sample is assigned to one of the four preset knowledge quadrants based on a preset allocation rule, wherein if the first output is consistent with the label information, it means that the query sample can be correctly answered by the parameter knowledge of the target model; if the first output is inconsistent with the label information, it means that the query sample cannot be correctly answered by the parameter knowledge of the target model; if the second output is consistent with the label information, it means that the query sample can be correctly answered by the retrieval knowledge of the external knowledge source; if the second output is inconsistent with the label information, it means that the query sample cannot be correctly answered by the retrieval knowledge of the external knowledge source. That is, by analyzing whether the query sample can be correctly answered by the parameter knowledge of the target model or the retrieval knowledge of the external knowledge source, the query sample is assigned to the corresponding knowledge quadrant. By dividing the query sample into four knowledge quadrants and constructing customized preference data for each knowledge quadrant, a high-quality data set for direct preference optimization (DPO) training is formed, which can effectively improve the target model's ability to refuse to answer outside the knowledge boundary.In some embodiments, based on a preset allocation rule, if the first output is consistent with the label information and the second output is consistent with the label information, the query sample is allocated to the first knowledge quadrant; if the first output is consistent with the label information and the second output is inconsistent with the label information, the query sample is allocated to the second knowledge quadrant; if the first output is inconsistent with the label information and the second output is consistent with the label information, the query sample is allocated to the third knowledge quadrant; if the first output is inconsistent with the label information and the second output is inconsistent with the label information, the query sample is allocated to the fourth knowledge quadrant. It should be noted that the above allocation rules are only examples and not limitations. Those skilled in the art should understand that any allocation rule can be included in the scope of protection of this specification, and this example embodiment does not specifically limit this. In some embodiments, the query sample is input into the target model. At this time, the target model will use an external knowledge source to perform an external search on the query sample input into the target model to obtain the corresponding retrieval result, and then the target model will generate and output the corresponding question and answer response (i.e., the third output) based on the query sample and the retrieval result. In some embodiments, a customized preference data set is constructed based on whether the third output is consistent with the label information and the knowledge quadrant to which the query sample belongs. The preference data set includes the query sample, the preferred response and the non-preferred response corresponding to the query sample. If the query sample belongs to the first knowledge quadrant, the rejection response information is used as the non-preferred response. If the third output is consistent with the label information, the third output is used as the preferred response. If not, the preferred response is defaulted. If the query sample belongs to the second knowledge quadrant, the rejection response information is used as the non-preferred response. If the third output is consistent with the label information, the third output is used as the preferred response. If not, the third output is also used as the non-preferred response. In this case, the preferred response is defaulted. If the query sample belongs to the third knowledge quadrant, the rejection response information is used as the non-preferred response. If the third output is consistent with the label information, the third output is used as the preferred response. If not, the third output is also used as the non-preferred response. In this case, the preferred response is defaulted. If the query sample belongs to the fourth knowledge quadrant, the rejection response information is used as the preferred response. In this case, regardless of whether the third output is consistent with the label information, the third output is used as the non-preferred response.

[0037] In some embodiments, constructing a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs includes: determining answer indication information corresponding to the third output based on whether the third output of the target model for the query sample is consistent with the label information, wherein the answer indication information is used to indicate whether the third output is the correct answer corresponding to the query sample; constructing rules based on the answer indication information and the preference data corresponding to the knowledge quadrant to which the query sample belongs, and constructing a preference data set based on the third output and the refusal response information. In some embodiments, based on whether the third output is consistent with the label information corresponding to the query sample, the answer indication information corresponding to the third output is determined, and the answer indication information is used to indicate whether the third output is the correct answer or the incorrect answer corresponding to the query sample. Then, based on the answer indication information and the preference data construction rule corresponding to the knowledge quadrant to which the query sample belongs, a preference data set is constructed according to the third output and the refusal response information. For example, the preference data construction rule corresponding to the first knowledge quadrant is "select the correct answer as the preferred response and take 'refusal to answer' as the non-preferred response", the preference data construction rule corresponding to the second indication quadrant is "select the correct answer as the preferred response and take the incorrect answer and 'refusal to answer' as the non-preferred responses", the preference data construction rule corresponding to the third knowledge quadrant is "select the correct answer as the preferred response and take the incorrect answer and 'refusal to answer' as the non-preferred responses", and the preference data construction rule corresponding to the fourth knowledge quadrant is "select 'refusal to answer' as the preferred response and take the incorrect answer and the correct answer as the non-preferred responses". It should be noted that the correspondence between the above-mentioned knowledge quadrants and the preference data construction rules is only for example and not for limitation. Those skilled in the art should understand that any correspondence can be included in the scope of protection of this specification, and this example embodiment does not make any special limitations on this. In some embodiments, if the query sample belongs to the first knowledge quadrant, the refusal response information is used as a non-preferred response. If the third output is the correct answer, the third output is used as a preferred response. If the third output is an incorrect answer, the preferred response is defaulted. If the query sample belongs to the second knowledge quadrant, the refusal response information is used as a non-preferred response. If the third output is the correct answer, the third output is used as a preferred response. If the third output is an incorrect answer, the third output is also used as a non-preferred response. In this case, the preferred response is defaulted. If the query sample belongs to the third knowledge quadrant, the refusal response information is used as a non-preferred response. If the third output is the correct answer, the third output is used as a preferred response. If the third output is an incorrect answer, the third output is also used as a non-preferred response. In this case, the preferred response is defaulted. If the query sample belongs to the fourth knowledge quadrant, the refusal response information is used as a preferred response. In this case, regardless of whether the third output is a correct answer or an incorrect answer, the third output is used as a non-preferred response.

[0038] In some embodiments, the method further includes: obtaining performance data of the target model in each of the four knowledge quadrants with respect to one or more preset indicators; if the performance data of the target model in at least one knowledge quadrant does not meet the preset conditions, cleaning the query samples assigned to the at least one knowledge quadrant. In some embodiments, performance data of the target model in each of the four preset knowledge quadrants with respect to one or more preset indicators are obtained, and the preset indicators include but are not limited to accuracy, rejection rate, rejection accuracy, etc., which are not specifically limited in this example embodiment. If the performance data of the target model in at least one knowledge quadrant does not meet the preset conditions, cleaning the query samples assigned to the at least one knowledge quadrant to ensure that the target model generates accurate answers within the knowledge boundary and correctly rejects answers outside the knowledge boundary, wherein the preset conditions include but are not limited to accuracy less than or equal to a preset threshold, rejection accuracy lower than or equal to a preset threshold, etc., which are not specifically limited in this example embodiment. In some embodiments, if the query samples corresponding to the target knowledge quadrant are cleaned, a new batch of query samples needs to be obtained, and each query sample in this new batch of query samples is assigned to the knowledge quadrant to which it belongs, until the number of query samples assigned to the target knowledge quadrant meets the preset conditions (for example, reaches a preset number threshold).

[0039] In some embodiments, the loss function used in the training process of the target model includes direct preference optimization loss and at least one of supervised fine-tuning loss and knowledge boundary classification loss. In some embodiments, in addition to direct preference optimization (DPO) loss, the loss function used in the training process of the target model also needs to be combined with supervised fine-tuning (SFT) loss and knowledge boundary classification loss to further comprehensively improve the performance of the target model through a multi-objective training method, so that it can generate accurate answers within the knowledge boundary and correctly reject answers outside the knowledge boundary. Among them, the SFT loss is used to improve the target model's ability to generate preference outputs, and the knowledge boundary classification loss is used to enhance the target model's perception of the knowledge boundary by predicting the knowledge boundary division results corresponding to the query.

[0040] In some embodiments, the method further includes: if the similarity between the second output and the label information is less than a first preset threshold, obtaining a fourth output obtained by performing a semantic matching retrieval on the query sample from the external knowledge source; if the similarity between the fourth output and the label information is greater than or equal to the first preset threshold, and the retrieval confidence corresponding to the fourth output is greater than or equal to a second preset threshold, determining that the second output is consistent with the label information; otherwise, determining that the second output is inconsistent with the label information. In some embodiments, if the similarity between the second output and the label information is less than the first preset threshold, the second output is not directly determined to be inconsistent with the label information. Instead, a semantic matching retrieval operation is performed on the query sample using the external knowledge source without inputting a target model to obtain a corresponding retrieval result (i.e., the fourth output). Semantic matching retrieval is a retrieval technology based on semantic understanding. It aims to find the most semantically relevant results to the query content from large-scale data, rather than relying solely on mechanical keyword matching. It improves retrieval accuracy and relevance by deeply understanding the semantic information of the text (such as intent, context, entity relationships, etc.). In some embodiments, if the similarity between the fourth output and the label information is still less than the first preset threshold, it is determined that the second output is inconsistent with the label information. If the similarity between the fourth output and the label information is greater than or equal to the first preset threshold, it is necessary to further determine whether the retrieval confidence corresponding to the fourth output is greater than or equal to the second preset threshold. If so, it can be determined that the second output is consistent with the label information. Otherwise, it is determined that the second output is inconsistent with the label information. Retrieval confidence is a quantitative evaluation indicator that measures the relevance or reliability of the retrieval results and the query content, and is usually presented in the form of a probability value, score or grade.

[0041] In some embodiments, the method further includes: if the similarity between the first output and the label information is less than a third preset threshold, obtaining a fifth output of the target model for at least one similar query that meets a preset similarity condition with the query sample without relying on an external knowledge source; if the similarity between the fifth output and the label information is greater than or equal to the third preset threshold, and the model confidence corresponding to the fifth output is greater than or equal to a fourth preset threshold, determining that the first output is consistent with the label information; otherwise, determining that the first output is inconsistent with the label information. In some embodiments, if the similarity between the first output and the label information is less than a third preset threshold, the first output is not directly determined to be inconsistent with the label information. Instead, at least one similar query that satisfies a preset similarity condition with the query sample (for example, a corresponding similarity greater than or equal to a preset threshold) is first obtained. Then, without relying on an external knowledge source, the at least one similar query is input into the target model to obtain at least one question-and-answer response (i.e., a fifth output) generated and output by the target model for the at least one similar query. If the at least one question-and-answer response does not contain a question-and-answer response with a similarity greater than or equal to the third preset threshold to the label information, the first output is determined to be inconsistent with the label information. If the at least one question-and-answer response contains a target question-and-answer response with a similarity greater than or equal to the third preset threshold to the label information, it is necessary to further determine whether the model confidence corresponding to the target question-and-answer response is greater than or equal to a fourth preset threshold. If so, the first output is determined to be consistent with the label information. Otherwise, the first output is determined to be inconsistent with the label information. Model confidence refers to the degree of certainty of the target model regarding its prediction results, and is typically presented in the form of a probability value, a score, or a confidence interval.

[0042] In some embodiments, the method further includes: if the similarity between the first output and the label information is greater than or equal to a fifth preset threshold, obtaining the reasoning step information of the target model regarding the first output; and determining whether the label information corresponding to the first output and the query sample are consistent based on the logical rationality of the reasoning step information for the query sample. In some embodiments, if the similarity between the first output and the label information is greater than or equal to a fifth preset threshold, it is not directly determined that the first output is consistent with the label information. Instead, it is necessary to first obtain the target model's reasoning step information about the first output. The reasoning step information includes a series of intermediate step-by-step reasoning steps in the process of the target model generating the first output based on the query sample, and then obtain the logical rationality of the reasoning step information for the query sample. If the logical rationality is greater than or equal to the preset threshold, it is determined that the first output is consistent with the label information. If the logical rationality is less than the preset threshold, it is determined that the first output is inconsistent with the label information. The logical rationality of the reasoning step information for the query sample can be obtained by manually (for example, a trainer) annotating the reasoning step information, or the reasoning step information can be input into a trained logical judgment model to obtain the logical rationality of the reasoning step information output by the logical judgment model for the query sample. It should be noted that the above-mentioned method of obtaining the logical rationality is only an example and not a limitation. Those skilled in the art should understand that any obtaining method can be included in the scope of protection of this specification, and this example embodiment does not specifically limit this.

[0043] In some embodiments, the method further includes: inputting a target query into the trained target model, and obtaining answer response information corresponding to the target query output by the trained target model, wherein if the target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary, the answer response information includes a rejection response information. In some embodiments, the target query to be answered is input into the trained target model. If the target query exceeds the parameter knowledge boundary of the target model and the retrieval knowledge boundary of the external knowledge source, the target model will output a rejection response information; otherwise (the target query is within the parameter knowledge boundary and within the retrieval knowledge boundary, or the target query is within the parameter knowledge boundary or within the retrieval knowledge boundary), the target model will generate and output the corresponding correct answer.

[0044] Figure 2 This is a schematic diagram of a model training device provided in an embodiment of this specification. This model training device (hereinafter referred to as "model training device 1") can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the model training device 1 includes an acquisition module 11, a knowledge boundary demarcation module 12, a dataset construction module 13, and a model training module 14.

[0045] An obtaining module 11 is configured to obtain a first output of a target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source; a knowledge boundary division module 12, configured to determine a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; a data set construction module 13, configured to construct a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference data set includes the query sample, the preferred response and the non-preferred response corresponding to the query sample, and one of the preferred response and the non-preferred response includes a rejection response; The model training module 14 is used to train the target model based on the preference data set through a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

[0046] In some embodiments, determining the knowledge boundary division result corresponding to the query sample based on the first output, the second output and the label information corresponding to the query sample includes: determining the knowledge boundary division result corresponding to the query sample based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information.

[0047] In some embodiments, determining the knowledge boundary division result corresponding to the query sample based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, includes: allocating the query sample to one of the four knowledge quadrants based on whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, wherein each knowledge quadrant corresponds to a different knowledge boundary division result; wherein, constructing a preference data set based on the third output of the target model for the query sample, the label information and the knowledge boundary division result, includes: constructing a preference data set based on the third output of the target model for the query sample, the label information and the knowledge quadrant to which the query sample belongs.

[0048] In some embodiments, constructing a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs includes: determining answer indication information corresponding to the third output based on whether the third output of the target model for the query sample is consistent with the label information, wherein the answer indication information is used to indicate whether the third output is the correct answer corresponding to the query sample; constructing rules based on the answer indication information and the preference data corresponding to the knowledge quadrant to which the query sample belongs, and constructing a preference data set based on the third output and the refusal response information.

[0049] In some embodiments, the model training device 1 is also used to: obtain performance data of the target model in each of the four knowledge quadrants regarding one or more preset indicators; if the performance data of the target model in at least one knowledge quadrant does not meet the preset conditions, clean the query samples assigned to the at least one knowledge quadrant.

[0050] In some embodiments, the loss function used in the training process of the target model includes direct preference optimization loss and at least one of supervised fine-tuning loss and knowledge boundary classification loss.

[0051] In some embodiments, the model training device 1 is also used to: if the similarity between the second output and the label information is less than a first preset threshold, obtain the fourth output obtained by the external knowledge source performing a semantic matching retrieval on the query sample; if the similarity between the fourth output and the label information is greater than or equal to the first preset threshold, and the retrieval confidence corresponding to the fourth output is greater than or equal to the second preset threshold, determine that the second output is consistent with the label information; otherwise, determine that the second output is inconsistent with the label information.

[0052] In some embodiments, the model training device 1 is also used to: if the similarity between the first output and the label information is less than a third preset threshold, obtain the fifth output of the target model for at least one similar query that meets the preset similarity condition with the query sample without relying on an external knowledge source; if the similarity between the fifth output and the label information is greater than or equal to the third preset threshold, and the model confidence corresponding to the fifth output is greater than or equal to a fourth preset threshold, determine that the first output is consistent with the label information; otherwise, determine that the first output is inconsistent with the label information.

[0053] In some embodiments, the model training device 1 is also used to: if the similarity between the first output and the label information is greater than or equal to a fifth preset threshold, obtain the reasoning step information of the target model regarding the first output; and determine whether the label information corresponding to the first output and the query sample is consistent based on the logical rationality of the reasoning step information for the query sample.

[0054] In some embodiments, the model training device 1 is also used to: input the target query into the trained target model, and obtain the answer response information corresponding to the target query output by the trained target model, wherein, if the target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary, the answer response information includes rejection response information.

[0055] The above device embodiments correspond to the aforementioned method embodiments. For detailed descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For detailed descriptions, please refer to the corresponding method embodiments.

[0056] The embodiments of this specification also provide a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the method of the embodiments of this specification.

[0057] An embodiment of the present specification further provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor to execute the method of the embodiment of the present specification.

[0058] The embodiments of this specification also provide Figure 3 The structural diagram of the electronic device shown in FIG. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the above method.

[0059] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0060] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0061] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0062] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0064] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0065] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0066] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0067] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible within the scope of the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for training a model, comprising: Obtaining a first output of the target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source; Determining a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is within the parameter knowledge boundary of the target model and within the retrieval knowledge boundary of the external knowledge source; constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference dataset includes the query sample, a preferred response and a non-preferred response corresponding to the query sample, and one of the preferred response and the non-preferred response includes a rejection response information; The target model is trained based on the preference data set by a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

2. The method according to claim 1, wherein determining a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample comprises: A knowledge boundary division result corresponding to the query sample is determined according to whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information.

3. The method according to claim 2, wherein determining a knowledge boundary division result corresponding to the query sample based on whether the first output is consistent with the label information corresponding to the query sample and whether the second output is consistent with the label information comprises: According to whether the first output is consistent with the label information corresponding to the query sample, and whether the second output is consistent with the label information, assigning the query sample to one of four knowledge quadrants, wherein each knowledge quadrant corresponds to a different knowledge boundary division result; The step of constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result includes: A preference data set is constructed according to the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs.

4. The method according to claim 3, wherein constructing a preference dataset based on the third output of the target model for the query sample, the label information, and the knowledge quadrant to which the query sample belongs comprises: Determining answer indication information corresponding to the third output according to whether the third output of the target model for the query sample is consistent with the label information, wherein the answer indication information is used to indicate whether the third output is a correct answer corresponding to the query sample; Based on the answer indication information and the preference data construction rule corresponding to the knowledge quadrant to which the query sample belongs, a preference data set is constructed according to the third output and the refusal response information.

5. The method according to claim 4, further comprising: Obtaining performance data of the target model in each of the four knowledge quadrants with respect to one or more preset indicators; If the performance data of the target model in at least one knowledge quadrant does not meet a preset condition, the query samples assigned to the at least one knowledge quadrant are cleaned.

6. The method according to claim 1, wherein the loss function used in the training process of the target model includes a direct preference optimization loss and at least one of a supervised fine-tuning loss and a knowledge boundary classification loss.

7. The method according to claim 2, further comprising: If the similarity between the second output and the label information is less than a first preset threshold, obtaining a fourth output obtained by performing semantic matching retrieval on the query sample by the external knowledge source: If the similarity between the fourth output and the label information is greater than or equal to the first preset threshold, and the retrieval confidence corresponding to the fourth output is greater than or equal to the second preset threshold, it is determined that the second output is consistent with the label information; otherwise, it is determined that the second output is inconsistent with the label information.

8. The method according to claim 2 or 7, further comprising: If the similarity between the first output and the label information is less than a third preset threshold, obtaining a fifth output of the target model for at least one similar query that meets a preset similarity condition with the query sample without relying on an external knowledge source; If the similarity between the fifth output and the label information is greater than or equal to the third preset threshold, and the model confidence corresponding to the fifth output is greater than or equal to the fourth preset threshold, it is determined that the first output is consistent with the label information; otherwise, it is determined that the first output is inconsistent with the label information.

9. The method according to claim 2, further comprising: If the similarity between the first output and the label information is greater than or equal to a fifth preset threshold, obtaining reasoning step information of the target model with respect to the first output; According to the logical rationality of the reasoning step information for the query sample, it is determined whether the first output is consistent with the label information corresponding to the query sample.

10. The method according to claim 1, further comprising: The target query is input into the trained target model to obtain answer response information corresponding to the target query output by the trained target model, wherein if the target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary, the answer response information includes rejection response information.

11. A device for training a model, comprising: An obtaining module, configured to obtain a first output of a target model for a query sample without relying on an external knowledge source, and a second output obtained by retrieving the query sample from the external knowledge source; a knowledge boundary division module, configured to determine a knowledge boundary division result corresponding to the query sample based on the first output, the second output, and the label information corresponding to the query sample, wherein the knowledge boundary division result is used to indicate whether the query sample is within the parameter knowledge boundary of the target model and whether it is within the retrieval knowledge boundary of the external knowledge source; a data set construction module, configured to construct a preference data set based on the third output of the target model for the query sample, the label information, and the knowledge boundary division result, wherein the preference data set includes the query sample, a preferred response and a non-preferred response corresponding to the query sample, and one of the preferred response and the non-preferred response includes a rejection response information; A model training module is used to train the target model based on the preference data set through a direct preference optimization method to obtain a trained target model, so that the trained target model can output the rejection response information when the input target query exceeds the parameter knowledge boundary and the retrieval knowledge boundary.

12. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 10.

14. A computer program product having at least one instruction stored thereon, characterized in that: When the at least one instruction is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Method and system for enhancing knowledge boundary perception ability of large language model

    CN118013049A

  • Method, device and equipment for training large language model

    CN118153624A

  • Large language model training method and device, equipment and storage medium

    CN119066155A

  • Large model processing method for task-oriented dialogue, electronic equipment and storage medium

    CN120067263A

  • Knowledge boundary determination method and device, storage medium and electronic equipment

    CN120105243A

Cited By

  • Model training method and device, copywriting generation method and device, equipment and medium

    CN121659026A