Model optimization method of large language model and related equipment

By constructing training samples and perturbing the order of candidate information, the large language model is optimized, which solves the instability problem in retrieval enhancement generation technology and improves the accuracy and stability of the model.

CN120973887APending Publication Date: 2025-11-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410610241.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing large language models exhibit instability in retrieval enhancement generation techniques. For example, they cannot be correctly cited when the correct candidate is located at the end of the list, the conclusion changes after the candidate order is changed, and semantically similar candidates are prone to being selected incorrectly.

Method used

By acquiring training text and candidate information sets, training samples are constructed, the order of candidate information is perturbed, predicted text response results are generated, result consistency evaluation is performed, preference data pairs are constructed, and the large language model is optimized.

Benefits of technology

It enhances the retrieval and generation capabilities of large language models, reduces model illusions, and improves the accuracy and stability of model responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973887A_ABST
    Figure CN120973887A_ABST
Patent Text Reader

Abstract

The invention discloses a model optimization method for a large language model and related equipment. According to the embodiment of the invention, a plurality of training texts and a candidate information set obtained through retrieval can be obtained; constructing a training sample set based on the plurality of training texts and the candidate information set; disturbing the arrangement sequence of multiple pieces of candidate information in the candidate information set to obtain a disturbed candidate information set; calling the to-be-optimized large language model to perform text reply on the training text based on the perturbed candidate information set, and generating a predicted text reply result; performing result consistency evaluation on the predicted text reply result and the training text reply result, and constructing a preference data pair corresponding to the training text based on an evaluation result; and based on the preference data pair corresponding to the training text, performing model optimization on the to-be-optimized large language model to obtain an optimized large language model. According to the method, the retrieval enhancement generation capability of the large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method for optimizing large language models and related equipment. Background Technology

[0002] Retrieval-enhanced generation (REGG) refers to techniques that retrieve relevant information from external knowledge bases before using a large language model to answer a question. REGGG has been shown to significantly improve the accuracy of model responses, particularly reducing erroneous outputs in knowledge-intensive tasks. By citing information sources, users can verify the accuracy of the answers, thereby increasing their trust in the model's output.

[0003] However, in practical applications, the use of retrieval enhancement generation technology may encounter unstable situations. For example, when the correct candidate is located too far back in the list, it may not be cited correctly; after changing the order of the candidates, the cited candidates may change, and even the final conclusion may change; when the semantics of the candidates are similar to the question, it is easy to select the wrong candidates to organize the answer. Summary of the Invention

[0004] This application provides a method for optimizing a large language model and related equipment. The related equipment may include a model optimization device for a large language model, electronic equipment, computer-readable storage medium, and computer program products, which can improve the retrieval enhancement generation capability of the large language model.

[0005] This application provides a method for optimizing a large language model, including:

[0006] Obtain multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval database, wherein the set of candidate information includes multiple candidate information arranged in a preset order;

[0007] Based on multiple training texts and the candidate information set, a training sample set is constructed. The training sample set includes multiple training samples, each training sample including the training text and the corresponding training text response result. The training text response result includes a reference marker, which is used to characterize the reference of the training text response result to the candidate information in the candidate information set.

[0008] The order of multiple candidate information in the candidate information set is perturbed to obtain a perturbed candidate information set, which includes multiple candidate information arranged in the perturbed order.

[0009] The large language model to be optimized is invoked to perform a text response on the training text based on the perturbated candidate information set, and a predicted text response result is generated.

[0010] A consistency evaluation is performed on the predicted text response results and the training text response results, and a preference data pair corresponding to the training text is constructed based on the evaluation results. The preference data pair includes the training text response results as correct data and the predicted text response results as incorrect data.

[0011] Based on the preference data pairs corresponding to the training texts, the large language model to be optimized is optimized to obtain the optimized large language model.

[0012] Accordingly, embodiments of this application provide a model optimization apparatus for a large language model, comprising:

[0013] The acquisition unit is used to acquire multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval library, wherein the set of candidate information includes multiple candidate information arranged in a preset order.

[0014] The construction unit is used to construct a training sample set based on multiple training texts and the candidate information set. The training sample set includes multiple training samples, each training sample includes the training text and the corresponding training text response result. The training text response result includes a reference marker, which is used to characterize the reference of the training text response result to the candidate information in the candidate information set.

[0015] A perturbation unit is used to perturb the arrangement order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set, wherein the perturbed candidate information set includes multiple candidate information arranged in the perturbed order;

[0016] The prediction unit is used to call the large language model to be optimized to perform a text response on the training text based on the perturbed candidate information set, and generate a predicted text response result;

[0017] An evaluation unit is used to evaluate the consistency between the predicted text response result and the training text response result, and to construct a preference data pair corresponding to the training text based on the evaluation result. The preference data pair includes the training text response result as correct data and the predicted text response result as incorrect data.

[0018] An optimization unit is used to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, so as to obtain an optimized large language model.

[0019] Optionally, in some embodiments of this application, the prediction unit may further include a first acquisition subunit, a first invocation subunit, and a first training subunit, as follows:

[0020] The first acquisition subunit is used to acquire an initial large language model and a preset text output instruction that instructs the initial large language model to generate results that meet the current business requirements.

[0021] The first calling subunit is used to call the initial large language model to perform text response on the training text based on the candidate information set, and generate an initial predicted text response result that satisfies the preset text output instruction;

[0022] The first training subunit is used to train the initial large language model based on the difference between the initial predicted text response results and the training text response results, until a large language model to be optimized is obtained that meets the training termination condition.

[0023] Optionally, in some embodiments of this application, the first training subunit may include a second training subunit, a second calling subunit, a denoising subunit, and a fine-tuning subunit, as follows:

[0024] The second training subunit is used to train the initial large language model based on the difference between the initial predicted text response result and the training text response result, until a fine-tuned large language model to be optimized is obtained that meets the training termination condition.

[0025] The second calling subunit is used to call the large language model to be optimized before fine-tuning to perform a text response to the training text based on the candidate information set, and generate a predicted text response result before denoising.

[0026] The denoising subunit is used to denoise the predicted text response result before denoising to obtain the predicted text response result after denoising.

[0027] The fine-tuning subunit is used to fine-tune the large language model to be optimized before fine-tuning based on the denoised predicted text response results, until the large language model to be optimized is obtained.

[0028] Optionally, in some embodiments of this application, the denoising subunit may be specifically used to delete erroneous text response results that do not follow the training text or are inconsistent with the facts from the predicted text response results before denoising, to obtain the text response results after deletion; and to remove response results that are inconsistent with the conclusion of the training text response results from the text response results after deletion, to obtain the predicted text response results after denoising.

[0029] Optionally, in some embodiments of this application, the evaluation unit may include a second acquisition subunit, a third invocation subunit, and a first construction subunit, as follows:

[0030] The second acquisition subunit is used to acquire the predicted text response result corresponding to the training text, and the training text response result;

[0031] The third calling subunit is used to call the consistency evaluation model to evaluate the consistency of the predicted text response result corresponding to the training text and the training text response result, and obtain the evaluation result corresponding to the training text.

[0032] The first construction subunit is used to construct the preference data pair corresponding to the training text based on the evaluation result corresponding to the training text, the predicted text response result corresponding to the training text, and the training text response result.

[0033] Optionally, in some embodiments of this application, the third calling subunit may be specifically used to obtain a consistency evaluation model and a preset result evaluation instruction that instructs the consistency evaluation model to generate evaluation results; call the consistency evaluation model to perform result consistency evaluation on the predicted text response result corresponding to the training text and the training text response result, and obtain the evaluation result corresponding to the training text that satisfies the preset result evaluation instruction.

[0034] Optionally, in some embodiments of this application, the first construction subunit may be specifically used to filter out training texts from the training texts whose evaluation result is the evaluation failure result; and construct preference data pairs corresponding to the training texts based on the predicted text response results corresponding to the training texts and the training text response results.

[0035] Optionally, in some embodiments of this application, the preset result evaluation instructions include conclusion consistency evaluation instructions and inductive correctness evaluation instructions.

[0036] Optionally, in some embodiments of this application, the optimization unit may include a second construction subunit and a fourth invocation subunit, as follows:

[0037] The second construction subunit is used to construct the loss function corresponding to the direct preference optimization process, and to construct the direct preference optimization model based on the loss function;

[0038] The fourth calling subunit is used to call the direct preference optimization model to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, so as to obtain the optimized large language model.

[0039] Optionally, in some embodiments of this application, the second construction subunit may be specifically used to obtain a first cumulative probability difference between the correct data and the initial predicted text response result, and a second cumulative probability difference between the incorrect data and the initial predicted text response result; based on the first cumulative probability difference and the second cumulative probability difference, a loss function corresponding to the direct preference optimization process is constructed.

[0040] Optionally, in some embodiments of this application, the acquisition unit may be specifically used to perform data vectorization processing on the raw data obtained from one or more data sources to obtain vectorized raw data, and construct a retrieval library based on the vectorized raw data; acquire multiple training texts; and retrieve a set of candidate information related to the training texts from the retrieval library by comparing the vectorized representations corresponding to the training texts with the vectorized raw data.

[0041] An electronic device provided in this application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the model optimization method for a large language model provided in this application.

[0042] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps in the model optimization method for a large language model provided in this application.

[0043] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the model optimization method for a large language model provided in embodiments of this application.

[0044] This application provides a method and related equipment for optimizing a large language model. It can acquire multiple training texts and a set of candidate information related to the training texts retrieved from a search database; construct a training sample set based on the multiple training texts and the candidate information set; perturb the order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating predicted text response results; evaluate the consistency between the predicted text response results and the training text response results, and construct preference data pairs corresponding to the training texts based on the evaluation results; optimize the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain an optimized large language model. This application can reduce model illusion by adding citation markers to the text response results, thereby improving the accuracy of the model's answers; and by perturbing the position of candidate information, it can improve the stability and final effect of the model without changing the retrieval module or adding new training data. Simultaneously, it utilizes preference data to further optimize the large language model to be optimized, thereby addressing the model's shortcomings and improving the retrieval enhancement generation capabilities of the large language model. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of a scenario for the model optimization method for a large language model provided in an embodiment of this application;

[0047] Figure 2 This is a first flowchart of the model optimization method for a large language model provided in the embodiments of this application;

[0048] Figure 3 This is a second flowchart of the model optimization method for a large language model provided in the embodiments of this application;

[0049] Figure 4 This is a flowchart of the model training process provided in the embodiments of this application;

[0050] Figure 5 This is a flowchart of the self-iterative optimization process provided in the embodiments of this application;

[0051] Figure 6 This is a flowchart of the model optimization process provided in the embodiments of this application;

[0052] Figure 7This is a schematic diagram of the structure of the model optimization device for a large language model provided in an embodiment of this application;

[0053] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] This application provides a method for optimizing a large language model and related equipment. The related equipment may include a model optimization device for a large language model, an electronic device, a computer-readable storage medium, and a computer program product. Specifically, the model optimization device for the large language model may be integrated into an electronic device, which may be a terminal or a server, etc.

[0056] It is understood that the model optimization method for the large language model in this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting this application.

[0057] like Figure 1 As shown, an example is a model optimization method for a large language model jointly executed by a terminal and a server. The model optimization system for a large language model provided in this application includes a terminal and a server, etc.; the terminal and the server are connected via a network, such as a wired or wireless network, etc., wherein the model optimization device for the large language model can be integrated into the server.

[0058] The server can be used to: acquire multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval database; construct a training sample set based on the multiple training texts and the candidate information set; perturb the order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating a predicted text response result; evaluate the consistency between the predicted text response result and the training text response result, and construct preference data pairs corresponding to the training texts based on the evaluation results; and optimize the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain an optimized large language model. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In the large language model optimization method or apparatus disclosed in this application, multiple servers can form a blockchain, and the server is a node on the blockchain.

[0059] The terminal can be used to: collect multiple training texts and a set of candidate information related to the training texts retrieved from a search database, and send the training texts and the candidate information set to the server. The terminal can include mobile phones, smart voice interaction devices, smart home appliances, in-vehicle terminals, aircraft, tablets, laptops, or personal computers (PCs), etc. A client can also be set on the terminal, which can be an application client or a browser client, etc.

[0060] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0061] This embodiment will be described from the perspective of a model optimization device for a large language model. This device can be integrated into an electronic device, such as a server or terminal. This embodiment can be applied to various scenarios including cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0062] It is understood that in the specific implementation of this application, data related to user information (such as training text and training text response results corresponding to the training text) are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0063] like Figure 2 As shown, the specific process of the model optimization method for this large language model can be as follows:

[0064] S201. Obtain multiple training texts and a set of candidate information related to the training texts retrieved from the retrieval database.

[0065] Large Language Models (LLMs) are a natural language processing technique that uses deep learning algorithms to understand and generate human language. LLMs are typically based on neural network architectures and are trained on large amounts of text data to perform a wide range of tasks and capture and learn complex patterns and relationships within large volumes of text. In this embodiment, the specific structure of the LLM is not limited; it can be a structure such as GPT4.

[0066] Retrieval-enhanced generation (REGG) is a technique that uses information from private or proprietary data sources to assist in text generation. While large language models demonstrate powerful text processing capabilities in practical applications, there is still room for improvement in areas such as accuracy, knowledge update speed, and answer transparency. REGG refers to retrieving relevant information from external knowledge bases before using a large language model to answer a question. REGG has been proven to significantly improve the accuracy of large language model responses, especially in knowledge-intensive tasks, reducing erroneous outputs. By citing information sources, the accuracy of the answer can be verified, thereby increasing trust in the model's output.

[0067] The candidate information set consists of multiple candidate information pieces related to the training text, and these candidate pieces are arranged in a preset order. For example, if the training text is "Is xxx behavior a breach of contract?", then the candidate information related to the training text could be the corresponding contract clauses, and the candidate information set could be a set of information composed of multiple contract clauses arranged in the order of the clauses.

[0068] For example, typical user questions can be collected as training text through various methods such as vertical website crawling, expert editing, and log retrieval. For instance, the training text could be "Does xxx behavior constitute a breach of contract?". Then, the retrieval module in the retrieval enhancement system can retrieve a set of candidate information related to the training text from the retrieval database. For example, the candidate information set could be a collection of information composed of multiple contract clauses arranged in the order of the clauses.

[0069] Optionally, in one embodiment, the step "obtaining multiple training texts and a set of candidate information related to the training texts retrieved from the retrieval database" may include:

[0070] The raw data obtained from one or more data sources is vectorized to obtain vectorized raw data, and a retrieval library is built based on the vectorized raw data.

[0071] Obtain multiple training texts;

[0072] A set of candidate information related to the training text is retrieved from the retrieval database by comparing the vectorized representation of the training text with the original vectorized data.

[0073] For example, raw data can be obtained first through methods such as loading data in multiple formats and acquiring data from different data sources. Then, the obtained raw data is processed, including data filtering, data compression, and data formatting. The processed data is then segmented, and the segmented data is vectorized to obtain vectorized raw data. After vectorization, an index can be built and written to a database to obtain a usable retrieval library. After obtaining multiple training texts, the vectorized representations of the training texts can be compared with the vectorized raw data to retrieve a set of candidate information related to the training texts from the retrieval library. Retrieval methods can include similarity retrieval, full-text retrieval, etc., and multiple retrieval methods can be combined to improve recall.

[0074] S202. Based on multiple training texts and the candidate information set, construct a training sample set.

[0075] In one embodiment, for example, the training sample set includes multiple training samples, which include training text and corresponding training text response results. The training text response results are the correct response methods for the training text, and the training text response results include citation markers, which are used to characterize the citation of candidate information in the candidate information set by the training text response results. For example, if there are multiple clauses in the candidate information set, citation markers [1], [2], etc. can be added and marked at the end of the sentence describing the corresponding clause. If the same clause is mentioned multiple times in the answer, the same citation marker can be used.

[0076] Optionally, in one embodiment, the initial large language model can be trained using a training sample set. Before the step "calling the large language model to be optimized to perform text response on the training text based on the perturbed candidate information set, and generating predicted text response results", the following may be included:

[0077] Obtain the initial large language model and the preset text output instructions that instruct the initial large language model to generate results that meet the current business requirements;

[0078] The initial large language model is invoked to generate an initial predicted text response result that satisfies the preset text output instructions by responding to the training text based on the candidate information set.

[0079] Based on the difference between the initial predicted text response results and the training text response results, the initial large language model is trained until a large language model to be optimized is obtained that meets the training termination condition.

[0080] In one embodiment, an initial large language model can be obtained, which is an untrained large language model. A preset text output prompt is also obtained, instructing the initial large language model to generate results that meet current business requirements. Training text can then be input into the initial large language model. The initial large language model can then respond to the training text using a set of candidate information according to the preset text output prompt, generating an initial predicted text response that satisfies the preset text output prompt. Since the model is not fully trained, there is a difference between this initial predicted text response and the correct training text response. The initial large language model can be trained using the difference between the initial predicted text response and the training text response until a large language model to be optimized that meets the training termination condition is obtained. Applying the above steps ensures that the generated text response includes citation markers and cites content from the given set of candidate information, increasing the accuracy of model generation and reducing model illusions.

[0081] For example, the preset text output instruction could be: "You are a real estate agent. Here is a question about rent and a rental contract. You need to generate a response to the question based on the rental contract."

[0082] The specific requirements are as follows:

[0083] The reply body is a paragraph, following the format of "Based on" + "Summary". "Based on" refers to the specific content in the cited clauses that can solve the problem, and the reply summarizes the response to the user's question. Add citation marks [1], [2], ... at the end of the sentence describing the corresponding clause. When the same clause is mentioned multiple times in the reply, use the same citation mark.

[0084] The training text could be "Does behavior xxx constitute a breach of contract?"

[0085] The given set of candidate information includes "Contract Clause 1, Contract Clause 2, ...

[0086] The embodiment of this application aims to output the following result from the large language model to be optimized: "According to the provisions of contract clause xx, it can be determined that xxx, therefore, the behavior of xxx is a breach of contract / not a breach of contract."

[0087] Optionally, in one embodiment, since the return results of the large language model are not necessarily all correct and contain a lot of noise, a data denoising process is required. The step "training the initial large language model based on the difference between the initial predicted text response results and the training text response results until an optimized large language model that meets the training termination condition is obtained" may include:

[0088] Based on the difference between the initial predicted text response results and the training text response results, the initial large language model is trained until a fine-tuned large language model that meets the training termination condition is obtained.

[0089] The large language model to be optimized before fine-tuning is invoked to perform text responses to the training text based on the candidate information set, generating the predicted text response results before denoising.

[0090] The predicted text response results before denoising are denoised to obtain the predicted text response results after denoising.

[0091] Based on the denoised predicted text response results, the large language model to be optimized is fine-tuned until the large language model to be optimized is obtained.

[0092] For example, a denoising process can be performed before training the large language model to be optimized. That is, the large language model to be optimized is trained first. However, the return results of the large language model at this time may not be correct. Therefore, the predicted text response results before denoising can be denoised to obtain the predicted text response results after denoising. The model can then be fine-tuned using the predicted text response results after denoising, so that the large language model has the ability to generate a response in a specified format based on the input text and candidate information.

[0093] Optionally, in one embodiment, the step "denoising the predicted text response result before denoising to obtain the predicted text response result after denoising" may include:

[0094] The text response results after deletion are obtained by removing erroneous text responses that do not follow the training text or are not factual from the predicted text response results before denoising.

[0095] Remove responses from the deleted text responses that are inconsistent with the conclusions of the training text responses to obtain the denoised predicted text responses.

[0096] In one embodiment, denoising can include several methods. Firstly, denoising can be performed by removing reference illusions from the output, where large model illusions occur when the generated content does not match reality or the context. Secondly, denoising can be performed by determining whether the conclusions are consistent.

[0097] For example, since large language models may not strictly follow the instructions in their output, such data can be removed directly. That is, erroneous text responses that do not follow the training text or are not true can be deleted from the predicted text response results before denoising to obtain the text response results after deletion. This is the process of removing citation illusions from the output results.

[0098] Then, in order to further improve the accuracy of the model output, a consistency judgment can be made. That is, the answer corresponding to the training text is output by the large language model to be optimized before fine-tuning, and compared with the response result of the training text. If it is determined that the conclusions of the two results are inconsistent, the inconsistent result can be removed.

[0099] S203. Perturb the order of multiple candidate information in the candidate information set to obtain the perturbed candidate information set.

[0100] The perturbation-induced candidate information set includes multiple candidate information items arranged in the order after perturbation.

[0101] In one embodiment, by borrowing the idea of ​​rejection sampling, perturbation can be added to the model input to obtain multiple output results, and then problems that the model cannot solve can be filtered out. This perturbation of candidate information can be achieved using a random shuffle function. The random shuffle function has an iterator pointing to the first element of the sequence as its first parameter, a parameter pointing to the position after the last element of the sequence as its second parameter, and also includes a custom random number generator, thereby changing the order of multiple candidate information in the candidate information set.

[0102] Rejection sampling is a Monte Carlo method that estimates the probability of a target distribution by sampling from a simple reference distribution and then accepting or rejecting the samples based on the ratio of that distribution to a target distribution. Common rejection sampling methods involve changing model parameters, such as randomizing seed values, but these are not applicable to this embodiment because the training data has already been fitted, and changing the parameters will not yield differentiated results. This embodiment can randomize the order of the input candidate information as input perturbation information, thereby identifying unstable models.

[0103] In one embodiment, each candidate information in the original candidate information set is numbered, such as from 1 to 10. The positions of multiple candidate information in the candidate information set can be randomly shuffled, and the mapping from the original position to the new position of the candidate information is recorded. For example, the original candidate information set includes [0, 1, 2, ..., 9], and the perturbed candidate information set includes [7, 6, 5, ..., 1]. That is, the perturbation makes the candidate information originally at position [1] change to position [6].

[0104] S204. Call the large language model to be optimized to perform text response on the training text based on the perturbed candidate information set, and generate the predicted text response result.

[0105] For example, a trained large language model to be optimized can be used to generate a text response to the training text using a perturbed set of candidate information, and then generate a predicted text response result. In this case, after perturbing the order of multiple candidate information in the candidate information set, the reference markers in the generated predicted text response result must also be changed accordingly. When only the reference markers are changed, the conclusion of the correct answer remains unchanged. For example, the original reference marker [1] will be changed to the reference marker [6].

[0106] S205. Evaluate the consistency between the predicted text response results and the training text response results, and construct preference data pairs corresponding to the training text based on the evaluation results.

[0107] For example, the above steps yield two results: a predicted text response from the model and a correct training text response. If the model has a strong fitting ability, theoretically, the candidate information and conclusions cited in these two results should be consistent. However, this is not always the case. To improve the accuracy of the model's predictions, we can compare the consistency of the conclusions of the two results to determine if there are any problems with the model's predictions, allowing for further model optimization. This can be achieved by constructing preference data pairs, which include the correct training text response and the incorrect predicted text response.

[0108] Optionally, in one embodiment, the step of "evaluating the consistency between the predicted text response results and the training text response results, and constructing preference data pairs corresponding to the training texts based on the evaluation results" may include:

[0109] Obtain the predicted text response results corresponding to the training text, as well as the training text response results;

[0110] The consistency evaluation model is invoked to evaluate the consistency between the predicted text response results corresponding to the training text and the training text response results, and the evaluation results corresponding to the training text are obtained.

[0111] Based on the evaluation results corresponding to the training text, the predicted text response results corresponding to the training text, and the training text response results, preference data pairs corresponding to the training text are constructed.

[0112] For example, a consistency evaluation model can be applied to evaluate the consistency between the predicted text response results corresponding to the training text and the training text response results, thereby determining whether the conclusions of the two results are consistent.

[0113] Optionally, in one embodiment, the step of "calling the consistency evaluation model to perform a consistency evaluation on the predicted text response results corresponding to the training text and the training text response results, and obtaining the evaluation results corresponding to the training text" may include:

[0114] Obtain the consistency evaluation model and the preset result evaluation instructions that instruct the consistency evaluation model to generate evaluation results;

[0115] The consistency evaluation model is invoked to evaluate the consistency between the predicted text response results and the training text response results, and the evaluation results corresponding to the training text that meet the preset result evaluation instructions are obtained.

[0116] Optionally, in one embodiment, the preset result evaluation instructions include conclusion consistency evaluation instructions and induction correctness evaluation instructions.

[0117] For example, in order to instruct the consistency evaluation model to perform a conclusion consistency evaluation on two results, a preset result evaluation instruction can be set in advance. This preset result evaluation instruction can include a consistency evaluation instruction and an inductive correctness evaluation instruction. The consistency evaluation instruction is used to instruct the consistency evaluation model to perform a conclusion consistency evaluation on two results, that is, to determine whether the conclusions of the two results are consistent. The inductive correctness evaluation instruction is used to instruct the consistency evaluation model to determine whether the conclusion is derived from a correct induction.

[0118] Optionally, in one embodiment, the step of "constructing preference data pairs corresponding to the training text based on the evaluation results corresponding to the training text, the predicted text response results corresponding to the training text, and the training text response results" may include:

[0119] Select training texts that fail the evaluation and are deemed to have failed the evaluation;

[0120] Based on the predicted text response results that do not correspond to the training text, and the training text response results, construct preference data pairs that do not correspond to the training text.

[0121] The evaluation results include passing and failing results. Furthermore, incorrect conclusions and incorrect generalizations can also be defined as failing results.

[0122] For example, since discussion is only necessary for evaluating the failure results, in order to reduce the workload of the model, we can filter out the training texts whose evaluation results are considered failure results from the training texts, and construct preference data pairs only for this part of the training texts.

[0123] In one embodiment, a consistency evaluation model and a preset result evaluation prompt can be obtained. The preset result evaluation prompt is: "You are a real estate agent. You are given a housing rental-related question, a correct answer, and a candidate answer. You need to judge the quality of the candidate answer. Please output the following three steps, separating the output of different steps with newlines."

[0124] Step 1: Determine whether the conclusions of the candidate answer and the correct answer are consistent. If they are consistent, output "Conclusions are consistent"; otherwise, output "Conclusions are inconsistent" and give the reason.

[0125] Step 2: Determine whether the cited clauses of the candidate answer can lead to the final conclusion. If yes, output "Correct conclusion"; otherwise, output "Incorrect conclusion" and give the reason.

[0126] Step 3: If the results of the above two steps are "Conclusion Consistent" and "Induction Correct", output "Candidate Answer Passes" directly; otherwise, output "Candidate Answer Fails".

[0127] The training text is "Does behavior xxx constitute a breach of contract?"

[0128] The correct training text response is: "According to the contract clause xx, it can be determined that xxx, therefore xxx's behavior constitutes a breach of contract."

[0129] The predicted text response obtained by the model is: "According to the provisions of contract clause xx, it can be known that xxx, therefore, the behavior of xxx is not a breach of contract."

[0130] Inputting the two results above into the consistency evaluation model will output "Step 1: Conclusions are inconsistent".

[0131] Reason: The correct answer points out that the behavior xxx is a breach of contract, while the candidate answer argues that the behavior xxx is not a breach of contract.

[0132] Step 2: Summarize the errors

[0133] Reason: The candidate answers misunderstood the contract terms, leading to an incorrect induction.

[0134] Step 3: Candidate answers are rejected.

[0135] By following the steps above, each training text corresponds to an evaluation result, which can then be used to filter out unstable training data.

[0136] Then, training texts that fail the evaluation can be filtered out (data with inconsistent conclusions or inferences can also be filtered separately). The predicted text responses output by this part of the data are not entirely correct, so preference data pairs can be constructed. The format can be <correct data, incorrect data>. In the preference data pair, the correct data can be the training text response, and the incorrect data can be the predicted text response. This means that the response of the correct data is better than the response of the incorrect data.

[0137] The above methods can improve the stability of the model without reconstructing a new problem or new training data.

[0138] S206. Based on the preference data pairs corresponding to the training texts, optimize the large language model to be optimized to obtain the optimized large language model.

[0139] For example, after obtaining preference data pairs, a Direct Preference (DPO) model can be applied to optimize a large language model, resulting in an optimized model. The DPO algorithm directly optimizes user or expert preferences, rather than relying on traditional cumulative rewards. In the DPO algorithm, different decision sequences or strategies are compared, and the model is optimized based on user or expert preferences, so that the final strategy better matches the expected behavior. The DPO algorithm is typically used in scenarios where the reward function is difficult to define explicitly, or in applications where user preferences need to be directly encoded into the decision-making process. Implementing the DPO algorithm requires building a preference model that can learn from user or expert feedback. In practical applications, a mechanism may need to be designed to collect user preference data, such as through comparative queries or ranking feedback. This data is then used to train one or more models that can predict preference scores for a given decision sequence and optimize the strategy accordingly.

[0140] Optionally, in one embodiment, the step "optimizing the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain the optimized large language model" may include:

[0141] Construct a loss function corresponding to the direct preference optimization process, and build a direct preference optimization model based on the loss function;

[0142] The direct preference optimization model is invoked to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, resulting in the optimized large language model.

[0143] In one embodiment, preference data can be added to the large language model to be optimized and fed into a direct preference optimization model for training. Since the model can be explicitly provided with information on what constitutes a good or bad answer to different questions, it can learn these differences and improve upon its shortcomings. For example, ... Figure 6 As shown, the large language model to be optimized and preference data pairs can be input into the direct preference optimization model to output a more stable optimized large language model. Since a single loop may not be able to completely solve the instability problem, an iterative optimization process can be used.

[0144] Optionally, in one embodiment, the step "constructing the loss function corresponding to the direct preference optimization process" may include:

[0145] Obtain the first cumulative probability difference between the correct data and the initial predicted text response result, and the second cumulative probability difference between the incorrect data and the initial predicted text response result;

[0146] Based on the first and second cumulative probability differences, a loss function corresponding to the direct preference optimization process is constructed.

[0147] For example, a loss function can be constructed as shown in Equation 1:

[0148]

[0149] This can be simplified by expanding the fraction in the logarithm of Equation 1, setting β to 1, and temporarily ignoring the log sigmoid. Equation 1 can then be simplified to Equation 2 as follows:

[0150] [logp(y w )-logp ref (y w )]-[logp(y l )-logp ref (y l )]

[0151] Formula 2

[0152] Since the initial loss function includes a negative sign, the optimization objective is to maximize Equation 2, i.e., [logp(y w )-logp ref (y w )] and [logp(y l )-logpref (y l The greater the difference between [logp(y)], the better. w )-logp ref (y w [logp(y)] represents the first cumulative probability difference between the correct data and the initial predicted text response, where [logp(y)] represents the first cumulative probability difference between the correct data and the initial predicted text response. l )-logp ref (y l The second cumulative probability difference between the erroneous data and the initial predicted text response is represented by [logp(y). This means that when the probability of correct data increases and the probability of erroneous data decreases, [logp(y)] can be achieved. w )-logp ref (y w )] and [logp(y l )-logp ref (y l The differences between them become larger.

[0153] In one embodiment, β is a temperature hyperparameter. In practical applications, the value of β can be adjusted by iterating through the parameters to find the optimal value, thereby improving the performance of the model.

[0154] In addition, when training a direct preference model, not only can preference data pairs be added, but the correct data in the preference data pairs can also be used as fine-tuning data for simultaneous training. Different weights can be designed for the two parts of the data to jointly strengthen the learning of unstable samples.

[0155] As described above, this embodiment can obtain multiple training texts and a set of candidate information related to the training texts retrieved from the retrieval database; construct a training sample set based on the multiple training texts and the candidate information set; perturb the order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating predicted text response results; evaluate the consistency between the predicted text response results and the training text response results, and construct preference data pairs corresponding to the training texts based on the evaluation results; optimize the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain the optimized large language model. This application can reduce model illusion by adding citation markers to the text response results, thereby improving the accuracy of the model's answers; and by perturbing the position of candidate information, it can improve the stability and final effect of the model without changing the retrieval module or adding new training data. At the same time, it uses preference data to further optimize the large language model to be optimized, thereby addressing the model's shortcomings in a targeted manner to improve the retrieval enhancement generation capability of the large language model.

[0156] Based on the methods described in the preceding embodiments, the following will provide a more detailed explanation by taking the specific integration of the large language model optimization device into an electronic device as an example. This application provides a method for optimizing a large language model, such as... Figure 3 As shown, the specific process of the model optimization method for this large language model can be as follows:

[0157] S301. Electronic equipment acquires training sample set.

[0158] In one embodiment, such as Figure 4 As shown, the training sample set includes multiple training samples, which include training text and corresponding training text response results. The training text response results are the correct responses to the training text. Furthermore, the training text response results include reference markers, which are used to characterize the references made by the training text response results to candidate information in the candidate information set.

[0159] For example, electronic devices can collect typical user questions as training text through various methods such as vertical website crawling, expert editing, and log retrieval. For instance, the training text could be "Does xxx behavior constitute a breach of contract?". Then, the retrieval module in the retrieval enhancement system retrieves a set of candidate information related to the training text from the retrieval database. For example, the candidate information set could be a collection of multiple contract clauses arranged in order. Finally, the correct response to the training text can be obtained based on the training text itself, or the training text and the candidate information set together; that is, the training text response result.

[0160] S302. Electronic devices train an initial large language model using a training sample set to obtain a large language model to be optimized before fine-tuning.

[0161] In one embodiment, the electronic device can acquire an untrained initial large language model and preset text output instructions, and train the initial large language model using a training sample set. However, the output of the large language model to be optimized before fine-tuning after training still has room for improvement, so further steps such as denoising are required.

[0162] For example, the preset text output instruction could be: "You are a real estate agent. Here is a question about rent and a rental contract. You need to generate a response to the question based on the rental contract."

[0163] The specific requirements are as follows:

[0164] The reply body is a paragraph, following the format of "Based on" + "Summary". "Based on" refers to the specific content in the cited clauses that can solve the problem, and the reply summarizes the response to the user's question. Add citation marks [1], [2], ... at the end of the sentence describing the corresponding clause. When the same clause is mentioned multiple times in the reply, use the same citation mark.

[0165] The training text could be "Does behavior xxx constitute a breach of contract?"

[0166] The given set of candidate information includes "Contract Clause 1, Contract Clause 2, ...

[0167] The training text response result could be: "According to the contract clause xx, it can be determined that xxx, therefore xxx's behavior is a breach of contract / not a breach of contract."

[0168] S303. The electronic device performs denoising processing on the predicted text response result before denoising to obtain the predicted text response result after denoising.

[0169] For example, electronic devices can remove erroneous text responses that do not follow the training text or are factually inaccurate from the predicted text responses before denoising; this is the process of removing citation illusions from the output. Furthermore, the answer corresponding to the training text can be output by the large language model to be optimized before fine-tuning, and compared with the training text response corresponding to the training text. If the conclusions of the two results are inconsistent, the inconsistent result can be removed.

[0170] S304. Based on the denoised predicted text response results, the electronic device performs model fine-tuning on the large language model to be optimized before fine-tuning, and obtains the large language model to be optimized.

[0171] S305. The electronic device perturbs the order of multiple candidate information in the candidate information set to obtain the perturbed candidate information set.

[0172] For example, the original candidate information set includes [0,1,2,...,9], and the perturbation candidate information set includes [7,6,5,...,1]. That is to say, the perturbation makes the candidate information originally at position [1] change to position [6].

[0173] S306. The electronic device calls the large language model to be optimized to perform text responses to the training text and generates predicted text response results.

[0174] For example, such as Figure 5 As shown, when generating the predicted text response results, the corresponding citation markers will also change.

[0175] S307. The electronic device evaluates the consistency between the predicted text response results and the training text response results.

[0176] For example, an electronic device can acquire a consistency assessment model and a preset result assessment instruction. This preset instruction might be: "You are a real estate agent. You are given a housing rental-related question, a correct answer, and a candidate answer. You need to judge the quality of the candidate answer. Please output the following three steps, separating the different steps with newlines."

[0177] Step 1: Determine whether the conclusions of the candidate answer and the correct answer are consistent. If they are consistent, output "Conclusions are consistent"; otherwise, output "Conclusions are inconsistent" and give the reason.

[0178] Step 2: Determine whether the cited clauses of the candidate answer can lead to the final conclusion. If yes, output "Correct conclusion"; otherwise, output "Incorrect conclusion" and give the reason.

[0179] Step 3: If the results of the above two steps are "Conclusion Consistent" and "Induction Correct", output "Candidate Answer Passes" directly; otherwise, output "Candidate Answer Fails".

[0180] The training text is "Does behavior xxx constitute a breach of contract?"

[0181] The correct training text response is: "According to the contract clause xx, it can be determined that xxx, therefore xxx's behavior constitutes a breach of contract."

[0182] The predicted text response obtained by the model is: "According to the provisions of contract clause xx, it can be known that xxx, therefore, the behavior of xxx is not a breach of contract."

[0183] Inputting the two results above into the consistency evaluation model will output "Step 1: Conclusions are inconsistent".

[0184] Reason: The correct answer points out that the behavior xxx is a breach of contract, while the candidate answer argues that the behavior xxx is not a breach of contract.

[0185] Step 2: Summarize the errors

[0186] Reason: The candidate answers misunderstood the contract terms, leading to an incorrect induction.

[0187] Step 3: Candidate answers are rejected.

[0188] S308. The electronic device constructs preference data pairs corresponding to the training text based on the evaluation results.

[0189] For example, electronic devices can filter out training texts that fail the evaluation (and can also select data with inconsistent or incorrect conclusions to filter separately). The predicted text responses output by this part of the data are not entirely correct, so a pair of preference data can be constructed, which can be in the format of <correct data, incorrect data>. In the pair of preference data, the correct data can be the training text response, and the incorrect data can be the predicted text response, which means that the response method of the correct data is better than the response method of the incorrect data.

[0190] S309. The electronic device optimizes the large language model to be optimized based on preference data to obtain the optimized large language model.

[0191] For example, electronic devices can acquire a direct preference optimization model, whose loss function can be shown in Equation 1:

[0192]

[0193] Then, by inputting the large language model to be optimized and the preference data pairs into the direct preference optimization model, a more stable optimized large language model can be output.

[0194] As can be seen from the above, this embodiment can acquire a training sample set through an electronic device; train an initial large language model using the training sample set to obtain a large language model to be optimized before fine-tuning; denoise the predicted text response results before denoising to obtain denoised predicted text response results; fine-tune the large language model to be optimized before fine-tuning based on the denoised predicted text response results to obtain a large language model to be optimized; perturb the order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training text and generate predicted text response results; evaluate the consistency between the predicted text response results and the training text response results; construct preference data pairs corresponding to the training text based on the evaluation results; optimize the large language model to be optimized based on the preference data to obtain an optimized large language model. This application can reduce model illusion by adding citation markers to the text response results, thereby improving the accuracy of the model's answers; and by perturbing the position of candidate information, it can improve the stability of the model and the final effect without changing the retrieval module or adding new training data. At the same time, preference data is used to further optimize the large language model, thereby addressing its shortcomings and improving its retrieval and generation capabilities.

[0195] To better implement the above methods, embodiments of this application also provide a model optimization device for large language models, such as... Figure 7As shown, the model optimization device for this large language model may include an acquisition unit 701, a construction unit 702, a perturbation unit 703, a prediction unit 704, an evaluation unit 705, and an optimization unit 706, as follows:

[0196] The acquisition unit 701 is used to acquire multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval library. The set of candidate information includes multiple candidate information arranged in a preset order.

[0197] The construction unit 702 is used to construct a training sample set based on multiple training texts and the candidate information set. The training sample set includes multiple training samples, each training sample includes the training text and the training text response result corresponding to the training text. The training text response result includes a reference marker, which is used to characterize the reference of the training text response result to the candidate information in the candidate information set.

[0198] The perturbation unit 703 is used to perturb the arrangement order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set, wherein the perturbed candidate information set includes multiple candidate information arranged in the perturbed order;

[0199] Prediction unit 704 is used to call the large language model to be optimized to perform text response on the training text based on the perturbed candidate information set, and generate predicted text response results;

[0200] Evaluation unit 705 is used to evaluate the consistency between the predicted text response result and the training text response result, and to construct a preference data pair corresponding to the training text based on the evaluation result. The preference data pair includes the training text response result as correct data and the predicted text response result as incorrect data.

[0201] The optimization unit 706 is used to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, so as to obtain the optimized large language model.

[0202] Optionally, in some embodiments of this application, the prediction unit may further include a first acquisition subunit, a first invocation subunit, and a first training subunit, as follows:

[0203] The first acquisition subunit is used to acquire an initial large language model and a preset text output instruction that instructs the initial large language model to generate results that meet the current business requirements.

[0204] The first calling subunit is used to call the initial large language model to perform text response on the training text based on the candidate information set, and generate an initial predicted text response result that satisfies the preset text output instruction;

[0205] The first training subunit is used to train the initial large language model based on the difference between the initial predicted text response results and the training text response results, until a large language model to be optimized is obtained that meets the training termination condition.

[0206] Optionally, in some embodiments of this application, the first training subunit may include a second training subunit, a second calling subunit, a denoising subunit, and a fine-tuning subunit, as follows:

[0207] The second training subunit is used to train the initial large language model based on the difference between the initial predicted text response result and the training text response result, until a fine-tuned large language model to be optimized is obtained that meets the training termination condition.

[0208] The second calling subunit is used to call the large language model to be optimized before fine-tuning to perform a text response to the training text based on the candidate information set, and generate a predicted text response result before denoising.

[0209] The denoising subunit is used to denoise the predicted text response result before denoising to obtain the predicted text response result after denoising.

[0210] The fine-tuning subunit is used to fine-tune the large language model to be optimized before fine-tuning based on the denoised predicted text response results, until the large language model to be optimized is obtained.

[0211] Optionally, in some embodiments of this application, the denoising subunit may be specifically used to delete erroneous text response results that do not follow the training text or are inconsistent with the facts from the predicted text response results before denoising, to obtain the text response results after deletion; and to remove response results that are inconsistent with the conclusion of the training text response results from the text response results after deletion, to obtain the predicted text response results after denoising.

[0212] Optionally, in some embodiments of this application, the evaluation unit may include a second acquisition subunit, a third invocation subunit, and a first construction subunit, as follows:

[0213] The second acquisition subunit is used to acquire the predicted text response result corresponding to the training text, and the training text response result;

[0214] The third calling subunit is used to call the consistency evaluation model to evaluate the consistency of the predicted text response result corresponding to the training text and the training text response result, and obtain the evaluation result corresponding to the training text.

[0215] The first construction subunit is used to construct the preference data pair corresponding to the training text based on the evaluation result corresponding to the training text, the predicted text response result corresponding to the training text, and the training text response result.

[0216] Optionally, in some embodiments of this application, the third calling subunit may be specifically used to obtain a consistency evaluation model and a preset result evaluation instruction that instructs the consistency evaluation model to generate evaluation results; call the consistency evaluation model to perform result consistency evaluation on the predicted text response result corresponding to the training text and the training text response result, and obtain the evaluation result corresponding to the training text that satisfies the preset result evaluation instruction.

[0217] Optionally, in some embodiments of this application, the first construction subunit may be specifically used to filter out training texts from the training texts whose evaluation result is the evaluation failure result; and construct preference data pairs corresponding to the training texts based on the predicted text response results corresponding to the training texts and the training text response results.

[0218] Optionally, in some embodiments of this application, the preset result evaluation instructions include conclusion consistency evaluation instructions and inductive correctness evaluation instructions.

[0219] Optionally, in some embodiments of this application, the optimization unit may include a second construction subunit and a fourth invocation subunit, as follows:

[0220] The second construction subunit is used to construct the loss function corresponding to the direct preference optimization process, and to construct the direct preference optimization model based on the loss function;

[0221] The fourth calling subunit is used to call the direct preference optimization model to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, so as to obtain the optimized large language model.

[0222] Optionally, in some embodiments of this application, the second construction subunit may be specifically used to obtain a first cumulative probability difference between the correct data and the initial predicted text response result, and a second cumulative probability difference between the incorrect data and the initial predicted text response result; based on the first cumulative probability difference and the second cumulative probability difference, a loss function corresponding to the direct preference optimization process is constructed.

[0223] Optionally, in some embodiments of this application, the acquisition unit may be specifically used to perform data vectorization processing on the raw data obtained from one or more data sources to obtain vectorized raw data, and construct a retrieval library based on the vectorized raw data; acquire multiple training texts; and retrieve a set of candidate information related to the training texts from the retrieval library by comparing the vectorized representations corresponding to the training texts with the vectorized raw data.

[0224] As can be seen from the above, this embodiment can acquire multiple training texts and a set of candidate information related to the training texts retrieved from the retrieval database through the acquisition unit 701; construct a training sample set based on the multiple training texts and the candidate information set through the construction unit 702; perturb the order of multiple candidate information in the candidate information set through the perturbation unit 703 to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training texts based on the perturbed candidate information set through the prediction unit 704 to generate a predicted text response result; evaluate the consistency between the predicted text response result and the training text response result through the evaluation unit 705, and construct the preference data pair corresponding to the training text based on the evaluation result; and optimize the large language model to be optimized based on the preference data pair corresponding to the training text through the optimization unit 706 to obtain the optimized large language model. This application can reduce model illusion by adding citation markers to the text response result and improve the accuracy of the model's answer; and improve the stability and final effect of the model without changing the retrieval module or adding new training data by perturbing the position of candidate information. At the same time, preference data is used to further optimize the large language model, thereby addressing its shortcomings and improving its retrieval and generation capabilities.

[0225] This application also provides an electronic device, such as... Figure 8 The diagram shows a structural schematic of an electronic device involved in an embodiment of this application. This electronic device can be a terminal or a server, specifically:

[0226] The electronic device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0227] The processor 801 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes various functions and processes data by running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801.

[0228] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0229] The electronic device also includes a power supply 803 that supplies power to the various components. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0230] The electronic device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0231] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 802 according to the following instructions, and the processor 801 runs the applications stored in the memory 802 to realize various functions, as follows:

[0232] This application provides a method and related equipment for optimizing a large language model. The method involves acquiring multiple training texts and a set of candidate information related to the training texts retrieved from a database; constructing a training sample set based on the multiple training texts and the candidate information set; perturbing the order of multiple candidate information items in the candidate information set to obtain a perturbed candidate information set; calling the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating a predicted text response result; evaluating the consistency between the predicted text response result and the training text response result, and constructing preference data pairs corresponding to the training texts based on the evaluation results; and optimizing the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain an optimized large language model.

[0233] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0234] As described above, this embodiment can obtain multiple training texts and a set of candidate information related to the training texts retrieved from the retrieval database; construct a training sample set based on the multiple training texts and the candidate information set; perturb the order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set; call the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating predicted text response results; evaluate the consistency between the predicted text response results and the training text response results, and construct preference data pairs corresponding to the training texts based on the evaluation results; optimize the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain the optimized large language model. This application can reduce model illusion by adding citation markers to the text response results, thereby improving the accuracy of the model's answers; and by perturbing the position of candidate information, it can improve the stability and final effect of the model without changing the retrieval module or adding new training data. At the same time, it uses preference data to further optimize the large language model to be optimized, thereby addressing the model's shortcomings in a targeted manner to improve the retrieval enhancement generation capability of the large language model.

[0235] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0236] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the model optimization methods for large language models provided in embodiments of this application. For example, the instructions can execute the following steps:

[0237] This application provides a method and related equipment for optimizing a large language model. The method involves acquiring multiple training texts and a set of candidate information related to the training texts retrieved from a database; constructing a training sample set based on the multiple training texts and the candidate information set; perturbing the order of multiple candidate information items in the candidate information set to obtain a perturbed candidate information set; calling the large language model to be optimized to respond to the training texts based on the perturbed candidate information set, generating a predicted text response result; evaluating the consistency between the predicted text response result and the training text response result, and constructing preference data pairs corresponding to the training texts based on the evaluation results; and optimizing the large language model to be optimized based on the preference data pairs corresponding to the training texts to obtain an optimized large language model.

[0238] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0239] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0240] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the large language model optimization methods provided in the embodiments of this application, the beneficial effects that any of the large language model optimization methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0241] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various alternative implementations of the model optimization aspect of the aforementioned large language model.

[0242] The above provides a detailed description of a large language model optimization method and related equipment provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A model optimization method for large language models, characterized in that, include: Obtain multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval database, wherein the set of candidate information includes multiple candidate information arranged in a preset order; Based on multiple training texts and the candidate information set, a training sample set is constructed. The training sample set includes multiple training samples, each training sample including the training text and the corresponding training text response result. The training text response result includes a reference marker, which is used to characterize the reference of the training text response result to the candidate information in the candidate information set. The order of multiple candidate information in the candidate information set is perturbed to obtain a perturbed candidate information set, which includes multiple candidate information arranged in the perturbed order. The large language model to be optimized is invoked to perform a text response on the training text based on the perturbated candidate information set, and a predicted text response result is generated. The consistency between the predicted text response results and the training text response results is evaluated, and a preference data pair corresponding to the training text is constructed based on the evaluation results. The preference data pair includes the training text response results as correct data and the predicted text response results as incorrect data. Based on the preference data pairs corresponding to the training texts, the large language model to be optimized is optimized to obtain the optimized large language model.

2. The model optimization method for a large language model according to claim 1, characterized in that, Before calling the large language model to be optimized to perform text response on the training text based on the perturbed candidate information set and generating the predicted text response result, the method further includes: Obtain an initial large language model and a preset text output instruction that instructs the initial large language model to generate results that meet the current business requirements; The initial large language model is invoked to perform text response on the training text based on the candidate information set, generating an initial predicted text response result that satisfies the preset text output instruction; Based on the difference between the initial predicted text response results and the training text response results, the initial large language model is trained until a large language model to be optimized that meets the training termination condition is obtained.

3. The model optimization method for large language models according to claim 2, characterized in that, The step of training the initial large language model based on the difference between the initial predicted text response result and the training text response result until a large language model to be optimized that meets the training termination condition is obtained includes: Based on the difference between the initial predicted text response results and the training text response results, the initial large language model is trained until a fine-tuned large language model that meets the training termination condition is obtained. The large language model to be optimized before fine-tuning is invoked to perform a text response on the training text based on the candidate information set, generating a predicted text response result before denoising. The predicted text response result before denoising is denoised to obtain the predicted text response result after denoising. Based on the denoised predicted text response results, the large language model to be optimized before fine-tuning is fine-tuned until the large language model to be optimized is obtained.

4. The model optimization method for a large language model according to claim 3, characterized in that, The step of denoising the predicted text response result before denoising to obtain the predicted text response result after denoising includes: The text response results after deletion are obtained by removing erroneous text response results that do not follow the training text or are not in accordance with the facts from the predicted text response results before denoising. Remove responses that are inconsistent with the conclusions of the training text responses from the deleted text response results to obtain the denoised predicted text response results.

5. The model optimization method for a large language model according to claim 1, characterized in that, The step of evaluating the consistency between the predicted text response results and the training text response results, and constructing preference data pairs corresponding to the training text based on the evaluation results, includes: Obtain the predicted text response result corresponding to the training text, and the training text response result; The consistency evaluation model is invoked to evaluate the consistency between the predicted text response result corresponding to the training text and the training text response result, thereby obtaining the evaluation result corresponding to the training text. Based on the evaluation results corresponding to the training text, the predicted text response results corresponding to the training text, and the training text response results, a preference data pair corresponding to the training text is constructed.

6. The model optimization method for a large language model according to claim 5, characterized in that, The consistency evaluation model is invoked to evaluate the consistency between the predicted text response result corresponding to the training text and the training text response result, thereby obtaining the evaluation result corresponding to the training text, including: Obtain a consistency evaluation model and a preset result evaluation instruction that instructs the consistency evaluation model to generate evaluation results; The consistency evaluation model is invoked to evaluate the consistency between the predicted text response result corresponding to the training text and the training text response result, so as to obtain the evaluation result corresponding to the training text that satisfies the preset result evaluation instruction.

7. The model optimization method for a large language model according to claim 5, characterized in that, The evaluation results include pass and fail results. The construction of preference data pairs corresponding to the training text based on the evaluation results corresponding to the training text, the predicted text response results corresponding to the training text, and the training text response results includes: Select training texts from the training texts whose evaluation results are deemed as failing the evaluation; Based on the predicted text response results corresponding to the training text that did not pass the training text, and the training text response results, a preference data pair corresponding to the training text that did not pass the training text is constructed.

8. The model optimization method for a large language model according to claim 6, characterized in that, The preset result evaluation instructions include conclusion consistency evaluation instructions and induction correctness evaluation instructions.

9. The model optimization method for a large language model according to claim 3, characterized in that, The step of optimizing the large language model to be optimized based on the preference data pairs corresponding to the training text to obtain the optimized large language model includes: Construct a loss function corresponding to the direct preference optimization process, and build a direct preference optimization model based on the loss function; The direct preference optimization model is invoked to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, thereby obtaining the optimized large language model.

10. The model optimization method for a large language model according to claim 9, characterized in that, The loss function corresponding to the construction of the direct preference optimization process includes: Obtain the first cumulative probability difference between the correct data and the initial predicted text response result, and the second cumulative probability difference between the incorrect data and the initial predicted text response result; Based on the first cumulative probability difference and the second cumulative probability difference, a loss function corresponding to the direct preference optimization process is constructed.

11. The model optimization method for a large language model according to claim 1, characterized in that, The acquisition of multiple training texts and a set of candidate information related to the training texts retrieved from a search library includes: The raw data obtained from one or more data sources is processed into vectorized raw data, and a retrieval library is constructed based on the vectorized raw data. Obtain multiple training texts; A set of candidate information related to the training text is retrieved from the retrieval library by comparing the vectorized representation of the training text with the original vectorized data.

12. A model optimization device for a large language model, characterized in that, include: The acquisition unit is used to acquire multiple training texts and a set of candidate information related to the training texts retrieved from a retrieval library, wherein the set of candidate information includes multiple candidate information arranged in a preset order. The construction unit is used to construct a training sample set based on multiple training texts and the candidate information set. The training sample set includes multiple training samples, each training sample includes the training text and the corresponding training text response result. The training text response result includes a reference marker, which is used to characterize the reference of the training text response result to the candidate information in the candidate information set. A perturbation unit is used to perturb the arrangement order of multiple candidate information in the candidate information set to obtain a perturbed candidate information set, wherein the perturbed candidate information set includes multiple candidate information arranged in the perturbed order; The prediction unit is used to call the large language model to be optimized to perform a text response on the training text based on the perturbed candidate information set, and generate a predicted text response result; An evaluation unit is used to evaluate the consistency between the predicted text response result and the training text response result, and to construct a preference data pair corresponding to the training text based on the evaluation result. The preference data pair includes the training text response result as correct data and the predicted text response result as incorrect data. An optimization unit is used to optimize the large language model to be optimized based on the preference data pairs corresponding to the training text, so as to obtain an optimized large language model.

13. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is used to run the application program within the memory to perform the operations in the model optimization method for a large language model according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps in the model optimization method for a large language model according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps in the model optimization method for a large language model as described in any one of claims 1 to 11.