Large language model training method and device and text query method and device
By introducing prompt sample sets and index pool-based incremental training methods in large language models, the problems of insufficient accuracy and waste of computing resources during training and querying of large language models are solved, and more efficient and accurate model training and querying effects are achieved.
Patent Information
- Application Number
- CN202510112818.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
Existing large language models have problems of insufficient accuracy and waste of computing resources during training and querying, especially when the entire model needs to be trained from scratch.
By introducing relevant prompt sample sets and sample query text to input into a large language model, the model can learn reward functions based on positive and negative samples in the prompt sample set, thereby generating more accurate answers; and perform incremental training based on the updated index pool instead of training the entire model from scratch, reducing computing resources and time costs.
It improves the answer accuracy and model quality of large language models, and greatly reduces computing resources and time costs, achieving more efficient model training and query processes.
Smart Images

Figure CN120045666A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, particularly to the fields of deep learning, natural language processing, and large models. Specifically, it relates to a training method for a large language model, a text query method, and their devices. Background Art
[0002] Large Language Model (LLM) is a type of natural language processing model based on deep learning technology, with extremely high language understanding and generation capabilities. With the improvement of computing power and the accumulation of large-scale data, large language models have been widely applied in recent years, including multiple fields such as text generation, machine translation, automatic question answering, and information retrieval. Summary of the Invention
[0003] The present disclosure provides a training method for a large language model, a text query method, a device, a device, and a storage medium.
[0004] According to one aspect of the present disclosure, a training method for a large language model is provided, including: determining a sample query text, and matching at least one set of prompt samples related to the sample query text from a preset index pool, where the index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples; jointly inputting the set of prompt samples and the sample query text into the large language model to be trained to obtain a sample answer text; obtaining parameters related to the accuracy of the sample answer text, and updating the index pool according to the parameters related to the accuracy; performing incremental training on the large language model based on the updated index pool to obtain the trained target large language model.
[0005] In this application, by jointly inputting the relevant set of prompt samples and the sample query text into the large language model, the model can learn the reward function implicitly based on the positive and negative samples in the set of prompt samples, so as to generate more accurate answers; by performing incremental training based on the updated index pool instead of training the entire model from scratch, it can greatly reduce the computing resources and time costs, and help improve the quality of the model more efficiently.
[0006] According to another aspect of the present disclosure, a text query method is provided, including: obtaining a target query text; matching at least one target set of prompt samples related to the target query text from a target index pool, where the target index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples; jointly inputting the target set of prompt samples and the target query text into a target large language model to obtain a to-be-determined answer text; determining a target answer text corresponding to the target query text according to the to-be-determined answer text.
[0007] According to another aspect of the present disclosure, there is provided a training device for a large language model, including: a matching module configured to determine a sample query text and match at least one set of prompt samples related to the sample query text from a preset index pool, where the index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; a determination module configured to jointly input the set of prompt samples and the sample query text into the large language model to be trained to obtain a sample answer text; an update module configured to obtain parameters related to the accuracy of the sample answer text and update the index pool according to the parameters related to the accuracy; and a generation module configured to perform incremental training on the large language model based on the updated index pool to obtain a trained target large language model.
[0008] According to another aspect of the present disclosure, there is provided a text query device, including: an acquisition module configured to acquire a target query text; a matching module configured to match at least one target set of prompt samples related to the target query text from a target index pool, where the target index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; a generation module configured to jointly input the target set of prompt samples and the target query text into a target large language model to obtain a to-be-determined answer text; and a determination module configured to determine a target answer text corresponding to the target query text according to the to-be-determined answer text.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned training method for a large language model and text query method.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the above-mentioned training method for a large language model and text query method.
[0011] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the above-mentioned training method for a large language model and text query method.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0014] Figure 1 It is a schematic diagram of an exemplary implementation manner of a training method for a large language model according to an exemplary embodiment of the present disclosure.
[0015] Figure 2 It is a schematic diagram of a preset index pool according to an exemplary embodiment of the present disclosure.
[0016] Figure 3 It is a schematic diagram of an exemplary implementation manner of a training method for a large language model according to an exemplary embodiment of the present disclosure.
[0017] Figure 4 It is a schematic diagram of an exemplary implementation manner of a training method for a large language model according to an exemplary embodiment of the present disclosure.
[0018] Figure 5 It is a schematic diagram of an exemplary implementation manner of a text query method according to an exemplary embodiment of the present disclosure.
[0019] Figure 6 It is a schematic diagram of an exemplary implementation manner of a text query method according to an exemplary embodiment of the present disclosure.
[0020] Figure 7 It is a schematic diagram of a training device for a large language model according to an exemplary embodiment of the present disclosure.
[0021] Figure 8 It is a schematic diagram of a text query device according to an exemplary embodiment of the present disclosure.
[0022] Figure 9 It is a schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0024] Deep Learning (DL for short) is a new research direction in the field of Machine Learning (ML for short). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analytical learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed previous related technologies.
[0025] Artificial Intelligence (AI for short) is a discipline that studies how to make computers simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It has technologies at both the hardware level and the software level. Artificial intelligence hardware technologies generally include several aspects such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0026] Natural Language Processing (NLP) is an important branch in the field of Artificial Intelligence (AI). Its aim is to enable computers to understand, generate, and interact with human language. It involves multiple disciplines, including linguistics, computer science, and mathematics, aiming to enable computers to process and analyze large amounts of natural language data (such as text, speech, etc.). The goal of NLP is to enable computers to understand the meaning of language like humans, so as to achieve intelligent interaction with humans.
[0027] In the technical solutions of this disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0028] Figure 1 is a schematic diagram of an exemplary implementation manner of a training method for a large language model shown in this application, as Figure 1 shown, the training method for this large language model includes the following steps:
[0029] S101, determine a sample query text, and match at least one set of prompt samples related to the sample query text from a preset index pool, where the index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0030] For ease of understanding, Figure 2It is a schematic diagram of a preset index pool shown in this application. As Figure 2 shown, the preset index pool includes N candidate sample sets, and each candidate sample set includes positive samples and negative samples. For example, candidate sample set 1 includes positive samples corresponding to historical query text 1 (generated based on historical query text 1 and the correct answer corresponding to historical query text 1), negative samples (generated based on historical query text 1 and the wrong answer corresponding to historical query text 1); candidate sample set 2 includes positive samples and negative samples corresponding to historical query text 2, and so on.
[0031] In this application, there is no limit on the number of positive samples and negative samples included in each candidate sample set.
[0032] In this application, the number of matching prompt sample sets can be set according to the actual situation.
[0033] In this application, it is first necessary to determine a sample query text. Optionally, the sample query text can be the text input by the user during an online query, that is, model training is performed based on online data.
[0034] After determining the sample query text, it is necessary to match at least one prompt sample set related to the sample query text from the preset index pool. Exemplarily, for a certain sample query text, there can be 5 prompt sample sets corresponding to this sample query text, which are candidate sample set 1, candidate sample set 3, candidate sample set 5, candidate sample set 8, and candidate sample set 9 respectively.
[0035] S102, Input the prompt sample set and the sample query text into the large language model to be trained to obtain a sample answer text.
[0036] After determining the prompt sample set corresponding to the sample query text as described above, the prompt sample set can be used as the prompt word for this sample query text, and input it together with this sample query text into the large language model to be trained to obtain the sample answer text corresponding to this sample query text output by the large language model.
[0037] S103, Obtain the parameters related to the accuracy of the sample answer text, and update the index pool according to the parameters related to the accuracy.
[0038] In this application, the update operation of the index pool can include adding a new candidate sample set to the index pool and keeping the current index pool unchanged. That is, it can be understood that in this application, adding a new candidate sample set to the index pool and keeping the current index pool unchanged can both be regarded as the update operation of the index pool.
[0039] As an implementable approach, after obtaining the sample answer text corresponding to the sample query text, determine whether the sample answer text is correct (i.e., whether the sample answer text solves the problem raised by the sample query text). If the sample answer text is determined to be correct, it indicates that the current large language model can already accurately understand and answer the sample query text. At this time, keep the current index pool unchanged.
[0040] As another implementable approach, after obtaining the sample answer text corresponding to the sample query text, determine whether the sample answer text is correct. If the sample answer text is determined to be incorrect, it indicates that the current large language model cannot accurately understand and answer the sample query text. To improve the model's ability, it is necessary to create a negative sample corresponding to the sample query text based on the sample query text and the sample answer text, and re-enter the prompt sample set and the sample query text into the large language model together to obtain a new sample answer text, and then determine whether the new sample answer text is correct. If the new sample answer text is determined to be correct, create a positive sample corresponding to the sample query text based on the sample query text and the new sample answer, and create a candidate sample set corresponding to the sample query text based on the positive sample and the previously created negative sample corresponding to the sample query text, and supplement it to the index pool; if the new sample answer texts obtained multiple times are all incorrect, the correct answer text corresponding to the sample query text can be manually marked, and a positive sample is created based on the sample query text and its corresponding correct answer text, and a candidate sample set corresponding to the sample query text is created based on the positive sample and the previously created negative sample, and supplemented to the index pool.
[0041] S104, perform incremental training on the large language model based on the updated index pool to obtain the trained target large language model.
[0042] After the above update of the index pool, obtain the second sample query text, and repeat the above steps S101 - S103 based on the second sample query text on the basis of the updated index pool to obtain a new updated index pool, and then continue to obtain the third sample query text, and so on, to perform incremental training of the model until the training is completed to obtain the trained target large language model.
[0043] An embodiment of the present application provides a training method for a large language model, including: determining a sample query text, and matching at least one set of prompt samples related to the sample query text from a preset index pool, where the index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples; inputting the set of prompt samples and the sample query text into the large language model to be trained to obtain a sample answer text; obtaining an accuracy-related parameter of the sample answer text, and updating the index pool according to the accuracy-related parameter; performing incremental training on the large language model based on the updated index pool to obtain a trained target large language model. By introducing the relevant set of prompt samples and jointly inputting them into the large language model with the sample query text, the model can learn the reward function implicitly based on the positive and negative samples in the set of prompt samples, so as to generate more accurate answers; by performing incremental training based on the updated index pool instead of training the entire model from scratch, the computing resources and time costs can be greatly reduced, which helps to improve the quality of the model more efficiently.
[0044] Figure 3 is a schematic diagram of an exemplary embodiment of a training method for a large language model shown in the present application, as Figure 3 shown, the training method of the large language model includes the following steps:
[0045] S301, determine a sample query text, and match at least one set of prompt samples related to the sample query text from a preset index pool, where the index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0046] In some embodiments, multiple first initial sample sets related to the sample query text are matched from the index pool based on the term frequency-inverse document frequency algorithm (BM25 algorithm), and then the first vector similarity between the sample query text and each first initial sample set is calculated, and a set of prompt samples related to the sample query text is screened out from the first initial sample sets based on the first vector similarity (for example, sorting the first vector similarities in descending order, and selecting the first initial sample sets corresponding to the top K first vector similarities as the set of prompt samples related to the sample query text).
[0047] Among them, calculating the first vector similarity between the sample query text and each first initial sample set includes: converting the sample query text into a first text vector; converting each first initial sample set into a second text vector respectively; calculating the vector similarity between the first text vector and each second text vector respectively as the first vector similarity.
[0048] Among them, the BM25 algorithm has high computational efficiency when matching to obtain a set of hint samples related to the sample query text, and can quickly screen out a preliminary first initial sample set from the index pool, shortening the screening and processing time. By calculating the first vector similarity between the sample query text and each first initial sample set, a set of hint samples that are highly relevant to the query text at the semantic level can be screened out more precisely. The combination of the two enables efficient screening of the hint sample set while enhancing the semantic matching ability and improving the subsequent training effect.
[0049] Among them, the number of the first initial sample sets here can be jointly determined based on the input length limit of the large language model and the threshold of the term frequency-inverse document frequency algorithm. For example, the first parameter k can be determined based on the input length limit of the large language model, and the second parameter w can be determined based on the threshold of the term frequency-inverse document frequency algorithm. Then, the number of the first initial sample sets is determined to be w×k. It is not difficult to understand that the number of hint sample sets related to the sample query text screened out from the first initial sample sets based on the first vector similarity is less than w×k, and the number of the finally determined hint sample sets related to the sample query text can be set to k.
[0050] S302, jointly input the hint sample set and the sample query text into the large language model to be trained to obtain a sample answer text.
[0051] S303, perform an accuracy evaluation on the sample answer text to obtain an evaluation result.
[0052] As an implementable method, input the sample query text and the sample answer text into a pre-trained accuracy evaluation model to obtain the evaluation result of the sample answer text.
[0053] As another implementable method, match at least one evaluation sample set related to the sample answer text from the index pool (continuing with Figure 2 as an example, the determined evaluation sample sets can be candidate sample set 2, candidate sample set 3, candidate sample set 4, candidate sample set 5, candidate sample set 6); then perform an accuracy evaluation on the sample answer text based on the evaluation sample set to obtain an evaluation result.
[0054] In some embodiments, the process of matching at least one evaluation sample set related to the sample answer text from the index pool includes: matching multiple second initial sample sets related to the sample answer text from the index pool based on the term frequency-inverse document frequency algorithm; calculating the second vector similarity between the sample answer text and each second initial sample set, and screening out the evaluation sample set related to the sample answer text from the second initial sample sets based on the second vector similarity (for example, sorting the second vector similarities in descending order, and selecting the top K second initial sample sets corresponding to the second vector similarities as the evaluation sample set related to the sample answer text).
[0055] Among them, calculating the second vector similarity between the sample answer text and each second initial sample set includes: converting the sample answer text into a first text vector; converting each second initial sample set related to the sample answer text into a second text vector respectively; calculating the vector similarity between the first text vector and each second text vector respectively as the second vector similarity.
[0056] In this way, the BM25 algorithm has high computational efficiency when matching the evaluation sample set related to the sample answer text, can quickly screen out the preliminary second initial sample set from the index pool, shortening the screening and processing time, and by calculating the second vector similarity between the sample answer text and each second initial sample set, it is possible to more accurately screen out the evaluation sample set that is highly semantically related to the sample answer text.
[0057] Among them, the number of evaluation sample sets can be set according to the actual situation. The number of evaluation sample sets can be the same as or different from the number of prompt sample sets related to the sample query text.
[0058] Among them, when obtaining the evaluation result by performing accuracy evaluation on the sample answer text based on the evaluation sample set, it may include: jointly inputting the sample answer text and the evaluation sample set into the large language model to obtain the evaluation result output by the large language model. In this way, the finally trained target large language model has not only improved the inference performance in obtaining the answer text, but also improved the evaluation performance in evaluating the answer text.
[0059] S304, perform accuracy annotation on the sample answer text to obtain the true label.
[0060] Among them, the true label can be considered the most accurate label.
[0061] S305, update the index pool by combining the evaluation result and the true label.
[0062] In this application, the update operation of the index pool may include adding a new candidate sample set to the index pool and keeping the current index pool unchanged.
[0063] In some embodiments, if the evaluation result is consistent with the true label result and both indicate that the sample answer text is correct, it means that the current large language model has been able to accurately understand and answer the sample query text. At this time, keep the current index pool unchanged, and the next training can be carried out based on the next sample query text combined with the current index pool.
[0064] In some embodiments, if the evaluation result is inconsistent with the true label result, the evaluation result indicates that the sample answer text is wrong, and the true label indicates that the sample answer text is correct. Taking the true label as the standard, it is also considered that the current large language model has been able to accurately understand and answer the sample query text. At this time, keep the current index pool unchanged, and the next training can be carried out based on the next sample query text combined with the current index pool.
[0065] In some embodiments, if the evaluation result is inconsistent with the true label result, the evaluation result indicates that the sample answer text is correct, and the true label indicates that the sample answer text is wrong. Taking the true label as the standard, a negative sample can be established based on the sample query text and its corresponding sample answer text, and the negative sample can be stored in the buffer area. The positive sample corresponding to the sample query text and its corresponding correct answer established manually can be obtained in a timely manner, and a new candidate sample set can be created based on the positive and negative samples of the sample query text and added to the current index pool.
[0066] In some embodiments, if the evaluation result is consistent with the true label result and both indicate that the sample answer text is incorrect, it means that the current large language model cannot accurately understand and answer the sample query text. Then, a negative sample is established based on the sample query text and its corresponding sample answer text (this negative sample is a very meaningful negative sample, which helps to improve the model's ability), and the negative sample is stored in the buffer area. Then, the sample query text and its corresponding set of prompt samples are re-input into the large language model to obtain the second sample answer text corresponding to the sample query text again. If the evaluation result corresponding to the second sample answer text is consistent with the true label result and both indicate that the sample answer text is correct, a positive sample corresponding to the sample query text is established based on the sample query text and the second sample answer text, and a new candidate sample set is created based on this positive sample and the negative sample established during the first evaluation and added to the current index pool. If the evaluation result corresponding to the second sample answer text is not consistent with the true label result indicating that the sample answer text is correct, the third sample answer text corresponding to the sample query text is obtained, and so on. If after N iterations, the evaluation result and the true label result are never consistent and both indicate that the sample answer text is correct for the N sample answer texts, the negative samples established corresponding to the sample query text during the N processes are saved, and the positive sample corresponding to the sample query text and its correct answer established manually can be obtained in a timely manner. Then, a new candidate sample set is created based on the positive and negative samples of the sample query text and added to the current index pool.
[0067] S306. Perform incremental training on the large language model based on the updated index pool to obtain the trained target large language model.
[0068] In the embodiment of the present application, by introducing the relevant set of prompt samples and jointly inputting them into the large language model with the sample query text, the model can learn the reward function implicitly based on the positive and negative samples in the set of prompt samples, so as to generate more accurate answers. By jointly determining the iconic positive and negative samples to be added to the index pool based on the evaluation result and the true label, the index pool can guide the large language model to generate more accurate answer texts in the future. By performing incremental training based on the updated index pool instead of training the entire model from scratch, the computing resources and time costs can be greatly reduced, which helps to improve the quality of the model more efficiently.
[0069] Figure 4 is a schematic diagram of an exemplary embodiment of a training method for a large language model shown in the present application. As Figure 4 shown, the training method for the large language model includes the following steps:
[0070] S401. Determine a plurality of historical query texts.
[0071] S402. For any historical query text, obtain the first thought chain between the historical query text and its corresponding correct answer text, and obtain the second thought chain between the historical query text and its corresponding wrong answer text.
[0072] Among them, each thought chain can be represented as a trajectory sequence s1, a1, s2, a2, …, and this trajectory follows the Markov assumption, that is, the future state only depends on the current state and is independent of the past state. The specific construction steps are as follows:
[0073] s1 refers to the problem description: This state represents the input historical query text and is the starting point of the entire reasoning process.
[0074] a1 refers to Plan 1: This is the first reasoning or planning step obtained based on s1, usually manifested as a preliminary analysis of the problem, providing guidance for the subsequent answer generation.
[0075] s2 refers to the answer generated by the problem description + Plan 1 + LLM according to calculation 1. In this state, the historical query text is combined with the first reasoning step, and after being inferred by the large language model, a preliminary answer is generated.
[0076] a2 refers to Plan 2, which is used for subsequent planning and correction. This is the next reasoning or correction step generated based on s2, and may involve correcting the preliminary answer or further reasoning, aiming to optimize the model's response, and so on.
[0077] S403. Based on the first thought chain and the second thought chain, establish positive samples and negative samples corresponding to the historical query text respectively, and establish a candidate sample set corresponding to the historical query text based on the positive samples and negative samples.
[0078] Among them, both the positive samples and the negative samples include the entire thought chain process and the positive sample label or the negative sample label.
[0079] S404. Based on the candidate sample sets corresponding to multiple historical query texts respectively, establish a preset index pool.
[0080] At this time, the number of candidate sample sets generally included in the preset index pool is not too large, and during the subsequent incremental training process, the index pool will also be incrementally expanded.
[0081] S405. Determine the sample query text, and match at least one set of prompt samples related to the sample query text from the preset index pool, where the index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0082] S406, input the prompt sample set and the sample query text into the large language model to be trained to obtain the sample response text.
[0083] S407, match at least one evaluation sample set related to the sample response text from the index pool.
[0084] S408, evaluate the accuracy of the sample response text based on the evaluation sample set to obtain the evaluation result.
[0085] S409, perform accuracy annotation on the sample response text to obtain the true label.
[0086] S410, update the index pool by combining the evaluation result and the true label.
[0087] S411, perform incremental training on the large language model based on the updated index pool to obtain the trained target large language model.
[0088] Regarding the specific implementation manners of steps S405 - S411, reference can be made to the specific introduction of the relevant parts in the above embodiments, and details will not be elaborated here.
[0089] The embodiments of this application introduce the pre - construction process of the index pool. Using the chain of thought enables more information to be included in the candidate sample set; by introducing the relevant prompt sample set and the sample query text and inputting them into the large language model together, the model can learn the reward function implicitly based on the positive and negative samples in the prompt sample set, thereby generating more accurate answers; by jointly determining the iconic positive and negative samples based on the evaluation result and the true label and adding them to the index pool, the index pool can guide the large language model to generate more accurate response texts in the future; by performing incremental training based on the updated index pool instead of training the entire model from scratch, it can greatly reduce the computing resources and time costs, and contribute to more efficiently improving the quality of the model.
[0090] Figure 5 is a schematic diagram of an exemplary implementation manner of a text query method shown in this application. As Figure 5 shown, this text query method includes the following steps:
[0091] S501, obtain the target query text.
[0092] Among them, the target query text is generally the text actually input by online users.
[0093] S502, match at least one target prompt sample set related to the target query text from the target index pool, where the target index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0094] In some embodiments, multiple initial sample sets related to the target query text are matched from the target index pool based on the term frequency-inverse document frequency algorithm (BM25 algorithm), and then, based on the vector similarity between the target query text and each initial sample set, a target prompt sample set related to the target query text is screened out from the initial sample sets.
[0095] S503, Input the target prompt sample set and the target query text into the target large language model together to obtain a to-be-determined answer text.
[0096] Among them, the target large language model can be trained and generated by the training method introduced in the above embodiments.
[0097] S504, Determine the target answer text corresponding to the target query text according to the to-be-determined answer text.
[0098] As an implementable way, after obtaining the to-be-determined answer text, perform adaptive editing and modification on the to-be-determined answer text to obtain the final target answer text.
[0099] As another implementable way, the to-be-determined answer text can be directly used as the target answer text, which can reduce the calculation amount.
[0100] The embodiment of the present application proposes a text query method, including: obtaining a target query text; matching at least one target prompt sample set related to the target query text from the target index pool, where the target index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples; inputting the target prompt sample set and the target query text into the target large language model together to obtain a to-be-determined answer text; determining the target answer text corresponding to the target query text according to the to-be-determined answer text. By introducing the relevant target prompt sample set and the target query text and inputting them into the trained target large language model together, the target large language model can help the model implicitly learn the reward function based on the positive samples and negative samples in the target prompt sample set, so as to generate a more accurate answer text.
[0101] Figure 6 It is a schematic diagram of an exemplary embodiment of a text query method shown in the present application, as Figure 6 shown, the text query method includes the following steps:
[0102] S601, Obtain a target query text.
[0103] S602, Match at least one target prompt sample set related to the target query text from the target index pool, where the target index pool includes multiple candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0104] S603. Input the target prompt sample set and the target query text into the target large language model to obtain a to-be-determined response text.
[0105] For steps S601 - S603, reference can be made to the specific introduction of the relevant parts in the above embodiments, and details will not be elaborated here.
[0106] S604. Evaluate the accuracy of the to-be-determined response text to obtain the evaluation result of the to-be-determined response text.
[0107] As an implementable way, input the to-be-determined response text and the target query text into a pre-trained accuracy evaluation model to obtain the evaluation result of the to-be-determined response text.
[0108] As another implementable way, match at least one target evaluation sample set related to the to-be-determined response text from the index pool; then evaluate the accuracy of the to-be-determined response text based on the target evaluation sample set to obtain the evaluation result of the to-be-determined response text.
[0109] S605. In response to the evaluation result being passed, use the to-be-determined response text as the target response text.
[0110] S606. In response to the evaluation result being not passed, repeat the steps of jointly inputting the target prompt sample set and the target query text into the target large language model and subsequent steps until a target response text that passes the evaluation is obtained.
[0111] In this application, by jointly inputting the relevant target prompt sample set and the target query text into the trained target large language model, the target large language model can learn the reward function implicitly based on the positive and negative samples in the target prompt sample set, so as to generate a more accurate response text; through the evaluation of the generated to-be-determined response text, the target response text presented to the user finally is a response text with high accuracy, improving user satisfaction.
[0112] Figure 7 is a schematic diagram of a training device for a large language model according to an exemplary embodiment of the present disclosure, as Figure 7 shown, the training device 700 of the large language model includes a matching module 701, a determination module 702, an update module 703, and a generation module 704, where:
[0113] The matching module 701 is configured to determine a sample query text, and match at least one prompt sample set related to the sample query text from a preset index pool, where the index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0114] A determination module 702, configured to jointly input the prompt sample set and the sample query text into a large language model to be trained, and obtain a sample answer text.
[0115] An update module 703, configured to obtain parameters related to the accuracy of the sample answer text, and update the index pool according to the parameters related to the accuracy.
[0116] A generation module 704, configured to perform incremental training on the large language model based on the updated index pool to obtain a trained target large language model.
[0117] By jointly inputting the relevant prompt sample set and the sample query text into the large language model, the device enables the model to learn the reward function implicitly based on the positive and negative samples in the prompt sample set, so as to generate more accurate answers; by performing incremental training based on the updated index pool instead of training the entire model from scratch, it can greatly reduce the computing resources and time costs, and help improve the quality of the model more efficiently.
[0118] Further, the update module 703 is further configured to: perform accuracy evaluation on the sample answer text to obtain an evaluation result; perform accuracy annotation on the sample answer text to obtain a true label; and update the index pool in combination with the evaluation result and the true label.
[0119] Further, the update module 703 is further configured to: match at least one evaluation sample set related to the sample answer text from the index pool; and perform accuracy evaluation on the sample answer text based on the evaluation sample set to obtain the evaluation result.
[0120] Further, the matching module 701 is further configured to: match multiple first initial sample sets related to the sample query text from the index pool based on the term frequency-inverse document frequency algorithm; calculate the first vector similarity between the sample query text and each first initial sample set, and screen out the prompt sample set related to the sample query text from the first initial sample sets based on the first vector similarity.
[0121] Further, the update module 703 is further configured to: match multiple second initial sample sets related to the sample answer text from the index pool based on the term frequency-inverse document frequency algorithm; calculate the second vector similarity between the sample answer text and each second initial sample set, and screen out the evaluation sample set related to the sample answer text from the second initial sample sets based on the second vector similarity.
[0122] Further, the matching module 701 and the updating module 703 are further configured to: for any one of the sample query text and the sample answer text, convert the any one of the sample texts into a first text vector; convert the initial sample set corresponding to the any one of the sample texts into a second text vector respectively; and calculate the vector similarity between the first text vector and each of the second text vectors.
[0123] Further, the training device 700 of the large language model further includes a building module, and the building module is configured to: determine a plurality of historical query texts; for any one of the historical query texts, obtain a first thought chain between the historical query text and its corresponding correct answer text, and obtain a second thought chain between the historical query text and its corresponding wrong answer text; establish a positive sample and a negative sample corresponding to the historical query text based on the first thought chain and the second thought chain respectively, and establish a candidate sample set corresponding to the historical query text based on the positive sample and the negative sample; and establish the preset index pool based on the candidate sample sets corresponding to the plurality of historical query texts respectively.
[0124] Further, the updating module 703 is further configured to: input the sample answer text and the evaluation sample set into the large language model together to obtain the evaluation result output by the large language model.
[0125] Figure 8 is a schematic diagram of a text query device according to an exemplary embodiment of the present disclosure, as Figure 8 shown, the text query device 800 includes an obtaining module 801, a matching module 802, a generating module 803, and a determining module 804, wherein:
[0126] The obtaining module 801 is configured to obtain a target query text.
[0127] The matching module 802 is configured to match at least one target prompt sample set related to the target query text from the target index pool, wherein the target index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples.
[0128] The generating module 803 is configured to input the target prompt sample set and the target query text into a target large language model together to obtain a pending answer text.
[0129] The determining module 804 is configured to determine a target answer text corresponding to the target query text according to the pending answer text.
[0130] This device inputs a relevant set of target prompt samples and a target query text into a trained target large language model, enabling the target large language model to learn a reward function implicitly based on the positive and negative samples in the target prompt sample set, thereby generating more accurate response text.
[0131] Further, the determination module 804 is further configured to: evaluate the accuracy of the to-be-determined response text to obtain an evaluation result of the to-be-determined response text; in response to the evaluation result being a passed evaluation, use the to-be-determined response text as the target response text; in response to the evaluation result being a failed evaluation, repeatedly execute the steps of jointly inputting the target prompt sample set and the target query text into the target large language model and subsequent steps until the target response text that passes the evaluation is obtained.
[0132] Further, the determination module 804 is further configured to: use the to-be-determined response text as the target response text.
[0133] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0134] Figure 9 A schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0135] As Figure 9 shown, the device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0136] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as a keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as a disk, optical disc, etc.; and communication unit 909, such as a network card, modem, wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0137] Computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 901 executes the various methods and processes described above, such as the training method of a large language model and the text query method. For example, in some embodiments, the training method of a large language model and the text query method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the training method of the large language model and the text query method described above can be executed. Alternatively, in other embodiments, computing unit 901 can be configured to execute the training method of the large language model and the text query method in any other suitable manner (e.g., by means of firmware).
[0138] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0139] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.
[0140] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0142] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0143] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain. It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0144] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for training a large language model, comprising: Determine a sample query text, and obtain at least one prompt sample set related to the sample query text by matching from a preset index pool, wherein the index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; Inputting the prompt sample set and the sample query text into the large language model to be trained to obtain a sample answer text; Acquiring accuracy-related parameters of the sample answer text, and updating the index pool according to the accuracy-related parameters; Incrementally train the large language model based on the updated index pool to obtain a trained target large language model.
2. The method according to claim 1, wherein: The obtaining of accuracy-related parameters of the sample answer text and updating the index pool according to the accuracy-related parameters includes: Performing accuracy evaluation on the sample answer text to obtain an evaluation result; Annotating the sample answer text for accuracy to obtain a true label; The index pool is updated in combination with the evaluation result and the true label.
3. The method according to claim 2, wherein: The accuracy evaluation of the sample answer text is performed to obtain an evaluation result, including: Matching and obtaining at least one evaluation sample set related to the sample answer text from the index pool; The sample answer text is evaluated for accuracy based on the evaluation sample set to obtain the evaluation result.
4. The method according to claim 1, wherein: The matching and obtaining at least one prompt sample set related to the sample query text from a preset index pool includes: Matching and obtaining a plurality of first initial sample sets related to the sample query text from the index pool based on a term frequency inverse document frequency algorithm; A first vector similarity between the sample query text and each of the first initial sample sets is calculated, and a prompt sample set related to the sample query text is screened out from the first initial sample set based on the first vector similarity.
5. The method according to claim 3, wherein: The matching and obtaining at least one evaluation sample set related to the sample answer text from the index pool includes: Matching and obtaining a plurality of second initial sample sets related to the sample answer text from the index pool based on a word frequency inverse document frequency algorithm; A second vector similarity between the sample answer text and each of the second initial sample sets is calculated, and an evaluation sample set related to the sample answer text is screened out from the second initial sample set based on the second vector similarity.
6. The method according to claim 4 or 5, wherein: The calculating of the first vector similarity between the sample query text and each of the first initial sample sets, or the calculating of the second vector similarity between the sample answer text and each of the second initial sample sets, comprises: For any sample text of the sample query text and the sample answer text, convert the any sample text into a first text vector; Converting the initial sample sets corresponding to any of the sample texts into second text vectors respectively; The vector similarity between the first text vector and each of the second text vectors is calculated and obtained.
7. The method according to claim 6, wherein: The method for establishing the preset index pool includes: Determine multiple historical query texts; For any of the historical query texts, obtaining a first thought chain between the historical query text and a corresponding correct answer text, and obtaining a second thought chain between the historical query text and a corresponding incorrect answer text; Based on the first thought chain and the second thought chain, respectively, a positive sample and a negative sample corresponding to the historical query text are established, and based on the positive sample and the negative sample, a candidate sample set corresponding to the historical query text is established; The preset index pool is established based on the candidate sample sets respectively corresponding to the multiple historical query texts.
8. The method according to claim 3, wherein: The step of performing accuracy evaluation on the sample answer text based on the evaluation sample set to obtain the evaluation result includes: The sample answer text and the evaluation sample set are input into the large language model together to obtain the evaluation result output by the large language model.
9. A text query method, comprising: Get the target query text; Matching and obtaining at least one target prompt sample set related to the target query text from a target index pool, wherein the target index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; Inputting the target prompt sample set and the target query text into a target large language model to obtain a pending answer text; The target answer text corresponding to the target query text is determined according to the pending answer text.
10. The method according to claim 9, wherein: The step of determining the target answer text corresponding to the target query text according to the pending answer text includes: Performing accuracy evaluation on the pending answer text to obtain an evaluation result of the pending answer text; In response to the evaluation result being that the evaluation is passed, taking the pending answer text as the target answer text; In response to the evaluation result being that the evaluation fails, the target prompt sample set and the target query text are input into the target large language model and subsequent steps are repeatedly executed until the target answer text that passes the evaluation is obtained.
11. The method according to claim 9, wherein: The step of determining the target answer text corresponding to the target query text according to the pending answer text includes: The pending answer text is used as the target answer text.
12. A large language model training device, comprising: A matching module, used to determine a sample query text, and to match and obtain at least one prompt sample set related to the sample query text from a preset index pool, wherein the index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; A determination module, used for inputting the prompt sample set and the sample query text into a large language model to be trained to obtain a sample answer text; An updating module, used for obtaining accuracy-related parameters of the sample answer text, and updating the index pool according to the accuracy-related parameters; A generation module is used to perform incremental training on the large language model based on the updated index pool to obtain a trained target large language model.
13. The device according to claim 12, wherein: The update module is further used for: Performing accuracy evaluation on the sample answer text to obtain an evaluation result; Annotating the sample answer text for accuracy to obtain a true label; The index pool is updated in combination with the evaluation result and the true label.
14. The device according to claim 13, wherein: The update module is further used for: Matching and obtaining at least one evaluation sample set related to the sample answer text from the index pool; The sample answer text is evaluated for accuracy based on the evaluation sample set to obtain the evaluation result.
15. The device according to claim 12, wherein: The matching module is further used for: Matching and obtaining a plurality of first initial sample sets related to the sample query text from the index pool based on a term frequency inverse document frequency algorithm; A first vector similarity between the sample query text and each of the first initial sample sets is calculated, and a prompt sample set related to the sample query text is screened out from the first initial sample set based on the first vector similarity.
16. The device according to claim 14, wherein: Update module, also used to: Matching and obtaining a plurality of second initial sample sets related to the sample answer text from the index pool based on a word frequency inverse document frequency algorithm; A second vector similarity between the sample answer text and each of the second initial sample sets is calculated, and an evaluation sample set related to the sample answer text is screened out from the second initial sample set based on the second vector similarity.
17. The device according to claim 15 or 16, wherein: The matching module and the updating module are further used for: For any sample text of the sample query text and the sample answer text, convert the any sample text into a first text vector; Converting the initial sample sets corresponding to any of the sample texts into second text vectors respectively; The vector similarity between the first text vector and each of the second text vectors is calculated and obtained.
18. The device according to claim 17, wherein: The device also includes a building module, the building module is used to: Determine multiple historical query texts; For any of the historical query texts, obtaining a first thought chain between the historical query text and a corresponding correct answer text, and obtaining a second thought chain between the historical query text and a corresponding incorrect answer text; Based on the first thought chain and the second thought chain, respectively, a positive sample and a negative sample corresponding to the historical query text are established, and based on the positive sample and the negative sample, a candidate sample set corresponding to the historical query text is established; The preset index pool is established based on the candidate sample sets respectively corresponding to the multiple historical query texts.
19. The device according to claim 14, wherein: The update module is further used for: The sample answer text and the evaluation sample set are input into the large language model together to obtain the evaluation result output by the large language model.
20. A text query device, comprising: An acquisition module, used to acquire a target query text; A matching module, configured to match and obtain at least one target prompt sample set related to the target query text from a target index pool, wherein the target index pool includes a plurality of candidate sample sets, and the candidate sample sets include positive samples and negative samples; A generating module, used for inputting the target prompt sample set and the target query text into a target large language model to obtain a pending answer text; A determination module is used to determine a target answer text corresponding to the target query text according to the pending answer text.
21. The device according to claim 20, wherein: The determining module is further used for: Performing accuracy evaluation on the pending answer text to obtain an evaluation result of the pending answer text; In response to the evaluation result being that the evaluation is passed, taking the pending answer text as the target answer text; In response to the evaluation result being that the evaluation fails, the target prompt sample set and the target query text are input into the target large language model and subsequent steps are repeatedly executed until the target answer text that passes the evaluation is obtained.
22. The device according to claim 20, wherein: The determining module is further used for: The pending answer text is used as the target answer text.
23. An electronic device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8 or 9-11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8 or 9-11.
25. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1-8 or 9-11.