Model training method, device, equipment, medium and product
By using question-and-answer pairs in the language model for initial training and using the generated candidate reply text to construct partial order pairs for further training, the problem of insufficient inference performance in the professional technical field is solved, and the reply quality is significantly improved.
Patent Information
- Application Number
- CN202510220781.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
In the knowledge areas related to expertise, language models are difficult to effectively improve their inference performance, especially in generating high-quality replies.
By obtaining question-and-answer pairs related to the target knowledge field, training the initial language model, obtaining the first language model, and then using the first language model to generate candidate reply text for different problem texts, filtering out positive and negative example samples, constructing partial order pairs, performing the second model training, and obtaining the second language model.
This method significantly improves the language model's reply ability in the target knowledge field, and the generated reply quality is higher, avoiding the generation of low-quality reply.
Smart Images

Figure CN120144781A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of computer technology, language models have emerged. Language models usually have natural language processing capabilities and can handle different types of natural language tasks. To improve the inference performance of language models, the industry usually uses retrieval-augmented generation (RAG) technology to assist language models in task processing.
[0003] Specifically, the RAG technology is divided into two stages: a data retrieval stage and a response generation stage. Among them, the data retrieval stage is used to retrieve knowledge data related to the input text from the knowledge base, and the response generation stage is used to generate a response text by combining the knowledge data.
[0004] After applying the RAG technology, in general fields unrelated to professional technologies, language models can refer to the retrieved data to generate more reasonable and accurate response texts. However, in knowledge fields related to professional technologies, it is difficult to effectively improve the inference performance of language models. Summary of the Invention
[0005] This application provides a model training method. This method can improve the response ability of a language model in a target knowledge field. This application also provides an apparatus, an electronic device, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0006] In a first aspect, this application provides a model training method, which includes:
[0007] Obtain at least one question-and-answer pair related to a target knowledge field; wherein each of the question-and-answer pairs consists of a first question text and a first response text;
[0008] Use the at least one first question-and-answer pair to train an initial language model to obtain a first language model;
[0009] Obtain at least one second question text related to the target knowledge field; wherein the second question text is different from the first question text;
[0010] For each of the at least one second question text, perform the following steps: Use the first language model to generate responses to the second question text, obtaining multiple candidate response texts for the second question text; From the multiple candidate response texts, determine a second response text whose response quality represents a positive example sample and a third response text whose response quality represents a negative example sample; Based on the second question text, the second response text, and the third response text, construct a partial order pair for the second question text.
[0011] Use the partial order pairs respectively corresponding to the at least one second question text to train the first language model, obtaining a second language model.
[0012] In a second aspect, the present application provides a model training device, which includes:
[0013] A first acquisition module, configured to acquire at least one question-and-answer pair related to a target knowledge domain; wherein, each of the question-and-answer pairs consists of a first question text and a first response text.
[0014] A first training module, configured to use the at least one first question-and-answer pair to train an initial language model, obtaining a first language model.
[0015] A second acquisition module, configured to acquire at least one second question text related to the target knowledge domain; wherein, the second question text is different from the first question text.
[0016] A construction module, configured to perform the following steps for each of the at least one second question text: Use the first language model to generate responses to the second question text, obtaining multiple candidate response texts for the second question text; From the multiple candidate response texts, determine a second response text whose response quality represents a positive example sample and a third response text whose response quality represents a negative example sample; Based on the second question text, the second response text, and the third response text, construct a partial order pair for the second question text.
[0017] A second training module, configured to use the partial order pairs respectively corresponding to the at least one second question text to train the first language model, obtaining a second language model.
[0018] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory, so that the electronic device executes the model training method as described in the first aspect or any implementation manner of the first aspect.
[0019] Fourthly, the present application provides a computer-readable storage medium storing instructions for instructing an electronic device to execute the model training method according to the first aspect or any implementation manner of the first aspect as described above.
[0020] Fifthly, the present application provides a computer program product containing instructions, which when running on an electronic device, causes the electronic device to execute the model training method according to the first aspect or any implementation manner of the first aspect as described above.
[0021] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners.
[0022] As can be seen from the above technical solutions, the present application has the following advantages:
[0023] The present application provides a model training method. The method first obtains at least one question-and-answer pair related to a target knowledge domain, where each question-and-answer pair consists of a first question text and a first reply text, and uses the at least one first question-and-answer pair to train an initial language model to obtain a first language model. Then, at least one second question text related to the target knowledge domain is obtained, where the second question text is different from the first question text. For each second question text in the at least one second question text, the following steps are performed: using the first language model to reply to the second question text to obtain multiple candidate reply texts for the second question text, determining a second reply text with a reply quality representing a positive example sample and a third reply text with a reply quality representing a negative example sample from the multiple candidate reply texts, constructing a partial order pair for the second question text according to the second question text, the second reply text, and the third reply text, and using the partial order pairs corresponding to the at least one second question text respectively to train the first language model to obtain a second language model.
[0024] In this method, for the target knowledge domain, first, the first model training is performed using the question-and-answer pairs, then different question texts are re-obtained, the language model (i.e., the first language model) after the first model training is used to reply to the questions, positive example samples and negative example samples are screened out from the multiple candidate reply texts, partial order pairs are constructed, and the second model training is performed. Since the first model training is performed using the question-and-answer pairs first, the reply ability of the first language model in the target knowledge domain is initially improved, and the first language model can generate reply texts (i.e., positive example samples) with higher reply quality for the second model training. Thus, by constructing partial order pairs and using the partial order pairs for the second model training, the second language model can better generate high-quality replies that conform to the positive example samples and avoid generating low-quality replies similar to the negative example samples, thereby improving the reply ability of the language model in the target knowledge domain. Description of the Drawings
[0025] To more clearly illustrate the technical method of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below.
[0026] Figure 1 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0027] Figure 2 It is a schematic flowchart of a method for obtaining question-and-answer pairs provided by an embodiment of the present application;
[0028] Figure 3 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;
[0029] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0030] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0031] First, some technical terms and application scenarios involved in the embodiments of the present application will be introduced.
[0032] With the rapid development of computer technology, language models have emerged. Language models have powerful expressive and learning abilities and can handle various different types of natural language tasks, such as dialogue tasks, translation tasks, speech recognition tasks, text generation tasks, semantic analysis tasks, etc.
[0033] Generally, a language model can perform natural language processing on the input text of a user and generate a corresponding response text. Among them, the input text can indicate a question, an instruction, etc.
[0034] To improve the inference performance of language models, the industry usually adopts the retrieval-augmented generation (RAG) technique to assist language models in task processing. Specifically, the RAG technique is divided into two stages: the data retrieval stage and the response generation stage. In the data retrieval stage, for the input text of the user, by calling the recall model, relevant knowledge data is retrieved from the knowledge base. For example, knowledge data similar to the input text is retrieved from the knowledge base. In the response generation stage, the knowledge data obtained in the data retrieval stage and the input text are sent to the language model together, enabling the language model to analyze the input text in combination with the knowledge data and then generate a response text.
[0035] After applying the RAG technique, in general domains unrelated to professional technologies, the retrieved knowledge data related to the input text can assist the language model in response generation. Therefore, the language model can generate more reasonable and accurate response texts. However, in knowledge domains related to professional technologies, due to the strong professionalism of the knowledge data related to the input text, it is difficult for the language model to quickly and accurately understand the knowledge data, difficult to effectively improve the reasonableness and accuracy of the response text, and difficult to effectively improve the inference performance of the language model.
[0036] For example, when the knowledge domain related to the professional technology field is the mathematics field, the input text is "What is a limit", and the knowledge data related to the input text can be data in the mathematics field such as advanced mathematics textbooks. Due to the strong professionalism of the mathematics field, for the input text "What is a limit", it is difficult for the language model to give a highly accurate response text after analyzing the data in the mathematics field such as advanced mathematics textbooks.
[0037] To address the above problems, the industry usually adopts the data annotation method to improve the inference performance of language models in knowledge domains. Among them, the data annotation method refers to that researchers in a specific knowledge domain manually annotate the knowledge data in that knowledge domain. Since the researchers have professional knowledge in that knowledge domain, through manual annotation, Q&A data in that knowledge domain can be obtained, and then the language model is trained based on the data annotation results to enhance the capabilities and performance of the language model in specific knowledge domains.
[0038] However, the above data annotation method relies on manual data annotation, making it difficult to obtain high-quality data annotation results and resulting in poor training effects.
[0039] In view of this, the present application provides a model training method. The method first obtains at least one question-and-answer pair related to the target knowledge domain, where each question-and-answer pair consists of a first question text and a first answer text. Using the at least one first question-and-answer pair, an initial language model is trained to obtain a first language model. Then, at least one second question text related to the target knowledge domain is obtained, where the second question text is different from the first question text. For each second question text in the at least one second question text, the following steps are performed: using the first language model to reply to the second question text to obtain multiple candidate reply texts for the second question text, determining a second reply text that represents a positive example sample of the reply quality and a third reply text that represents a negative example sample of the reply quality from the multiple candidate reply texts, constructing a partial order pair for the second question text according to the second question text, the second reply text, and the third reply text, and using the partial order pairs corresponding to the at least one second question text respectively to train the first language model to obtain a second language model.
[0040] In this method, for the target knowledge domain, first, the first model training is performed using the question-and-answer pairs. Then, different question texts are obtained again. The language model after the first model training (i.e., the first language model) is used to reply to the questions. The positive example samples and negative example samples are screened out from the multiple candidate reply texts, and partial order pairs are constructed for the second model training. Since the first model training is performed using the question-and-answer pairs first, the reply ability of the first language model in the target knowledge domain is initially improved. The first language model can generate reply texts with relatively high reply quality (i.e., positive example samples) for the second model training. In this way, by constructing partial order pairs and using the partial order pairs for the second model training, the second language model can better generate high-quality replies that conform to the positive example samples and avoid generating low-quality replies similar to the negative example samples, thereby improving the reply ability of the language model in the target knowledge domain.
[0041] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the following will be described with reference to the accompanying drawings. Refer to Figure 1 The flow diagram of a model training method shown in the figure, the method specifically includes:
[0042] S101: Obtain at least one question-and-answer pair related to the target knowledge domain.
[0043] The knowledge domain can be understood as a concept corresponding to the general domain. The general domain is usually not related to professional technologies. For example, the general domain can include the field of common sense in life, the field of language communication, etc. The knowledge domain is usually related to professional technologies. For example, the knowledge domain can include the field of mathematics, the field of physics, etc.
[0044] In other words, for the general domain, users can understand the knowledge in the general domain without specific professional knowledge. For the knowledge domain, users need specific professional knowledge to understand the knowledge in the knowledge domain. For example, users need mathematical professional knowledge to understand the knowledge in the mathematical domain, and users need physical professional knowledge to understand the knowledge in the physical domain.
[0045] In actual business scenarios, the knowledge domain can also include domains related to specific businesses. For example, in the product sales scenario, knowledge related to product sales (such as knowledge related to product introductions, knowledge related to product sales cases, knowledge related to potential customers, etc.) can constitute the product sales knowledge domain. Another example is that in the store management scenario, knowledge related to store management (such as knowledge related to store information, knowledge related to business conditions, etc.) can constitute the store management knowledge domain.
[0046] In the embodiments of the present application, the target knowledge domain can be understood as any knowledge domain related to professional technology, and at least one question-and-answer pair related to the target knowledge domain can be understood as a question-and-answer pair under the target knowledge domain. Each question-and-answer pair consists of a first question text and a first reply text.
[0047] In some embodiments, the reply quality of the first reply text meets the high-quality reply conditions, where the high-quality reply conditions can be understood as the conditions for measuring whether the reply text belongs to a high-quality reply. That is to say, each question-and-answer pair includes a first question text and a first reply text that belongs to a high-quality reply.
[0048] In some possible implementation manners, at least one question-and-answer pair related to the target knowledge domain can be obtained from actual business. For example, when the target knowledge domain is related to a specific business, question-and-answer pairs are obtained from the operation data of the business, and reply texts that meet the high-quality reply conditions are screened out from the reply texts of the question-and-answer pairs as at least one question-and-answer pair related to the target knowledge domain. For example, the reply quality of the reply text of the question-and-answer pair is evaluated through a set reply quality evaluation algorithm to obtain the reply quality score of the reply text, and the reply text with a reply quality score greater than the reply quality threshold is determined as the first reply text of the first question text. Another example is to use a language model to determine the reply text with the best reply quality from the reply texts of the question-and-answer pairs, and the reply text with the best reply quality is determined as the first reply text of the first question text.
[0049] In other possible implementation manners, at least one question-and-answer pair related to the target knowledge domain can also be extracted from the knowledge base. Specifically, refer to Figure 2 A schematic flowchart of a process for obtaining a question-and-answer pair provided, and this process includes:
[0050] A1: Obtain first sample data from the knowledge base of the target knowledge domain.
[0051] The process of obtaining the first sample data can be understood as the process of collecting data from the knowledge base. Among them, the knowledge base generally refers to a data set related to knowledge, such as a data set related to the knowledge of different knowledge domains. In the embodiments of the present application, the knowledge base can be related to the target knowledge domain, and the knowledge base can store data related to the target knowledge domain.
[0052] Specifically, the knowledge base can be a vector knowledge base. Collect data related to the target knowledge domain from different source channels, perform chunking processing on the data related to the target knowledge domain, and perform vectorization processing on the chunked data to generate a vectorized representation corresponding to the data related to the target knowledge domain, and store the vectorized representation corresponding to the data related to the target knowledge domain to construct a knowledge base. In this way, unified storage of different types of data is realized, which is convenient for subsequent data retrieval.
[0053] In the embodiments of the present application, any number and any type of data can be selected from the knowledge base related to the target knowledge domain as the first sample data. For example, when the target knowledge domain is the product sales knowledge domain, the sample data related to the target knowledge domain can include product introduction data, customer case data of multiple enterprises, introduction data of multiple enterprises, etc.
[0054] The first sample data can be knowledge related to the professional technology of the target knowledge domain. For example, when the target knowledge domain is the mathematics domain, the first sample data can be mathematical knowledge; when the target knowledge domain is the product sales knowledge domain, the first sample data can be knowledge related to product sales. In the embodiments of the present application, the knowledge related to the professional technology of the target knowledge domain can be used as a sample for extracting problem texts, so it is called sample data. In addition, the first sample data can include various types of data. For example, the first sample data can be data of document type, data of web page type, etc.
[0055] A2: Determine at least one first problem text from the first sample data.
[0056] In the embodiments of the present application, for each piece of first sample data, determine at least one first problem text related to the target knowledge domain. Among them, at least one first problem text is associated with the first sample data, and the first sample data can be used to answer at least one first problem text. In other words, at least one first problem text can be understood as the problem text extracted based on the content described in the first sample data.
[0057] For example, if the target knowledge domain is the product sales knowledge domain, and the first sample data A is the customer case data of Company A. The content described in the first sample data is that "after Company A uses Product A, its work efficiency is improved in many aspects". At this time, the first question text can be "How does Product A improve the efficiency of the enterprise in recruitment?", "How does Product A help the enterprise improve the online collaboration efficiency?", etc. The first sample data A can answer the above first question text. Another example, if the target knowledge domain is the product sales knowledge domain, and the first sample data B is the introduction data of Product A. The content described in the first sample data is "multiple functions of Product A". At this time, the first question text can be "How does Product A optimize the recruitment process?", "How does Product A integrate with functional module A to facilitate recruitment?", "What kind of working method is functional module B of Product A?", etc. The first sample data B can answer the above first question text.
[0058] Since the first sample data is related to the target knowledge domain, based on the content described in the first sample data, the first question text extracted can also be a first question text related to the target knowledge domain. That is to say, at least one first question text is a question involved in the target knowledge domain, which can be understood as the questions that users may have and the questions they may ask in the target knowledge domain.
[0059] The embodiments of the present application do not limit the method for determining the first question text from the first sample data. For example, the user (such as an annotator) can analyze the first sample data to extract the first question text in the first sample data. Another example is that a language model can also be used to analyze the first sample data to automatically extract the first question text in the first sample data.
[0060] A3: For at least one first question text, perform the following steps: determine the reply text of the first question text; optimize the reply text of the first question text to obtain the first reply text corresponding to the first question text; construct a question-and-answer pair according to the first question text and the first reply text corresponding to the first question text.
[0061] After extracting the first question text from the first sample data, for each first question text, the corresponding reply text can be determined. The embodiments of the present application do not limit the method for determining the reply text corresponding to each first question text. For example, different users (such as annotators) can respectively reply to each first question text multiple times to determine the reply text corresponding to each first question text. Another example is that a language model can also be used to reply to each first question text multiple times to determine the reply text corresponding to each first question text.
[0062] Next, considering that the response text for each first question text may not be accurate enough and the response quality does not meet the conditions for high-quality responses, for each first question text, a general language model is used to optimize the response text of the first question text to improve the response quality of the response text, obtaining the first response text corresponding to the first question text, and then constructing a question-and-answer pair.
[0063] Among them, the language model can have natural language processing capabilities, be able to understand the meaning of natural language, and process different types of natural language tasks. For example, the language model can be a deep learning model trained using text data. The general language model can be understood as a language model that does not use data in the knowledge domain to enhance the model's performance. In other words, the general language model has similar performance in model reasoning in various knowledge domains and does not have outstanding reasoning performance for a specific knowledge domain.
[0064] The language model (such as a general language model) can optimize the response text of the first question text based on the way of prompt learning. Among them, the prompt (prompt) can be used to guide the language model to perform specific outputs in generative tasks (such as text generation tasks, question-and-answer tasks, dialogue tasks). By configuring the prompt, it helps the language model understand the background and requirements of the task, enabling the language model to process different types of natural language processing tasks without retraining the language model, increasing the scalability and flexibility of the language model.
[0065] Specifically, generate a first prompt, send the first prompt to the general language model, and receive the first response text corresponding to the first question text returned by the general language model.
[0066] Among them, the first prompt includes: the first question text, the response text of the first question text, and information indicating to optimize the response text of the first question text. Since the first prompt includes the above information, the general language model can optimize the response text of the first question text based on the prompting ability of the first prompt, combined with the first question text. In this way, the response quality is improved, so that the response quality of the optimized response text (i.e., the first response text) meets the conditions for high-quality responses.
[0067] It should be noted that the embodiments of the present application do not limit the number of times of optimizing the response text of the first question text. For example, after one optimization of the response text of the first question text, if the optimized response text still does not meet the conditions for high-quality responses, it can be further optimized for the optimized response text. In this case, the first prompt input to the general language model can include the first question text, the optimized response text of the first question text, and information indicating to optimize the optimized response text of the first question text again.
[0068] S102: Train an initial language model using at least one first question-and-answer pair to obtain a first language model.
[0069] Since it is necessary to improve the reasoning ability of the language model in the target knowledge domain, at least one question-and-answer pair related to the target knowledge domain is used as training data to train the initial language model, and the trained initial language model is called the first language model.
[0070] Among them, the initial language model can be understood as the language model to be trained. In some possible implementation manners, the initial language model is trained using the low-rank adaptation (LoRA) technique. Among them, the LoRA technique injects a trainable low-rank decomposition matrix into each layer of the transformer architecture of the initial language model, fine-tunes some parameters in the initial language model, and while reducing the number of adjusted parameters, maintains the initial performance of the initial language model to achieve efficient model training.
[0071] In this way, by using multiple "first question text - first reply text" question-and-answer pairs in the target knowledge domain as training data, the reasoning ability of the trained first language model in the target knowledge domain is initially improved, so that the output of the first language model in the target knowledge domain can be close to the high-quality first reply text, and the reply quality of the first language model is improved to a certain extent.
[0072] S103: Obtain at least one second question text related to the target knowledge domain.
[0073] After the first model training is completed, use the first language model with a certain reasoning ability in the target knowledge domain to generate training data for the second model training, so as to use the training data generated by the first language model to perform the second model training on the first language model to further improve the reasoning ability in the target knowledge domain and enhance the accuracy of reply generation.
[0074] In the embodiments of the present application, the second question text may be different from the first question text. That is to say, in the first model training process, the question-and-answer pair is determined through the first question text, and the initial language model is trained using the question-and-answer pair as training data to obtain the first language model. In the second model training process, the partial order pair is determined through the second question text, and the first language model is trained using the partial order pair as training data to obtain the second language model.
[0075] In some possible implementations, the first question text in at least one question-and-answer pair is determined based on first sample data. In this case, second sample data different from the first sample data is obtained from the knowledge base of the target knowledge domain, and at least one second question text is determined from the second sample data.
[0076] In other words, during two model training processes, different sample data (i.e., first sample data and second sample data) is obtained from the knowledge base of the target knowledge domain. Furthermore, question extraction is performed from the different sample data to obtain different question texts (i.e., first question text and second question text).
[0077] For example, the first sample data is obtained from the first half of the knowledge base of the target knowledge domain, and the second sample data is obtained from the second half of the knowledge base of the target knowledge domain, ensuring that the first sample data is different from the second sample data.
[0078] In some embodiments, a general language model is used to extract question texts from the second sample data. Specifically, a second prompt is generated, the second prompt is sent to the general language model, and at least one second question text returned by the general language model is received.
[0079] Among them, the second prompt may include: the second sample data and information for indicating generating question texts based on the second sample data. Since the second prompt includes the above information, the general language model can extract questions from the second sample data based on the prompting ability of the second prompt and output the second question text. In this way, a large number of question texts related to the target knowledge domain during the second model training process are collected.
[0080] In this way, at least one second question text related to the target knowledge domain during the second model training process is determined, so as to generate training data during the second model training process based on at least one second question text.
[0081] It should be noted that other content related to S103 is similar to the content in parts A1 and A2 in the previous text, and will not be elaborated here.
[0082] S104: For each second question text in at least one second question text, perform the following steps: use the first language model to reply to the second question text to obtain multiple candidate reply texts for the second question text; determine a second reply text whose reply quality represents a positive example sample and a third reply text whose reply quality represents a negative example sample from the multiple candidate reply texts; construct a partial order pair for the second question text according to the second question text, the second reply text, and the third reply text.
[0083] After determining at least one second question text, for each second question text, multiple candidate response texts corresponding thereto can be determined by using a first language model that has undergone one model training.
[0084] In some possible implementation manners, the first language model is directly used to reply to each second question text. Specifically, when implementing, a third prompt word is generated, and the third prompt word is sent to the first language model multiple times, and the candidate response texts returned by the first language model are respectively received to determine multiple candidate response texts for the second question text.
[0085] Among them, the third prompt word may include: the second question text and information for indicating a reply to the second question text. Since the third prompt word includes the above information, the first language model can reply to the second question text based on the prompting ability of the third prompt word to generate a candidate response text for the second question text. Since the temperature parameter of the first language model is usually not set to 0, when the same third prompt word is sent to the first language model multiple times, the first language model can generate different candidate response texts, and thus multiple candidate response texts for the second question text can be determined.
[0086] In some other possible implementation manners, considering that it is difficult to obtain accurate and comprehensive candidate response texts when the first language model directly replies to the second question text, in the embodiments of the present application, candidate response texts for each second question text are generated by recalling reference data and combining rich reference data.
[0087] Among them, the reference data can be understood as data related to a certain second question text and used to assist in generating candidate response texts for the second question text, and the reference data can be data in a knowledge base. Specifically, when implementing, for each second question text among at least one second question text, the following steps are further performed: obtaining reference data related to the second question text from a knowledge base in a target knowledge domain.
[0088] For example, the similarity between the second question text and candidate data in the knowledge base is determined from a knowledge base in a target knowledge domain, and the candidate data whose similarity meets the set conditions is determined as the reference data related to the second question text. Among them, the candidate data can be understood as data stored in a knowledge base related to a target knowledge domain.
[0089] In other words, the process of obtaining reference data related to the second question text can be understood as a process of retrieving data from a knowledge base. Considering that the candidate data in the knowledge base can be stored in the form of vectorized representations, the second question text can be vectorized to obtain the vectorized representation corresponding to the second question text, and then the similarity between the vectorized representation corresponding to the second question text and the vectorized representations corresponding to the candidate data in the knowledge base can be determined, thereby determining the reference data related to the second question text.
[0090] In the embodiments of the present application, there is no limitation on the method for determining similarity. For example, cosine similarity, Euclidean distance and other measurement methods can be used to determine the similarity between the second question text and the candidate data in the knowledge base.
[0091] The set condition can be understood as the condition for determining candidate data as reference data. For example, the set condition can be that the similarity is greater than a similarity threshold, that is, the candidate data with a similarity greater than the similarity threshold is determined as the reference data related to the second question text. Another example is that the set condition can be the top-K of the similarity ranking result, that is, the candidate data is sorted in descending order of similarity, and the candidate data ranked in the top K is determined as the reference data related to the second question text, where K is an integer greater than 0.
[0092] In this case, using the first language model, based on the reference data related to the second question text, a response to the second question text is generated to obtain multiple candidate response texts for the second question text. In other words, the first language model can incorporate the reference data to generate a response to the second question text, generating multiple candidate response texts.
[0093] Specifically, multiple fourth prompt words are generated, the multiple fourth prompt words are respectively sent to the first language model, the candidate response texts returned by the first language model are received, and multiple candidate response texts for the second question text are determined.
[0094] Among them, the fourth prompt word may include: the second question text, the reference data related to the second question text, and information for instructing a response to the second question text, and the arrangement order of the reference data related to the second question text in the multiple fourth prompt words is different.
[0095] That is to say, multiple candidate response texts can be generated by the first language model based on different prompt words. Since the reference data can include multiple pieces of data, by adjusting the arrangement order of the reference data, multiple fourth prompt words can be generated. For example, the reference data includes data A, data B, and data C. The arrangement order of the reference data in the fourth prompt word A can be data A, data B, and data C. The arrangement order of the reference data in the fourth prompt word B can be data A, data C, and data B. The arrangement order of the reference data in the fourth prompt word C can be data B, data A, and data C. The arrangement order of the reference data in the fourth prompt word D can be data B, data C, and data A. The arrangement order of the reference data in the fourth prompt word E can be data C, data A, and data B. The arrangement order of the reference data in the fourth prompt word F can be data C, data B, and data A.
[0096] Since there are differences in the arrangement order of the reference data among multiple fourth prompt words, after sending multiple fourth prompt words to the first language model, the first language model can combine the reference data with different arrangement orders to generate candidate response texts with certain differences and obtain multiple candidate response texts.
[0097] In the embodiments of the present application, the fourth prompt word may further include one or more of the information indicating the format of the candidate response text, the information indicating the content that needs to be included in the candidate response text, and the information indicating that all the content included in the candidate response text needs to be derived from the reference data. In this way, the quality of multiple candidate response texts is improved.
[0098] The embodiments of the present application do not limit the method for adjusting the arrangement order of the reference data in the fourth prompt word. For example, the arrangement order of the reference data in the fourth prompt word can be adjusted by setting a random seed.
[0099] For each second question text, after obtaining multiple candidate response texts for the second question text, the second response text whose response quality represents a positive example sample and the third response text whose response quality represents a negative example sample can be screened from the multiple candidate response texts according to the response quality of the candidate response texts.
[0100] Among them, the second response text whose response quality represents a positive example sample can be understood as a response text that can be used as a positive example sample (i.e., the response quality represents a high-quality response) in model training. The third response text whose response quality represents a negative example sample can be understood as a response text that can be used as a negative example sample (i.e., the response quality represents a low-quality response) in model training. In other words, for each second question text, from multiple candidate response texts, the response texts with response quality at two extremes (i.e., high quality and low quality) are screened out as the second response text and the third response text.
[0101] In some embodiments, a reply quality evaluation value of a plurality of candidate reply texts is determined. The candidate reply texts whose reply quality index values meet the positive example sample conditions are determined as the second reply texts representing positive example samples of the reply quality, and the candidate reply texts whose reply quality index values meet the negative example sample conditions are determined as the third reply texts representing negative example samples of the reply quality.
[0102] Among them, the reply quality evaluation value is used to measure the reply quality of the candidate reply texts. In other words, the reply quality evaluation value can quantitatively evaluate a plurality of candidate reply texts. For example, the reply quality evaluation value can be a reply quality score.
[0103] The positive example sample conditions can be understood as the conditions for the second reply texts representing positive example samples of the reply quality, and the negative example sample conditions can be understood as the conditions for the third reply texts representing negative example samples of the reply quality. In some possible implementation manners, the positive example sample conditions can be that the reply quality evaluation value is greater than a first threshold, the reply quality evaluation value is ranked among the top N after being sorted from high to low, etc., and the negative example sample conditions can be that the reply quality evaluation value is less than a second threshold, the reply quality evaluation value is ranked among the last M after being sorted from high to low, etc.
[0104] Since the reply quality evaluation value can be used to measure the reply quality of the candidate reply texts, the candidate reply texts that meet the positive example sample conditions can, to a certain extent, represent the reply texts that the user hopes to obtain, and the candidate reply texts that meet the negative example sample conditions can, to a certain extent, represent the reply texts that the user does not hope to obtain. Therefore, the candidate reply texts that meet the positive example sample conditions can be used as positive example samples, and the candidate reply texts that meet the negative example sample conditions can be used as negative example samples.
[0105] In specific implementation, a fifth prompt word is generated and sent to the general language model, and the second reply texts representing positive example samples of the reply quality returned by the general language model are received. Among them, the fifth prompt word can include: the second question text, a plurality of candidate reply texts, the positive example sample conditions, and information for indicating to select the reply texts that meet the positive example sample conditions from the plurality of candidate reply texts.
[0106] Since the above information is included in the fifth prompt word, the general language model can evaluate the reply quality of the plurality of candidate reply texts based on the prompting ability of the fifth prompt word, determine the reply quality evaluation values of the plurality of candidate reply texts, and then combine the positive example sample conditions to determine the second reply texts that meet the positive example sample conditions from the plurality of candidate reply texts.
[0107] In some possible implementation manners, the positive example sample condition may be that the response quality evaluation value ranks first after being sorted from high to low, that is, the positive example sample condition is to select the response text with the best response quality. The information in the fifth prompt word for indicating to select the response text that meets the positive example sample condition from multiple candidate response texts may be the information for indicating to select the response text that meets the positive example sample condition based on the best of N samples (BoN) manner.
[0108] In this case, the general language model may determine the second response text that meets the positive example sample condition by pairwise comparison of multiple candidate response texts. That is, the general language model may first compare candidate response text A and candidate response text B to obtain Bo2, then compare candidate response text C and Bo2 to obtain Bo3, and so on, until finally obtaining BoN, and take BoN as the second response text.
[0109] Furthermore, pairwise comparison of candidate response texts may be implemented through the good-same-bad (GSB) evaluation index, that is, the general language model first compares the GSB of candidate response text A and candidate response text B to obtain Bo2, then compares the GSB of candidate response text C and Bo2 to obtain Bo3, and so on, until finally obtaining BoN.
[0110] After determining the second response text that represents the positive example sample of the response quality, the second response text may be used as a benchmark (i.e., the standard answer, the correct answer), and the third response text that represents the negative example sample of the response quality may be determined from multiple candidate response texts. Specifically, a sixth prompt word is generated, the sixth prompt word is sent to the general language model, and the third response text that represents the negative example sample of the response quality returned by the general language model is received.
[0111] Among them, the sixth prompt word may include: the second question text, multiple candidate response texts, the second response text that represents the positive example sample of the response quality, the negative example sample condition, and the information for indicating to select the response text that meets the negative example sample condition from multiple candidate response texts in combination with the second response text that represents the positive example sample of the response quality.
[0112] Since the sixth prompt word includes the above information, the general language model may evaluate the response quality of multiple candidate response texts based on the prompting ability of the sixth prompt word, and combine the second response text that represents the positive example sample of the response quality and the negative example sample condition to determine the third response text that meets the positive example sample condition from multiple candidate response texts.
[0113] In some possible implementation manners, the negative example sample condition may be that the reply quality evaluation value is ranked last after being sorted from high to low. That is, the negative example sample condition is to select the reply text with the worst reply quality. The information in the sixth prompt word for indicating to select the reply text that meets the negative example sample condition from multiple candidate reply texts may be the information for indicating to select the reply text that meets the negative example sample condition based on the worst of N samples (WoN) manner.
[0114] In this case, the general language model may determine the third reply text that meets the negative example sample condition by pairwise comparison of multiple candidate reply texts. That is, the general language model may first compare candidate reply text A and candidate reply text B to obtain Wo2, then compare candidate reply text C and Wo2 to obtain Wo3, and so on, until WoN is finally obtained, and WoN is used as the third reply text.
[0115] Furthermore, pairwise comparison of candidate reply texts may be implemented through the good-same-bad (GSB) evaluation index. That is, the general language model first compares the GSB of candidate reply text A and candidate reply text B to obtain Wo2, then compares the GSB of candidate reply text C and Wo2 to obtain Wo3, and so on, until WoN is finally obtained.
[0116] For each second question text, after determining the second reply text representing the positive example of the reply quality and the third reply text representing the negative example of the reply quality corresponding to the second question text, the second question text, the second reply text, and the third reply text are constructed into a partial order pair of the second question text.
[0117] In the embodiments of the present application, a partial order pair is composed of "second question text - second reply text and third reply text", indicating that when the second question text is used as the model input, the second reply text representing the positive example of the reply quality is the more preferred and user-demand-satisfying model output, and the third reply text representing the negative example of the reply quality is the non-preferred and user-demand-unsatisfying model output.
[0118] S105: Train the first language model by using the partial order pairs respectively corresponding to at least one second question text to obtain a second language model.
[0119] Use the partial order pair corresponding to each second question text as the training data in the second model training process to perform model training on the first language model, and the trained first language model is called the second language model.
[0120] In some possible implementation manners, the direct preference optimization (DPO) technology is used to perform reinforcement learning on the first language model. Among them, the DPO technology directly fine-tunes the first language model through preference data (that is, the second reply text whose reply quality represents a positive example sample and the third reply text whose reply quality represents a negative example sample), so that its output better conforms to user preferences, and trains the first language model without constructing a reward model.
[0121] In this way, by using multiple "second question text - second reply text whose reply quality meets user requirements - third reply text whose reply quality does not meet user requirements" partial order pairs in the target knowledge domain as training data, the trained second language model can tend to generate reply texts that meet user requirements in the target knowledge domain, avoid generating reply texts that are not preferred, and further improve the reply quality and reasoning ability of the second language model.
[0122] In this method, for the target knowledge domain, first, the first model training is performed using question-and-answer pairs, then different question texts are re-obtained, and the language model (that is, the first language model) after the first model training is used to reply to the questions. Positive example samples and negative example samples are screened out from multiple candidate reply texts, partial order pairs are constructed, and the second model training is performed. Since the first model training is first performed using question-and-answer pairs, the reply ability of the first language model in the target knowledge domain is initially improved, and the first language model can generate reply texts (that is, positive example samples) with higher reply quality for the second model training. In this way, by constructing partial order pairs and using the partial order pairs for the second model training, the second language model can better generate high-quality replies that conform to the positive example samples, avoid generating low-quality replies similar to the negative example samples, and improve the reply ability of the language model in the target knowledge domain.
[0123] As described above in combination with Figure 1 and Figure 2 the model training method provided by the embodiments of the present application has been introduced in detail. Next, the devices and equipment provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0124] See Figure 3 the structural schematic diagram of the model training device shown. The device 30 includes:
[0125] A first acquisition module 301, configured to acquire at least one question-and-answer pair related to the target knowledge domain; wherein each of the question-and-answer pairs is composed of a first question text and a first reply text;
[0126] A first training module 302, configured to train an initial language model using the at least one first question-and-answer pair to obtain a first language model;
[0127] A second acquisition module 303, configured to acquire at least one second question text related to the target knowledge domain; wherein, the second question text is different from the first question text;
[0128] A construction module 304, configured to perform the following steps for each second question text in the at least one second question text: use the first language model to reply to the second question text to obtain multiple candidate reply texts for the second question text; determine a second reply text whose reply quality represents a positive example sample and a third reply text whose reply quality represents a negative example sample from the multiple candidate reply texts; construct a partial order pair for the second question text according to the second question text, the second reply text, and the third reply text;
[0129] A second training module 305, configured to train the first language model by using the partial order pairs respectively corresponding to the at least one second question text to obtain a second language model.
[0130] In some possible implementation manners, the first acquisition module 301 is specifically configured to:
[0131] Acquire first sample data from the knowledge base of the target knowledge domain;
[0132] Determine at least one first question text from the first sample data;
[0133] For the at least one first question text, perform the following steps: determine a reply text for the first question text; optimize the reply text for the first question text to obtain a first reply text corresponding to the first question text; construct a question-and-answer pair according to the first question text and the first reply text corresponding to the first question text.
[0134] In some possible implementation manners, the first acquisition module 301 is specifically configured to:
[0135] Generate a first prompt; wherein, the first prompt includes: the first question text, the reply text for the first question text, and information for indicating optimizing the reply text for the first question text;
[0136] Send the first prompt to a general language model, and receive the first reply text corresponding to the first question text returned by the general language model.
[0137] In some possible implementation manners, the first question text in the at least one question-and-answer pair is determined based on first sample data; the second acquisition module 303 is specifically configured to:
[0138] Obtain second sample data different from the first sample data from the knowledge base of the target knowledge domain;
[0139] Determine at least one second question text from the second sample data.
[0140] In some possible implementation manners, the second obtaining module 303 is specifically configured to:
[0141] Generate a second prompt; wherein, the second prompt includes: the second sample data and information for indicating generating a question text based on the second sample data;
[0142] Send the second prompt to the general language model, and receive at least one second question text returned by the general language model.
[0143] In some possible implementation manners, the constructing module 304 is specifically configured to:
[0144] Generate a third prompt; wherein, the third prompt includes: the second question text and information for indicating replying to the second question text;
[0145] Send the third prompt to the first language model multiple times, respectively receive candidate reply texts returned by the first language model, and determine multiple candidate reply texts for the second question text.
[0146] In some possible implementation manners, the constructing module 304 is further configured to:
[0147] Obtain reference data related to the second question text from the knowledge base of the target knowledge domain;
[0148] The constructing module 304 is specifically configured to:
[0149] Use the first language model to reply to the second question text based on the reference data related to the second question text, and obtain multiple candidate reply texts for the second question text.
[0150] In some possible implementation manners, the constructing module 304 is specifically configured to:
[0151] Generate multiple fourth prompts; wherein, the fourth prompt includes: the second question text, reference data related to the second question text, and information for indicating replying to the second question text, and the arrangement order of the reference data related to the second question text in the multiple fourth prompts is different;
[0152] Send the multiple fourth prompt words to the first language model respectively, receive the candidate response texts returned by the first language model, and determine multiple candidate response texts for the second question text.
[0153] In some possible implementation manners, the building module 304 is specifically configured to:
[0154] Determine the response quality evaluation values of the multiple candidate response texts; wherein, the response quality evaluation values are used to measure the response quality of the candidate response texts;
[0155] Determine the candidate response texts whose response quality index values meet the positive example sample conditions as the second response texts representing positive example samples of response quality, and determine the candidate response texts whose response quality index values meet the negative example sample conditions as the third response texts representing negative example samples of response quality.
[0156] In some possible implementation manners, the building module 304 is specifically configured to:
[0157] Generate a fifth prompt word; wherein, the fifth prompt word includes: the second question text, the multiple candidate response texts, the positive example sample conditions, and information for indicating to select the response texts that meet the positive example sample conditions from the multiple candidate response texts;
[0158] Send the fifth prompt word to the general language model, and receive the second response texts representing positive example samples of response quality returned by the general language model;
[0159] Generate a sixth prompt word; wherein, the sixth prompt word includes: the second question text, the multiple candidate response texts, the second response texts representing positive example samples of response quality, the negative example sample conditions, and information for indicating to select the response texts that meet the negative example sample conditions from the multiple candidate response texts in combination with the second response texts representing positive example samples of response quality;
[0160] Send the sixth prompt word to the general language model, and receive the third response texts representing negative example samples of response quality returned by the general language model.
[0161] The model training device 30 according to the embodiment of the present application can correspondingly execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the model training device 30 are respectively for implementing Figure 1 The corresponding processes of the respective methods in the illustrated embodiments are not described herein again for the sake of brevity.
[0162] The embodiment of the present application also provides an electronic device. This electronic device is specifically used to implement as Figure 3The functions of the model training device 30 in the illustrated embodiment.
[0163] Figure 4 A structural schematic diagram of an electronic device 400 is provided, as Figure 4 shown. The electronic device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other via the bus 401.
[0164] The bus 401 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0165] The processor 402 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0166] The communication interface 403 is used for external communication. For example, the communication interface 403 can be used to communicate with a terminal.
[0167] The memory 404 can include volatile memory, such as random access memory (RAM). The memory 404 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0168] The memory 404 stores executable code, and the processor 402 executes the executable code to perform the foregoing model training method.
[0169] Specifically, in the case of implementing Figure 3 the illustrated embodiment, and Figure 3When each module or unit of the model training device 30 described in the embodiments is implemented by software, the software or program code required to execute the functions of each module / unit in Figure 3 can be partially or entirely stored in the memory 404. The processor 402 executes the program code corresponding to each unit stored in the memory 404 to execute the foregoing model training method.
[0170] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the foregoing model training method applied to the model training device 30.
[0171] The embodiments of the present application also provide a computer program product that includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, they wholly or partially generate the processes or functions described in the embodiments of the present application.
[0172] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, or data center to another website, computer, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0173] When the computer program product is executed by a computer, the computer executes any one of the foregoing model training methods. The computer program product can be a software installation package. In the case where any one of the foregoing model training methods needs to be used, the computer program product can be downloaded and executed on the computer.
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0175] The units involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.
[0176] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0177] In the context of the embodiments of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] It should be noted that the embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0179] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means an "or" relationship between the associated objects before and after. "At least one (one)" or its similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0180] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0181] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the technical field.
[0182] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model training method, characterized in that: The method comprises: Acquire at least one question-answer pair related to the target knowledge domain; wherein each of the question-answer pairs consists of a first question text and a first answer text; Using the at least one first question-answer pair, training an initial language model to obtain a first language model; Acquire at least one second question text related to the target knowledge domain; wherein the second question text is different from the first question text; For each second question text in the at least one second question text, the following steps are performed: using the first language model to reply to the second question text to obtain a plurality of candidate reply texts for the second question text; determining, from the plurality of candidate reply texts, a second reply text representing a positive sample of reply quality characterization and a third reply text representing a negative sample of reply quality characterization; constructing a partial order pair of the second question text according to the second question text, the second reply text and the third reply text; The first language model is trained using the partial order pairs respectively corresponding to the at least one second question text to obtain a second language model.
2. The method according to claim 1, characterized in that The obtaining of at least one question-answer pair related to the target knowledge domain comprises: Acquire first sample data from a knowledge base in the target knowledge domain; Determining at least one first question text from the first sample data; For the at least one first question text, the following steps are performed: determining a reply text for the first question text; optimizing the reply text for the first question text to obtain a first reply text corresponding to the first question text; and constructing a question-answer pair based on the first question text and the first reply text corresponding to the first question text.
3. The method according to claim 2, characterized in that The step of optimizing the reply text to the first question text to obtain a first reply text corresponding to the first question text includes: Generate a first prompt word; wherein the first prompt word includes: the first question text, the reply text of the first question text, and information for indicating that the reply text of the first question text is optimized; The first prompt word is sent to a universal language model, and a first reply text corresponding to the first question text returned by the universal language model is received.
4. The method according to claim 1, characterized in that: The first question text in the at least one question-answer pair is determined based on first sample data; and the obtaining of at least one second question text related to the target knowledge domain comprises: Acquire second sample data different from the first sample data from a knowledge base in the target knowledge domain; At least one second question text is determined from the second sample data.
5. The method according to claim 4, characterized in that The step of determining at least one second question text from the second sample data comprises: Generate a second prompt word; wherein the second prompt word includes: the second sample data and information for indicating that a question text is generated based on the second sample data; The second prompt word is sent to a general language model, and at least one second question text returned by the general language model is received.
6. The method according to claim 1, characterized in that The step of using the first language model to reply to the second question text to obtain a plurality of candidate reply texts for the second question text includes: Generate a third prompt word; wherein the third prompt word includes: the second question text and information for indicating a reply to the second question text; The third prompt word is sent to the first language model multiple times, and candidate reply texts returned by the first language model are received respectively to determine multiple candidate reply texts for the second question text.
7. The method according to claim 1, characterized in that For each second question text in the at least one second question text, the following steps are further performed: Acquire reference data related to the second question text from a knowledge base in the target knowledge domain; The step of using the first language model to reply to the second question text to obtain a plurality of candidate reply texts for the second question text includes: The first language model is used to respond to the second question text based on the reference data related to the second question text to obtain a plurality of candidate response texts for the second question text.
8. The method according to claim 7, characterized in that The using the first language model to reply to the second question text based on the reference data related to the second question text to obtain multiple candidate reply texts for the second question text includes: Generate a plurality of fourth prompt words; wherein the fourth prompt words include: the second question text, reference data related to the second question text, and information for indicating a reply to the second question text, and the reference data related to the second question text in the plurality of fourth prompt words are arranged in different orders; The plurality of fourth prompt words are respectively sent to the first language model, candidate reply texts returned by the first language model are received, and a plurality of candidate reply texts for the second question text are determined.
9. The method according to any one of claims 1 to 8, characterized in that: The step of determining, from the plurality of candidate reply texts, a second reply text representing a positive sample of reply quality characterization and a third reply text representing a negative sample of reply quality characterization comprises: Determining reply quality evaluation values of the multiple candidate reply texts; wherein the reply quality evaluation values are used to measure the reply quality of the candidate reply texts; The candidate reply text whose reply quality index value meets the positive sample condition is determined as the second reply text of the reply quality characterizing the positive sample, and the candidate reply text whose reply quality index value meets the negative sample condition is determined as the third reply text of the reply quality characterizing the negative sample.
10. The method according to claim 9, characterized in that The step of determining the candidate reply text whose reply quality index value satisfies the positive sample condition as the second reply text representing the reply quality positive sample, and determining the candidate reply text whose reply quality index value satisfies the negative sample condition as the third reply text representing the reply quality negative sample comprises: Generate a fifth prompt word; wherein the fifth prompt word includes: the second question text, the multiple candidate reply texts, the positive sample condition, and information for indicating to select a reply text that meets the positive sample condition from the multiple candidate reply texts; Sending the fifth prompt word to a general language model, and receiving a second reply text of a positive example representing a reply quality returned by the general language model; Generate a sixth prompt word; wherein the sixth prompt word includes: the second question text, the multiple candidate reply texts, the second reply text of the reply quality characterization positive sample, the negative sample condition, and information for indicating that a reply text that meets the negative sample condition is selected from the multiple candidate reply texts in combination with the second reply text of the reply quality characterization positive sample; The sixth prompt word is sent to the general language model, and a third reply text of a negative example representing reply quality returned by the general language model is received.
11. A model training device, characterized in that: The device comprises: A first acquisition module is used to acquire at least one question-answer pair related to the target knowledge field; wherein each question-answer pair consists of a first question text and a first reply text; A first training module, configured to train an initial language model using the at least one first question-answer pair to obtain a first language model; A second acquisition module is used to acquire at least one second question text related to the target knowledge field; wherein the second question text is different from the first question text; The construction module is used to perform the following steps for each second question text in the at least one second question text: using the first language model to reply to the second question text to obtain multiple candidate reply texts for the second question text; determining a second reply text representing a positive sample of reply quality characterization and a third reply text representing a negative sample of reply quality characterization from the multiple candidate reply texts; and constructing a partial order pair of the second question text according to the second question text, the second reply text and the third reply text; The second training module is used to train the first language model using the partial order pairs corresponding to the at least one second question text to obtain a second language model.
12. An electronic device, characterized in that: The electronic device comprises a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The method comprises instructions, wherein the instructions instruct an electronic device to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The computer program product comprises computer readable instructions for implementing the method according to any one of claims 1 to 10.
Citation Information
Cited By
Prompt generation model training method and information processing method
CN121478910A