Domain question and answer large model training and question and answer method, related equipment and program product
By using the referee model to score and screen the answers generated by the initial big model and the answers in the domain question and answer data, the problem of high manual proofreading cost in the training of the domain question and answer big model is solved, and an efficient training process and highly professional question and answer ability are achieved.
Patent Information
- Application Number
- CN202510447052.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing field Q&A big model training requires a large amount of manual proofreading instruction data, resulting in high cost and low training efficiency.
The referee model obtained by using the question and answer data marked with preference sort is scored separately for the answers generated by the initial big model and the answers in the domain question and answer data, and the answers that meet the preference requirements are selected and the domain questions are composed of the target training data, and the target training data is used to iteratively train the initial big model.
It reduces the workload of manual proofreading, reduces costs, and improves training efficiency. The obtained field Q&A big model has stronger professionalism and practicality.
Smart Images

Figure CN119961422A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and more specifically, to a domain question-answering large model training and question-answering method, related equipment and program products. Background Art
[0002] General large models are usually trained based on a wide range of public literature and network data, lacking professional knowledge and industry data accumulation, and therefore lack industry specificity and accuracy. However, users have high requirements for professional services of industry large models, and their fault tolerance is low. Once industry large models provide wrong information to the public, serious consequences may occur. Retraining and fine-tuning general large models with industry data and building a highly available question-answering system can solve the above problems to a certain extent.
[0003] The technical standards and domain text data in the industry data can only be used for pre-training of the model. What is needed to build a question-and-answer system for subsequent question-and-answering is instruction data, that is, the industry's question-and-answer data. The existing technology generally uses some artificial intelligence methods to generate some instruction data. The generated instruction data may have some defects such as factual errors. In order to ensure the quality of the instruction data, a lot of manual proofreading is required before it can be applied to the model training. Therefore, the current training process of the domain question-and-answer large model still requires a lot of manual proofreading of instruction data. Summary of the invention
[0004] In view of the above problems, this application is proposed to provide a domain question-answering large model training and question-answering method, related equipment and program products to reduce the manual proofreading of instruction data during the domain question-answering large model training process, reduce costs and improve training efficiency. The specific plan is as follows:
[0005] In the first aspect, a method for training a domain question-answering large model is provided, comprising:
[0006] Obtain an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings;
[0007] Acquire a domain knowledge base, and extract domain question-answering data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information;
[0008] The initial large model is iteratively trained using the domain question and answer data. In each round of training, the referee model scores the first answer corresponding to the domain question generated by the initial large model and the second answer corresponding to the domain question in the domain question and answer data respectively. Based on the scoring results, the answer that meets the preference requirements and the domain question are selected to form target training data. The initial large model is trained using the target training data, and the domain question and answer large model is obtained after the iterative training.
[0009] In one possible design, in another implementation of the first aspect of the embodiments of the present application, during the iterative training of the initial large model, different domain question and answer data are used in different training rounds.
[0010] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of performing any round of training on the initial large model using the domain question and answer data includes:
[0011] Sending the domain question in the domain question-answering data into the initial large model to obtain the first answer generated by the initial large model;
[0012] Scoring the first answer and the second answer respectively by the referee model, and selecting answers that meet the preference requirements and the domain question to form target training data based on the scoring results;
[0013] The target training data is used to continue training the initial large model obtained in the previous round of training to obtain the initial large model after this round of training.
[0014] In a possible design, in another implementation of the first aspect of the embodiment of the present application, the process of selecting answers that meet the preference requirements and the domain questions to form target training data based on the scoring results includes:
[0015] The answers in the scoring results are sorted according to their scores, and the top N answers in the sorting are selected together with the domain question to form the target training data, where N is an integer greater than or equal to 1.
[0016] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the domain knowledge base is an extensible domain knowledge base, and the method further includes:
[0017] When new knowledge information is detected in the extensible domain knowledge base, new domain question and answer data is extracted based on the new knowledge information, and the domain question and answer big model is updated and trained using the new domain question and answer data to obtain an updated domain question and answer big model.
[0018] In one possible design, in another implementation of the first aspect of the embodiments of the present application, the referee model is a referee large model obtained by supervised training of a general large model using question and answer data labeled with preference rankings.
[0019] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the process of extracting domain question and answer data based on the domain knowledge base includes:
[0020] For each modality of domain knowledge information in the domain knowledge base, vectorization processing is performed respectively to obtain a vector representation corresponding to the domain knowledge information and store it in a vector database;
[0021] The configured query question is vectorized, and the top K candidate domain knowledge information with the highest similarity to the query question is retrieved based on vector similarity, where K is a set positive integer;
[0022] The first K candidate domain knowledge information and the query question are added to the first prompt instruction prompt, and submitted to the general large model to instruct the general large model to generate the answer that best matches the query question, and then form domain question and answer data with the query question.
[0023] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the process of extracting domain question and answer data based on the domain knowledge base includes:
[0024] Calling a general big model to instruct the general big model to extract a number of questions and corresponding answers from the provided domain knowledge information;
[0025] The questions extracted by the general large model and the corresponding answers constitute the domain question and answer data.
[0026] In a second aspect, a question-answering method is provided, comprising:
[0027] Get domain questions;
[0028] Send the domain question to the configured domain question-answering model to obtain the answer output by the model;
[0029] Among them, the domain question and answer big model is trained using the domain question and answer big model training method described in any one of the implementation methods of the first aspect of the embodiments of the present application.
[0030] In a third aspect, a domain question answering large model training device is provided, comprising:
[0031] A model acquisition unit, used to acquire an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings;
[0032] A knowledge base processing unit, used to obtain a domain knowledge base and extract domain question and answer data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information;
[0033] A model training unit is used to iteratively train the initial large model using the domain question and answer data. In each round of training, the referee model scores the first answer corresponding to the domain question generated by the initial large model and the second answer in the domain question and answer data respectively. Based on the scoring results, answers that meet the preference requirements and the domain question are selected to form target training data. The initial large model is trained using the target training data until the domain question and answer large model is obtained after the last round of training.
[0034] In a fourth aspect, an electronic device is provided, comprising: a memory and a processor;
[0035] The memory is used to store programs;
[0036] The processor is used to execute the program to implement the various steps of the domain question and answer large model training method described in any one of the implementation methods of the first aspect of the embodiments of the present application, or to implement the various steps of the question and answer method described in the second aspect.
[0037] In a fifth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the various steps of the domain question and answer large model training method described in any one of the implementation methods of the first aspect of the embodiments of the present application are implemented, or the various steps of the question and answer method described in the second aspect are implemented.
[0038] In the sixth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the various steps of the domain question and answer big model training method described in any implementation method of the first aspect of the embodiment of the present application, or implements the various steps of the question and answer method described in the second aspect.
[0039] By means of the above technical scheme, the present application uses the initial large model with question-answering capability as the base of the domain question-answering large model, and introduces a referee model, which is trained by question-answering data marked with preference ranking, and can give preference scores to different input answers and evaluate the quality of different answers. The present application is configured with a domain knowledge base, and domain question-answering data can be extracted based on the domain knowledge information in the domain knowledge base. When the initial large model is iteratively trained, the present application does not directly use the extracted domain question-answering data for training, but the first answer corresponding to the domain question generated by the initial large model and the second answer corresponding to the domain question in the domain question-answering data are scored respectively by the referee large model, and the answers and domain questions that meet the preference requirements are selected based on the scoring results to form the target training data. Obviously, after the referee large model scores the first answer and the second answer, higher quality answers and domain questions can be selected to form high-quality target training data, and then the target training data can be used to train the initial large model, so that after one or more iterative trainings, the final domain question-answering large model can be obtained, which uses high-quality domain question-answering training data for training, better understands the semantics and norms of the domain, and provides domain professionalism and more practical question-answering capabilities. Furthermore, the present application solution can obtain high-quality target training data without the need for manual proofreading of the extracted domain question and answer data, thereby saving labor costs and improving training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0041] Figure 1 A schematic diagram of an implementation system architecture of the large model training method and question-answering method provided in the embodiments of the present application;
[0042] Figure 2 A flowchart of a method for training a domain question-answering large model provided in an embodiment of the present application;
[0043] Figure 3 A schematic diagram illustrating a method of obtaining the labeled data required for training a referee model;
[0044] Figure 4 A schematic diagram of a processing flow for extracting domain question-answering data from a domain knowledge base is illustrated;
[0045] Figure 5 A flowchart of another method for training a large model for domain question answering is shown;
[0046] Figure 6A schematic diagram of the structure of a domain question-answering large model training device provided in an embodiment of the present application;
[0047] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0049] The domain big model can also be called the industry big model. It is generally defined as a general big model retrained and fine-tuned with domain data to solve domain problems. There are two types of uses for domain data. The first type is to use domain data to continue training and fine-tuning the general model, and other methods to change the model weights; the second type is to not change the weight of the general big model, but use the in-context learning capability to inject domain knowledge through prompts, or use external databases. The former can be called training / fine-tuning a domain big model because it changes the weight of the model, while the latter can basically only be considered an application of the general big model.
[0050] The technical standards and domain text data in the domain data can only be used for pre-training of the model. To build a question-and-answer system for subsequent question-and-answer, instruction data is needed. The existing technology generally uses some artificial intelligence methods to generate some instruction data (domain question-and-answer data), but in order to ensure the factuality, a lot of manual proofreading is still required. High-quality industry instruction data is the bottleneck of model fine-tuning. This application introduces a referee model to score the extracted domain question-and-answer data and the feedback data of the initial large model respectively, and selects high-quality answers and domain questions based on the scoring results to form the target training data, so as to use the target training data to optimize the initial large model to obtain the final domain question-and-answer large model.
[0051] This application provides a training method for a domain question-answering big model and a question-answering method based on the domain question-answering big model. The method of this application can be applied to a variety of fields, such as the smart car field, the medical field, the education field, etc.
[0052] The method provided in this application can be divided into a training phase and a reasoning phase. The training phase is the phase of training the domain question-answering big model, and the reasoning phase is the process of using the trained domain question-answering big model to perform the domain question-answering task. The training phase and the reasoning phase can be deployed in the same device or in different devices.
[0053] Combination Figure 1 The system structure shown in FIG. 1 may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 A server is included as an example for explanation).
[0054] For example, the training phase can be deployed in the server 200, and the inference phase can be deployed in the terminal 100, such as a mobile phone, a tablet, a smart car, a robot, or a wearable device.
[0055] For ease of understanding, this application introduces the processes of the training phase and the reasoning phase respectively.
[0056] 1. Training Phase
[0057] Reference Figure 2 , a flowchart of a method for training a domain question-answering large model provided in an embodiment of the present application may specifically include the following steps:
[0058] Step S100, obtaining an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings.
[0059] The initial large model serves as the base for the domain question-answering large model to be trained in this application. In order to improve the training efficiency, the initial large model obtained in this embodiment is a large model with question-answering capabilities, that is, the initial large model can answer user instructions.
[0060] For the currently available general large models, some of them already have the ability to answer questions, so in this embodiment, such general large models can be directly used as the initial large models. For other general large models that do not have the ability to answer questions, in this embodiment, the collected public data sets can also be used to construct question-answering data, and the question-answering data can be used to supervise the general large models for training, so as to obtain the initial large models with question-answering capabilities.
[0061] For public data sets, they can be obtained from various channels such as the Internet or professional databases. Taking the maternal and infant industry as an example, the Maternal and Infant (MATINF) Dataset is a large-scale jointly annotated data set that can be used for classification, question-answering, and summarization in the field of maternal and infant care in Chinese. This data set collects nearly 2 million question-answer pairs from a large maternal and infant care question-answering website and is constructed through automatic and manual data cleaning. In this embodiment, the public data set can be used to perform supervised training (Supervised Fine-Tuning, SFT) on the general large model base to obtain a model with basic question-answering capabilities, which is called the initial large model, or sft-model.
[0062] Furthermore, in this embodiment, a referee model is also obtained. The referee model is trained using question and answer data marked with preference rankings. The referee model can score the answers corresponding to the input questions. The scoring results can evaluate the quality of the answers. For example, the higher the score, the higher the quality of the answer, which means it is more inclined to the user's preference for the answer.
[0063] The referee model in this embodiment can also be called a reward model, which is trained by question and answer data labeled with preference rankings. Different answers in the labeled data can be obtained through an initial large model, a general large model, a domain knowledge base, etc., and different answers are further labeled with preference rankings by artificial experts to obtain the final question and answer data labeled with preference rankings.
[0064] Reference Figure 3 , which illustrates a method for obtaining labeled data required for training a referee model.
[0065] Taking the question “What is hypersensitivity reaction” as an example, the answer to the question can be obtained through a variety of channels. Figure 3 The figure shows different answers obtained through a large model and multiple knowledge bases, which are defined as answers A, B, and C from left to right.
[0066] Experts can then annotate the answers A, B, and C with a preference ranking, for example, the ranking result is B>C>A. The annotated question and answer data can then be used to train the referee model.
[0067] This application focuses on professional question-answering systems based on specialized knowledge / data in vertical fields, and is committed to solving the needs of specific scenarios in vertical fields. Therefore, experts are required to focus on professional verification and authenticity verification when annotating data, and give partial order relationships of different answers based on the above two dimensions. Among them, professional verification requires the model to find, integrate and infer relevant professional knowledge based on the questions raised by users to provide accurate, comprehensive and in-depth answers; authenticity verification determines whether the answers output by the model contain fictitious, erroneous or inconsistent statements with reality, or whether the model can correctly identify untrue information.
[0068] After the experts annotate the answers according to the above-mentioned standards for professional verification and authenticity verification, the annotated question and answer data is obtained. After the referee model is trained with the question and answer data, the answer scores given by the referee model can meet the standards for professional verification and authenticity verification, that is, a more accurate quality score can be given.
[0069] The referee model in this embodiment can adopt a variety of neural network structures, such as using multiple types of regression models. In a possible implementation, in order to improve the performance of the referee model and reduce the time spent on training the referee model, the referee model can be trained on the base of the general large model to obtain a trained referee large model. The base capacity of the general large model can be fully utilized to improve the performance of the referee large model. At the same time, there is no need to train a completely new model architecture from scratch, which can improve the training efficiency of the referee model.
[0070] When the referee model is a large referee model, it actually uses the large model to do a regression task, that is, adding a multi-layer perceptron MLP at the output layer [CLS] position of the model, and the output value is the score given by the large referee model.
[0071] In one possible implementation, the process of training the referee model on the base of the general large model can adopt LORA or other possible training strategies.
[0072] Step S110: Acquire a domain knowledge base, and extract domain question and answer data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information.
[0073] In order to enable the domain question-answering big model to better understand the semantics, norms and knowledge of the domain and provide a more professional and practical question-answering capability, a domain knowledge base is configured in this embodiment, which contains knowledge information of the target domain. The target domain is the domain to which the question-answering method is applied, such as the smart car field, medical field, education field, etc.
[0074] The domain knowledge information in the domain knowledge base can cover a variety of different sources, including but not limited to: community QA data obtained from online forums, document QA data extracted from domain professional documents, knowledge-based QA data extracted from knowledge graphs, etc. In actual applications, document QA is the main focus, and users are supported to upload documents and generate personal knowledge bases.
[0075] In a possible implementation, the domain knowledge base of the present application may include domain knowledge information of one or more modalities, including but not limited to: text modality, image modality, audio, video modality, etc.
[0076] The domain knowledge base in this embodiment can also support expansion, that is, it is an extensible domain knowledge base. Users can dynamically add new knowledge information to the domain knowledge base.
[0077] In this embodiment, domain question and answer data can be extracted based on the knowledge information in the domain knowledge base. When extracting the domain question and answer data, a rule-based extraction method can be used, or an extraction method based on an artificial intelligence algorithm can be used. A large amount of paired domain question and answer data can be obtained and applied to the subsequent model training process.
[0078] Step S120: Use the domain question and answer data to iteratively train the initial large model. In each round of training, the referee model scores the answers generated by the initial large model and the answers in the domain question and answer data respectively, and guides the training of the initial large model based on the scoring results to obtain the final trained domain question and answer large model.
[0079] Among them, domain question and answer data includes domain questions and corresponding answers (defined as the second answer).
[0080] This embodiment adopts a training strategy similar to reinforcement learning. In the iterative training process of the initial large model, in each round of training:
[0081] The referee model scores the answer corresponding to the domain question generated by the initial large model (defined as the first answer) and the second answer in the domain question and answer data respectively. Based on the scoring results, the answers and domain questions that meet the preference requirements are selected to form the target training data. The initial large model is trained using the target training data, and the domain question and answer large model is obtained after iterative training.
[0082] In an optional implementation, the process of performing any round of training on the initial large model using the domain question-answering data includes:
[0083] S1. Send the domain question in the domain question-answering data into the initial large model to obtain the first answer generated by the initial large model.
[0084] S2. The first answer and the second answer are scored respectively by the referee model, and based on the scoring results, the answers and domain questions that meet the preference requirements are selected to form the target training data.
[0085] S3. Continue training the initial large model obtained from the previous round of training using the target training data to obtain the initial large model after this round of training.
[0086] Among them, the initial large model after each round of training is used as the initial large model for the next round, that is, the initial large model is continuously updated with the iterative rounds of training. After multiple rounds of iterative training, the trained initial large model is used as the domain question and answer large model.
[0087] Since the referee model scores the first answer output by the initial large model and the second answer in the domain question and answer data in each round of training, it can measure the quality of the two answers, select the answer that meets the preference requirements from the first answer and the second answer, and form the target training data with the domain question for training the initial large model. The initial large model can continuously improve its domain question and answer capabilities based on high-quality target training data. In this way, in the next round of training, the initial large model can generate a higher quality first answer, and after being scored and evaluated by the referee model, a higher quality answer can be screened out, thereby continuously promoting the training effect of the initial large model.
[0088] In a possible implementation, the process of selecting answers that meet the preference requirements and the domain questions to form target training data based on the scoring results includes:
[0089] The answers in the scoring results are sorted by their scores, and the top N answers and domain questions are selected to form the target training data, where N is an integer greater than or equal to 1. When the preference requirement is set to select the answer with the highest score, N is equal to 1; when the preference requirement is set to select several answers with high scores, N is greater than 1.
[0090] Since the second answer is the answer corresponding to the domain question in the domain question and answer data, and the domain question and answer data is extracted from the domain knowledge base, it is possible that one or more answers are extracted from a domain question, that is, the second answer corresponding to a domain question may be one or more.
[0091] When N=1, the answer with the highest score is selected from the first answer and the second answer to form the target training data together with the domain question.
[0092] When N>1, more than two answers can be selected from the first answer and the second answer to form target training data together with the domain question.
[0093] It is understandable that the domain knowledge base can contain a large amount of domain knowledge information, so a large amount of domain question and answer data can be extracted for different domain knowledge information. When the extracted domain question and answer data is used to iteratively train the initial large model, the domain question and answer data used in different training rounds can be the same or different.
[0094] For example, in each round of training, M pairs of domain question and answer data are used to train the initial large model. After the current round of training is completed, new M pairs of domain question and answer data are selected as the next round of training data.
[0095] Another example is to extract domain question and answer data for domain knowledge information of different items in the domain knowledge base, and apply the extracted domain question and answer data to a round of training process, and then use the domain question and answer data extracted from domain knowledge information of another item in the domain knowledge base in the next round of training process.
[0096] The process of iterative training of the initial large model can have various training end conditions, such as: all domain question and answer data have been used up, or the number of iterations has reached a set number, or the training time has reached a set time, or the performance of the trained initial large model meets the set requirements, etc.
[0097] The domain question and answer big model training method provided in the embodiment of the present application uses the initial big model with question and answer capabilities as the base of the domain question and answer big model, and introduces a referee model, which is trained by question and answer data marked with preference sorting, and can give preference scores to different input answers and evaluate the quality of different answers. The present application is configured with a domain knowledge base, and domain question and answer data can be extracted based on the domain knowledge information in the domain knowledge base. When the initial big model is iteratively trained, the present application does not directly use the extracted domain question and answer data for training, but the first answer corresponding to the domain question generated by the initial big model and the second answer in the domain question and answer data are scored respectively by the referee big model, and the answers and domain questions that meet the preference requirements are selected based on the scoring results to form the target training data. Obviously, after the first answer and the second answer are scored by the referee big model, higher quality answers and domain questions can be selected to form high-quality target training data, and then the target training data can be used to train the initial big model, so that after one or more iterative trainings, the final domain question and answer big model can be obtained, which uses high-quality domain question and answer training data for training, better understands the semantics and specifications of the domain, and provides domain professionalism and more practical question and answer capabilities. Furthermore, the present application solution can obtain high-quality target training data without the need for manual proofreading of the extracted domain question and answer data, thereby saving labor costs and improving training efficiency.
[0098] In actual scenarios, the process of domain data generation itself is fluid. In terms of data timeliness, the training cycle of the large model itself determines that it is difficult to be timely, so supplementing domain data with strong timeliness is a necessary condition for the large domain model. Taking the financial field as an example, data timeliness is one of the challenges of large model implementation. How to input real-time data such as emergencies and financial information into the large model is directly related to whether the financial large model can accurately perform analysis and decision-making. To this end, an extensible domain knowledge base is provided in this embodiment.
[0099] After the domain question-answering model is trained based on the extensible domain knowledge base, it can be deployed in the application scenario. Subsequently, users can also add knowledge information to the extensible domain knowledge base periodically or in real time.
[0100] In one possible implementation, when the big model training method of the present application detects new knowledge information in the extensible domain knowledge base, it can further extract new domain question and answer data based on the new knowledge information, and use the new domain question and answer data to update and train the domain question and answer big model to obtain an updated domain question and answer big model.
[0101] In this embodiment, the process of updating and training the domain question and answer big model with new domain question and answer data is the same as the process of iteratively training the initial big model with domain question and answer data in the aforementioned step S120. It is only necessary to use the domain question and answer big model to be updated as the initial big model in step S120, and then the initial big model can be iteratively updated and trained with new domain question and answer data. In each round of training, the referee model scores the answers generated by the initial big model and the answers in the domain question and answer data respectively, and the training of the initial big model is guided by the scoring results. After iterative training, the updated domain question and answer big model is obtained.
[0102] The timing for updating and training the domain question-answering model can be set by the user, including but not limited to any one or more of the following combinations:
[0103] The amount of new knowledge information in the extensible domain knowledge base reaches the set quantity threshold, the importance of the new knowledge information reaches the set importance level, and the set large model update cycle is reached.
[0104] This embodiment configures an extensible domain knowledge base to quickly update knowledge and improve the timeliness of the domain question-answering large model response. At the same time, the existence of the extensible domain knowledge base can largely eliminate the illusion of a large model and improve the accuracy of the model response.
[0105] In some embodiments of the present application, the process of extracting domain question and answer data based on the domain knowledge base in the aforementioned step S110 is described.
[0106] The domain knowledge base may include domain knowledge information in multiple modes, such as text, image, audio, video, library table and other formats. In a possible implementation, the process of extracting domain question and answer data based on the domain knowledge base may include the following steps:
[0107] S1. For each modality of domain knowledge information in the domain knowledge base, vectorization processing is performed separately to obtain the vector representation corresponding to the domain knowledge information and store it in the vector database.
[0108] Specifically, a multimodal alignment technique can be used to align vectorized representations of domain knowledge information of different modalities into the same vector space, obtain vector representations corresponding to each piece of domain knowledge information, and store them in a vector database.
[0109] S2. Vectorize the configured query question and retrieve the top K (topK) candidate domain knowledge information with the highest similarity to the query question based on vector similarity.
[0110] The query question can be pre-set or extracted from domain knowledge information. The query question can be a domain question, so as to ensure that matching domain knowledge information can be retrieved in the domain knowledge base. The query question is vectorized, and then the similarity is calculated with each vector in the vector database. The candidate domain knowledge information corresponding to the topK vectors with the highest similarity is selected as the candidate answer corresponding to the query question.
[0111] S3. Add the topK candidate domain knowledge information and query questions to the first prompt instruction prompt, and submit them to the general large model to instruct the general large model to generate the answer that best matches the query question, and then form domain question and answer data with the query question.
[0112] The top K candidate answers obtained by the vector similarity retrieval may not be accurate enough. Therefore, in this embodiment, the capability of the large model can be called to generate the answer that best matches the query question.
[0113] In one possible implementation, a first prompt instruction prompt format template is obtained, wherein the prompt format template includes a task instruction, a question slot, and a candidate answer slot. The task instruction is used to instruct the large model to determine the answer that best matches the query question from the candidate answers provided in the candidate answer slot based on the query question in the question slot.
[0114] Fill the query question into the question slot, fill the topK candidate domain knowledge information into the candidate answer slot as candidate answers, obtain the first prompt instruction prompt, input the first prompt instruction prompt into the general large model, obtain the final answer output by the large model, and form the domain question and answer data with the query question.
[0115] Combination Figure 4 As shown, it illustrates a processing flow for extracting domain question and answer data from a domain knowledge base.
[0116] For various types of domain knowledge information contained in the domain knowledge base, including but not limited to: text data, image data, audio data, library table data, etc., these domain knowledge information are vectorized and aligned to a unified vector space to obtain a vector representation of each piece of knowledge data and store it in the vector database.
[0117] For the configured query question Q, it can be a question based on a template setting or a question extracted from domain knowledge information. The query question is vectorized to obtain a query question vector. The query question vector is used to perform vector similarity retrieval in the vector database to obtain the topK knowledge data with the highest similarity.
[0118] In order to improve the accuracy of the final extracted domain question and answer data, in this embodiment, the topK pieces of knowledge data and the query question can be filled into the prompt template to form a prompt instruction prompt. The prompt is used to instruct the large model to combine the query question and the topK pieces of knowledge data to determine the answer that best matches the query question, and obtain the final answer A corresponding to the query question output by the large model. The query question Q and the final answer A form a pair of domain question and answer data.
[0119] The domain question and answer data extraction solution in the above example combines vector retrieval technology and large model technology, which can improve the quality of the obtained domain question and answer data.
[0120] In another implementation, another possible implementation scheme for extracting domain question and answer data based on a domain knowledge base is introduced.
[0121] In this embodiment, the big model can be directly called to extract domain question and answer data. Exemplarily, the general big model is called to instruct the general big model to extract a number of questions and corresponding answers from the provided domain knowledge information. The questions and corresponding answers extracted by the general big model constitute the domain question and answer data.
[0122] The following is an example of a possible prompt command prompt:
[0123] Extract several questions and corresponding answers from the following information, which must be logical and factual, and give the results in the form of "Question:" and "Answer:". [Domain knowledge information].
[0124] Domain knowledge information can be extracted from the domain knowledge base and filled into the above prompt, and then the big model can be called to extract questions and answers.
[0125] It is understandable that, considering the input length limit of the large model, the domain knowledge information in the domain knowledge base can be segmented, and a part of the domain knowledge information can be sent into the large model each time for extracting questions and answers.
[0126] This embodiment injects domain knowledge information into the prompt and instructs the big model to extract questions and corresponding answers from the domain knowledge information. The ability of the big model can be used to extract domain question and answer data with higher efficiency.
[0127] Reference Figure 5 , providing a training process for a large domain question-answering model.
[0128] Based on the general large model, the general large model is fine-tuned by using the public data set to extract question-answering data to obtain the SFT large model, which has preliminary question-answering capabilities. The SFT large model can be used as the initial large model to be trained.
[0129] Furthermore, for multiple sources of different answers to the same question and answer, expert annotation can be used to obtain question and answer data annotated with preference rankings. The annotated question and answer data can then be used to supervise the training of the general large model RM to obtain the trained referee large model (also called reward model). The referee large model can score the answers corresponding to the input domain questions, and the scores can measure the quality of the answers.
[0130] This embodiment may also configure an extensible domain knowledge base, which stores domain knowledge information and supports dynamic update. Domain question and answer data QA may be extracted from the domain knowledge base.
[0131] In the iterative training process of the initial large model, the domain question Q in the domain question and answer data is sent to the initial large model in each round of training to obtain the answer A' generated by the initial large model. The domain question Q, answer A and answer A' are sent to the referee large model together. The referee large model gives the preference scores of answers A and A', and selects the answer with the highest score to form the target training data with the domain question Q. The initial large model is trained using the target training data.
[0132] After multiple rounds of iterative training, the initial large model after the last round of training is used as the final domain question-answering large model.
[0133] The trained domain question-answering model in this embodiment supports private deployment, reducing the risk of data leakage and security vulnerabilities, and providing higher credibility and protection for tasks involving sensitive information. The private deployment of the domain question-answering model enables enterprises to better manage and control data governance.
[0134] The existence of an extensible domain knowledge base can largely eliminate model illusions and improve the accuracy of the responses of the domain question-and-answer large model. At the same time, the extensible domain knowledge base can quickly update knowledge and improve the timeliness of the responses of the domain question-and-answer large model. The introduction of the referee large model enriches the source of high-quality domain question-and-answer data, maximizes the role of professional knowledge and industry data accumulation, and eliminates the need for manual proofreading of domain question-and-answer data, saving labor costs.
[0135] 2. Reasoning Stage
[0136] Based on the domain question-answering model obtained after training in the above embodiments, this embodiment further provides a question-answering method, which specifically includes:
[0137] First, get the domain question.
[0138] Furthermore, the domain questions are sent to the configured domain question-answering model to obtain the answers output by the model.
[0139] Specifically, in the domain question and answer scenario, users can interact with the domain question and answer big model, and send the user's interaction content as domain questions into the domain question and answer big model. The big model generates the corresponding answers and outputs them to the users.
[0140] Based on the training method of the domain question-answering big model introduced in the aforementioned embodiment, it can be seen that the model training process uses extensible domain knowledge information, so that the model can better understand the semantics, norms and knowledge of the domain, and thus can provide question-answering capabilities with stronger domain professionalism and practicality.
[0141] The following is a description of the domain question and answer big model training device provided in an embodiment of the present application. The domain question and answer big model training device described below and the domain question and answer big model training method described above can be referenced to each other.
[0142] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a domain question-answering large model training device disclosed in an embodiment of the present application.
[0143] like Figure 6 As shown, the device may include:
[0144] A model acquisition unit 11 is used to acquire an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings;
[0145] A knowledge base processing unit 12 is used to obtain a domain knowledge base and extract domain question and answer data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information;
[0146] The model training unit 13 is used to iteratively train the initial large model using the domain question and answer data. In each round of training, the referee model scores the first answer corresponding to the domain question generated by the initial large model and the second answer corresponding to the domain question in the domain question and answer data respectively. Based on the scoring results, the answer that meets the preference requirements and the domain question are selected to form target training data. The initial large model is trained using the target training data, and the domain question and answer large model is obtained after the iterative training.
[0147] In a possible implementation, when the model training unit iteratively trains the initial large model, different domain question and answer data are used in different training rounds.
[0148] In a possible implementation, the process of the model training unit performing any round of training on the initial large model using the domain question and answer data includes:
[0149] Sending the domain question in the domain question-answering data into the initial large model to obtain the first answer generated by the initial large model;
[0150] Scoring the first answer and the second answer respectively by the referee model, and selecting answers that meet the preference requirements and the domain question to form target training data based on the scoring results;
[0151] The target training data is used to continue training the initial large model obtained in the previous round of training to obtain the initial large model after this round of training.
[0152] In a possible implementation, the model training unit selects answers that meet the preference requirements and the domain questions to form target training data based on the scoring results, including:
[0153] The answers in the scoring results are sorted according to their scores, and the top N answers in the sorting are selected together with the domain question to form the target training data, where N is an integer greater than or equal to 1.
[0154] In a possible implementation, the domain knowledge base is an extensible domain knowledge base, and the device of the present application may further include:
[0155] The model update training unit is used to extract new domain question and answer data based on the newly added knowledge information when new knowledge information is detected in the extensible domain knowledge base, and use the new domain question and answer data to update and train the domain question and answer big model to obtain an updated domain question and answer big model.
[0156] In one possible implementation, the referee model acquired by the model acquisition unit is a referee large model obtained by supervised training of a general large model using question and answer data labeled with preference rankings.
[0157] In a possible implementation, the process of extracting domain question and answer data based on the domain knowledge base by the knowledge base processing unit includes:
[0158] For each modality of domain knowledge information in the domain knowledge base, vectorization processing is performed respectively to obtain a vector representation corresponding to the domain knowledge information and store it in a vector database;
[0159] Vectorize the configured query question and retrieve the top K candidate domain knowledge information matching the query question based on vector similarity;
[0160] The topK candidate domain knowledge information and the query question are added to the first prompt instruction prompt, and submitted to the general large model to instruct the general large model to generate the answer that best matches the query question, and then form domain question and answer data with the query question.
[0161] In another possible implementation, the process of extracting domain question and answer data based on the domain knowledge base by the knowledge base processing unit includes:
[0162] Calling a general big model to instruct the general big model to extract a number of questions and corresponding answers from the provided domain knowledge information;
[0163] The questions extracted by the general large model and the corresponding answers constitute the domain question and answer data.
[0164] The present application also provides an electronic device in an embodiment. Figure 7 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, tablet computers, teaching large screens, wearable devices, etc. Figure 7 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0165] like Figure 7As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603, so as to implement the domain question-answering large model training method of the aforementioned embodiment of the present application, or the question-answering method of the aforementioned embodiment. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in RAM 603. The processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0166] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0167] Also provided in an embodiment of the present application is a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the field question-and-answer big model training methods provided in the embodiments of the present application, or the question-and-answer method of the aforementioned embodiments.
[0168] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any one of the field question-and-answer large model training methods provided in the embodiments of the present application, or the question-and-answer method of the aforementioned embodiments.
[0169] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0170] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0171] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0172] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0173] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
Claims
1. A method for training a large domain question answering model, characterized in that: include: Obtain an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings; Acquire a domain knowledge base, and extract domain question-answering data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information; The initial large model is iteratively trained using the domain question and answer data. In each round of training, the referee model scores the first answer corresponding to the domain question generated by the initial large model and the second answer corresponding to the domain question in the domain question and answer data respectively. Based on the scoring results, the answer that meets the preference requirements and the domain question are selected to form target training data. The initial large model is trained using the target training data, and the domain question and answer large model is obtained after the iterative training.
2. The method according to claim 1, characterized in that: During the iterative training of the initial large model, different domain question and answer data are used in different training rounds.
3. The method according to claim 1, characterized in that The process of performing any round of training on the initial large model using the domain question-answering data includes: Sending the domain question in the domain question-answering data into the initial large model to obtain the first answer generated by the initial large model; Scoring the first answer and the second answer respectively by the referee model, and selecting answers that meet the preference requirements and the domain question to form target training data based on the scoring results; The target training data is used to continue training the initial large model obtained in the previous round of training to obtain the initial large model after this round of training.
4. The method according to claim 1, characterized in that The process of selecting answers that meet the preference requirements and the field questions to form target training data based on the scoring results includes: The answers in the scoring results are sorted according to their scores, and the top N answers in the sorting are selected together with the domain question to form the target training data, where N is an integer greater than or equal to 1.
5. The method according to claim 1, characterized in that The domain knowledge base is an extensible domain knowledge base, and the method further includes: When new knowledge information is detected in the extensible domain knowledge base, new domain question and answer data is extracted based on the new knowledge information, and the domain question and answer big model is updated and trained using the new domain question and answer data to obtain an updated domain question and answer big model.
6. The method according to claim 1, characterized in that The referee model is a large referee model obtained by supervised training of a general large model using question-answer data labeled with preference rankings.
7. The method according to any one of claims 1 to 6, characterized in that: The process of extracting domain question-answering data based on the domain knowledge base includes: For each modality of domain knowledge information in the domain knowledge base, vectorization processing is performed respectively to obtain a vector representation corresponding to the domain knowledge information and store it in a vector database; The configured query question is vectorized, and the top K candidate domain knowledge information with the highest similarity to the query question is retrieved based on vector similarity, where K is a set positive integer; The first K candidate domain knowledge information and the query question are added to the first prompt instruction prompt, and submitted to the general large model to instruct the general large model to generate the answer that best matches the query question, and then form domain question and answer data with the query question.
8. The method according to any one of claims 1 to 6, characterized in that: The process of extracting domain question-answering data based on the domain knowledge base includes: Calling a general big model to instruct the general big model to extract a number of questions and corresponding answers from the provided domain knowledge information; The questions extracted by the general large model and the corresponding answers constitute the domain question and answer data.
9. A question-answering method, characterized in that: include: Get domain questions; Send the domain question to the configured domain question-answering model to obtain the answer output by the model; Wherein, the domain question and answer big model is trained using the domain question and answer big model training method described in any one of claims 1 to 8.
10. A domain question-answering large model training device, characterized in that: include: A model acquisition unit, used to acquire an initial large model with question-answering capabilities, and a referee model trained using question-answering data labeled with preference rankings; A knowledge base processing unit, used to obtain a domain knowledge base and extract domain question and answer data based on the domain knowledge base, wherein the domain knowledge base includes domain knowledge information; A model training unit is used to iteratively train the initial large model using the domain question and answer data. In each round of training, the referee model scores the first answer corresponding to the domain question generated by the initial large model and the second answer in the domain question and answer data respectively. Based on the scoring results, answers that meet the preference requirements and the domain question are selected to form target training data. The initial large model is trained using the target training data until the domain question and answer large model is obtained after the last round of training.
11. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the domain question and answer large model training method as described in any one of claims 1 to 8, or to implement the various steps of the question and answer method as described in claim 9.
12. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the various steps of the domain question-answering large model training method as described in any one of claims 1 to 8, or implements the various steps of the question-answering method as described in claim 9.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the various steps of the domain question-answering large model training method as described in any one of claims 1 to 8, or implements the various steps of the question-answering method as described in claim 9.
Citation Information
Patent Citations
Method and device for training generative large language model based on knowledge base feedback
CN117009490A
Communication large model construction method and device, equipment and storage medium
CN117527608A
Question and answer dialogue data generation system and method for constructing stomatology large model
CN118520082A
Multi-mode RAG knowledge question-answering method and device applied to vertical field
CN119150998A
Information processing method and device based on large model and intelligent agent
CN119271782A
Cited By
Large language model optimization method and device based on multi-modal feedback and reinforcement learning
CN120386849A
Risk detection method and system and storage medium
CN121146073A
Risk detection methods, systems and storage media
CN121146073B
Model knowledge injection method and device and electronic equipment
CN121436088A