Knowledge question and answer model training method and device, equipment and storage medium
By enhancing and adjusting the training dataset of large language models, the problem of insufficient professionalism and complexity in specific fields is solved, and more accurate prediction responses and multi-task comprehension capabilities are achieved.
Patent Information
- Application Number
- CN202410006582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing large language models lack professionalism and complexity in dealing with knowledge in specific fields, resulting in low quality and poor accuracy in predicted response results.
By obtaining the training data sets of multiple tasks, they are enhanced, they are generated, similar questions and labeled reply results are generated, the training data set after multi-task enhancement is constructed, the pre-trained large language model is adjusted, and the knowledge question and answer model is formed.
It improves the accuracy of large language models in specific fields, has the ability to understand multi-tasks, and can generate more accurate prediction response results, adapting to the needs of different tasks.
Smart Images

Figure CN120256551A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method, device, equipment, and storage medium for a knowledge Q&A model. Background Art
[0002] Currently, large language models have achieved important breakthroughs in the field of natural language processing. However, existing large language models still perform poorly when dealing with knowledge in specific fields. Specifically, when applied to a specific field, existing large language models often lack an understanding of the professionalism and complexity of the knowledge in that specific field, resulting in problems such as low quality and poor accuracy of the predicted reply results. Summary of the Invention
[0003] Embodiments of this application provide a training method, device, equipment, and storage medium for a knowledge Q&A model. The technical solutions provided by the embodiments of this application are as follows:
[0004] According to one aspect of the embodiments of this application, a training method for a knowledge Q&A model is provided. The method includes:
[0005] For each of multiple tasks, obtain the training data set of the task, where the training data set of the task includes multiple questions related to the task, and the labeled reply results corresponding to the multiple questions respectively;
[0006] Perform knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task, where the knowledge enhancement is used to increase the number of questions included in the training data set of the task;
[0007] Use the enhanced training data sets of the multiple tasks to adjust a pre-trained large language model, and use the adjusted large language model as a knowledge Q&A model, where the knowledge Q&A model is used to generate a predicted reply result corresponding to a question to be replied according to the question to be replied.
[0008] According to one aspect of the embodiments of this application, a training device for a knowledge Q&A model is provided. The device includes:
[0009] An obtaining module, configured to obtain the training data set of the task for each of multiple tasks, where the training data set of the task includes multiple questions related to the task, and the labeled reply results corresponding to the multiple questions respectively;
[0010] A obtaining module, configured to perform knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task, where the knowledge enhancement is used to increase the number of questions included in the training data set of the task;
[0011] An adjustment module, configured to adjust a pre-trained large language model by using the enhanced training data sets of the multiple tasks respectively, and use the adjusted large language model as a knowledge question-answering model, where the knowledge question-answering model is configured to generate a predicted reply result corresponding to the question to be replied according to the question to be replied.
[0012] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the training method of the above knowledge question-answering model.
[0013] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the training method of the above knowledge question-answering model.
[0014] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor reads and executes the computer program from the computer-readable storage medium to implement the training method of the above knowledge question-answering model.
[0015] The technical solution provided by the embodiments of the present application at least includes the following beneficial effects:
[0016] By adjusting the pre-trained large language model by using the enhanced training data sets of multiple tasks, the problems of insufficient professionalism and complexity of the existing large language model in dealing with problems in specific fields are solved. On the one hand, the knowledge question-answering model can generate a more accurate predicted reply result according to the question to be replied. On the other hand, by integrating the enhanced training sets of multiple tasks, the knowledge question-answering model has the ability to understand multiple tasks and can generate the corresponding predicted reply results for different tasks. Description of the Drawings
[0017] Figure 1 is a schematic diagram of the implementation environment of the solution provided by an embodiment of the present application;
[0018] Figure 2 is a flowchart of the training method of the knowledge question-answering model provided by an embodiment of the present application;
[0019] Figure 3 is a flowchart of obtaining the training data set provided by an embodiment of the present application;
[0020] Figure 4 is a flowchart of the training method of the knowledge question-answering model provided by another embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of large language model adjustment provided by an embodiment of the present application;
[0022] Figure 6 It is a schematic diagram of multi-task large language model adjustment provided by an embodiment of the present application;
[0023] Figure 7 It is a block diagram of a training device for a knowledge answering model provided by an embodiment of the present application;
[0024] Figure 8 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0026] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0027] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0028] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0029] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, Artificial Intelligence Generated Content (AIGC), conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0030] Pre-training Model, also known as the foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and the function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. Pre-trained models are important tools for outputting artificial intelligence-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0031] The solution provided in the embodiments of this application relates to technologies such as pre-trained model technology of artificial intelligence, and will be specifically described through the following embodiments.
[0032] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided by an embodiment of the present application. The solution implementation environment may include a model training device 110 and a model using device 120.
[0033] The model training device 110 may be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, an intelligent robot, an intelligent TV, a multimedia playback device, etc., or some other electronic device with strong computing power. The present application does not limit this. The model training device 110 is used to train the knowledge question-answering model.
[0034] In the embodiment of the present application, the knowledge question-answering model is a model obtained by further training a pre-trained model. Optionally, the model training device 110 may adopt a machine learning method to train the knowledge question-answering model to make it have better performance. Optionally, the training process of the knowledge question-answering model is as follows (only a brief description here, and the specific training process can be seen in the following embodiments): For each of multiple tasks, obtain the training data set of the task, perform knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task, and use the enhanced training data sets of the multiple tasks to adjust the pre-trained large language model, and use the adjusted large language model as the knowledge question-answering model.
[0035] The model using device 120 may be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, an intelligent robot, an intelligent TV, a multimedia playback device, etc., or some other electronic device with strong computing power. The present application does not limit this. The model using device 120 may adopt the knowledge question-answering model to generate a predicted reply result corresponding to the to-be-replied question according to the to-be-replied question.
[0036] The model training device 110 and the model using device 120 may be two independent devices or the same device.
[0037] For the method provided by the embodiment of the present application, the execution subject of each step may be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. Among them, when the computer device is a server, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The computer device may be Figure 1 the model training device 110 in, or the model using device 120.
[0038] In some embodiments, by constructing an enhanced training dataset related to the domain, the knowledge-based question answering model can answer questions more accurately and effectively in a specific domain. The technical solution proposed in this application can be applied to the knowledge-based question answering scenarios in the field of data analysis, the knowledge-based question answering scenarios in the medical field, and the knowledge-based question answering scenarios in the legal field. The application field of the solution is not limited in this application.
[0039] In some embodiments, considering the sensitivity of the enterprise's business data and cost optimization, if an internally trained model can be used, not only can business data be protected from leakage, but also the pre-trained large language model can be adjusted by constructing an enhanced training dataset for a specific domain to save operating costs, and all aspects of the model are relatively controllable. Exemplarily, an insurance company needs a large language model that can process customer claims. A general large language model can be selected for task processing, but this general large language model may not accurately understand the specific knowledge and terms in the insurance domain and may not be able to provide professional advice related to insurance claims. To solve this problem, the enterprise can use an internal dataset (such as historical claim cases, insurance regulations, etc.) to construct an enhanced training dataset, and then adjust the general large language model according to the enhanced training dataset, so that the adjusted large language model is more focused on the insurance domain and learns the claim process, regulations, and business characteristics, etc. Through the above method, the enterprise can train with internal data, protect business data from being leaked, and can construct a corresponding enhanced training dataset according to the actual situation to adjust the general large language model to meet the needs of a specific domain.
[0040] Please refer to Figure 2 , which shows a flowchart of a method for training a knowledge-based question answering model provided by an embodiment of this application. The execution subject of each step of this method can be a computer device. For example, this computer device can be Figure 1 the model training device 110 in the solution implementation environment shown. This method can include at least one of the following steps 210 to 230.
[0041] Step 210, for each of multiple tasks, obtain the training dataset of the task, where the training dataset of the task includes multiple questions related to the task and the annotated reply results corresponding to the multiple questions respectively.
[0042] A task refers to a specific job or activity that needs to be completed in a particular domain. The categories of tasks can be prediction, classification, data statistics, etc. Exemplarily, in the field of data analysis, a task can be predicting user consumption behavior. Based on data such as a customer's historical transactions and browsing, it is possible to predict the likely consumption patterns, consumption categories, and consumption amounts of the user within a certain period in the future. A task can also be predicting the demand for a product. By collecting and analyzing data such as historical sales volume, query volume, and elimination volume, the market demand for the product in the future can be predicted. Exemplarily, in the field of medical diagnosis, a task can be a disease classification task, a drug recommendation task, or a prediction task of a patient's condition. Exemplarily, in the field of natural language processing, a task can be translating from the current language to another language based on text, or obtaining a summary of the text based on the text.
[0043] The training dataset of a task refers to a dataset that contains multiple input data related to the completion of the task, as well as the output results corresponding to the multiple input data respectively. These input data can be used as the input of the model, and the output results are the expected outputs of the model.
[0044] In some embodiments, please refer to Figure 3 , which shows a schematic diagram of obtaining a training dataset provided by an embodiment of the present application. From the online log data related to the task, at least one first question related to the task is collected, and an annotated reply result corresponding to each of the at least one first question is generated through a second large language model; according to predefined rules, at least one second question related to the task is constructed, and an annotated reply result corresponding to each of the at least one second question is generated through a third large language model, and the predefined rules are used to specify the conditions satisfied by the questions related to the task; according to the at least one first question and the annotated reply result corresponding to each of the at least one first question, and the at least one second question and the annotated reply result corresponding to each of the at least one second question, the training dataset of the task is obtained.
[0045] The online log data related to the task refers to data logs such as user queries and operation records related to the process of completing the task. The first question refers to the actual questions or query statements strongly related to the task extracted from the online log data.
[0046] The second large language model refers to a pre-trained large language model with strong language understanding and generation capabilities. The purpose of the second large language model is to generate reliable reply results for the first question, which is collected from the online log data. The second large language model can generate a reply result for the first question based on the first question.
[0047] In some embodiments, the response results of the first question can be obtained by multiple second large language models respectively. For the above response results, through manual evaluation and screening, the labeled response results corresponding to the first question can be obtained.
[0048] The second question refers to an artificially synthesized question related to the task generated according to predefined rules. Since these questions are artificially constructed, their quality can be guaranteed, and they can effectively supplement the first questions extracted from the online log data. The predefined rules refer to the rules or templates manually formulated for constructing synthetic questions related to the task. These rules will clearly stipulate the conditions that the constructed second questions need to meet. Exemplarily, the predefined rules can stipulate that the length of the question statement should be between 20 and 50 words, or it can stipulate that the question statement must contain 1-2 keywords related to the task theme, such as keywords containing specific movie or TV drama names in the "movie and TV drama data analysis" task.
[0049] In some embodiments, the predefined rules can be a set of keywords or phrases related to purchase behavior and user needs. According to the predefined rules, the corresponding second questions can be obtained. Exemplarily, the keywords included in the predefined rules can include "purchase", "user feedback", etc. According to these keywords, the following second questions can be artificially constructed: What are the most frequently purchased product categories by users? What factors are often related to users' purchase preferences? What are the most frequently mentioned problems in user feedback? etc.
[0050] In some embodiments, the corresponding second questions can be obtained according to the predefined rules through a keyword matching algorithm. This application does not limit the method of obtaining the corresponding second questions according to the predefined rules.
[0051] The third large language model refers to a pre-trained large language model that can generate the response results of the second question. Constructing the second question can increase the diversity of questions in the training dataset. Through the third large language model, reliable response results of the second question can be obtained, thereby providing richer and more accurate samples for the training dataset, which helps to improve the generalization ability of the model to diverse question scenarios.
[0052] In some embodiments, the response results of the second question can be obtained by multiple third large language models respectively. For the above response results, through manual evaluation and screening, the labeled response results corresponding to the second question can be obtained.
[0053] The labeled response results refer to the correct or expected response results manually selected from the response results of the second large language model and the third large language model corresponding to the first question or the second question. Please refer to Figure 4, the above obtained the response results of the first question and the second question through the second large language model and the third large language model respectively. Since there may be certain errors or deficiencies in the large language model when generating responses, it is necessary to manually evaluate and screen the response results to obtain the labeled response results corresponding to the first question and the second question respectively. Through manual evaluation and screening, the labeled response results corresponding to the first question and the second question can be obtained, and these labeled response results can be used as reliable references to provide accurate and reasonable response results to users in actual applications.
[0054] In the above method, the first question is obtained from the online log data, and the second question is constructed through predefined rules. The combination of the two is used as the training data set of the task, making the training data set more complete and rich.
[0055] Step 220, perform knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task. Among them, knowledge enhancement is used to increase the number of questions included in the training data set of the task.
[0056] In actual applications, the first question and the second question may involve many business-related terms or domain knowledge that cannot be understood by general large language models. Therefore, it is necessary to further optimize the training data set. Specifically, it is to optimize the descriptions of the first question and the second question, that is, knowledge enhancement. By generating semantically similar questions to the first question and the second question, the expression methods and knowledge coverage of the questions are expanded. This strengthens the understanding of the questions by the large language model to be adjusted and lays a foundation for generating more accurate response results subsequently.
[0057] In some embodiments, the process of knowledge enhancement is as Figure 4 shown. Select at least one target question from multiple questions related to the task included in the training data set of the task; generate similar questions corresponding to the target question, where the similar questions have the same or similar semantics as the target question but different expressions; determine the labeled response result corresponding to the target question as the labeled response result corresponding to the similar question; add the similar question and the labeled response result corresponding to the similar question to the training data set of the task to obtain the enhanced training data set of the task.
[0058] The target question refers to any one of the first question or the second question. Determining the labeled response result corresponding to the target question as the labeled response result corresponding to the similar question means directly referring to the manually labeled response result of the target question as the labeled response result of the similar question with similar semantics.
[0059] In some embodiments, similar questions corresponding to the target question are generated by a pre-trained first large language model. Alternatively, similar questions corresponding to the target question can be obtained through manual construction methods, and the present application does not limit this.
[0060] Based on the first large language model, a target question can be input. This first large language model can analyze its semantic meaning and generate multiple similar questions that are similar in meaning or related to it as the similar question set for this target question. Generating similar questions through the first large language model, on the one hand, enhances the understanding of the question by the large language model to be adjusted, and on the other hand, increases the number of samples in the training data set.
[0061] In some embodiments, multiple first large language models can be used to separately obtain similar questions according to the target question, and the similar questions separately obtained by the multiple first large language models are combined as the similar question set for this target question.
[0062] Through the above method, by constructing similar questions related to the first question and the second question, the expression form of the questions can be diversified, promoting the pre-trained large language model to be adjusted to better learn the relationship between the questions and the annotated reply results.
[0063] Step 230: Use the training data sets enhanced by multiple tasks respectively to adjust the pre-trained large language model, and use the adjusted large language model as a knowledge-answering model. The knowledge-answering model is used to generate a predicted reply result corresponding to the question to be replied according to the question to be replied.
[0064] The pre-trained large language model refers to the model to be trained, such as the Bidirectional Encoder Representations from Transformers (BERT) series of models. The knowledge-answering model refers to the adjusted large language model that supports the question-answering service.
[0065] In some embodiments, the above-mentioned first large language model, second large language model, third large language model, and pre-trained large language model all refer to any large language model. They are all based on deep learning technology, have language understanding and generation capabilities, and can encode and generate text. The difference is that the first large language model: is used to generate multiple similar questions that are similar in meaning or related to the target question to enhance the diversity and coverage of the questions included in the training data set. The second large language model: is used to generate the reply result corresponding to the first question to construct the training data set. The third large language model: is used to generate the reply result corresponding to the second question to enrich the samples in the training data set. The pre-trained large language model: refers to the large language model that needs to be adjusted. Therefore, the above-mentioned first large language model, second large language model, third large language model, and pre-trained large language model can all be the same, can all be different, or at least two of them can be the same, and the present application does not limit this.
[0066] In some embodiments, the first large language model, the second large language model, the third large language model, and the pre-trained large language model can be any publicly available large language model, such as a natural language model based on the transformer architecture trained with a large amount of data. The large amount of data can reach the sample level of over 100 million, and the present application does not make any limitations in this regard.
[0067] Please refer to Figure 5 , which shows a schematic diagram of large language model adjustment provided by an embodiment of the present application. Among them, subfigure (a) represents parameter adjustment of the pre-trained large language model according to the enhanced dataset to obtain an adjusted large language model as a knowledge question-answering model.
[0068] Subfigure (b) obtains similar questions corresponding to the target questions in the training dataset. Each target question can correspond to multiple similar questions. In subfigure (b), N is an integer greater than 1. According to the labeled reply results in the training dataset, multiple knowledge sets can be obtained by random sampling. In subfigure (b), M is an integer greater than 1. According to the similar questions and the knowledge sets, an enhanced training dataset is obtained, and the pre-trained large language model is adjusted according to the enhanced training dataset. Specifically, when the model inputs a question, the model will retrieve multiple labeled reply results and select the labeled reply result that best matches the current input question from the multiple labeled reply results. The pre-trained large language model is adjusted based on the difference between the labeled reply result and the predicted reply result to obtain an adjusted large language model as a knowledge question-answering model. This knowledge question-answering model can not only better understand the knowledge in a specific field but also achieve the purpose of performing multiple tasks by fine-tuning the pre-trained large language model with the enhanced training dataset, and the accuracy of each task exceeds 80%.
[0069] In some embodiments, after constructing the multi-task dataset, optionally, methods such as LORA (Language Optimization for Retrieval and Adaption), full model fine-tuning, Prompt Tuning, etc. can be used to adjust the pre-trained large language model, and the present application does not make any limitations in this regard.
[0070] In some embodiments, please refer to Figure 6 , which shows a schematic diagram of multi-task large language model adjustment provided by an embodiment of the present application. Figure 6Where N is an integer greater than 1. For each of multiple tasks, according to the task-enhanced training dataset, generate the prompt information for the task, where the prompt information is used to indicate to select the response result corresponding to the question related to the task from the knowledge set of the task. The knowledge set of the task includes the labeled response results corresponding to multiple questions related to the task; input the prompt information of the task into the pre-trained large language model, and output the predicted response result corresponding to the question related to the task through the pre-trained large language model; adjust the pre-trained large language model according to the labeled response results and predicted response results corresponding to the questions respectively related to multiple tasks, and use the adjusted large language model as the knowledge Q&A model.
[0071] The knowledge set of the task can be randomly selected from the labeled response results corresponding to multiple questions related to the task, and one task includes multiple knowledge sets.
[0072] In some embodiments, generate the knowledge set of the task according to the labeled response results corresponding to multiple questions related to the task; generate the prompt information of the task according to the prompt template, the questions included in the task-enhanced training dataset, and the knowledge set of the task, where multiple tasks share the same prompt template, and the prompt template is used to define the format of the prompt information.
[0073] The prompt information is also called prompt, and the prompt template is also called prompt template. One task includes multiple prompt information, that is, one task includes multiple prompts.
[0074] Exemplarily, task 1 can be the analysis of film and television drama data. The first question related to task 1 can be what is the cumulative box office of movie A, the second question related to task 1 can be what is the release date of movie A, and the third question related to task 1 can be what is the cumulative box office of movie B. The labeled response result of the first question can include: the box office of movie A is x ten million. The labeled response result of the second question can include: the release date of movie A is xx year xx month xx day. The labeled response result of the third question can include: the box office of movie B is x hundred million. Then the knowledge set of task 1 can be randomly selected from the box office of movie A is x ten million, the release date of movie A is xx year xx month xx day, and the box office of movie B is x hundred million. Exemplarily, the knowledge set of task 1 can be [the box office of movie A is x ten million, the release date of movie A is xx year xx month xx day]. The knowledge set 2 of task 1 can also be [the box office of movie B is x hundred million, the box office of movie A is x ten million]. The above is only an example, and the composition of the knowledge set can have multiple different extraction and combination methods, and this application does not make specific limitations.
[0075] Exemplarily, Task 2 can be video data analysis. The first question related to Task 1 can be how many clicks video A has, and the second question related to Task 1 can be how long is the average viewing duration of video A. The labeled reply results for the first question can include: The number of clicks of video A is 50,000 times, and the number of viewers is 20,000 people. The labeled reply results for the second question can include: The average viewing duration of video A is 1 minute and 20 seconds, and the completion rate of playback is 60%. The knowledge set of Task 2 can be ["The number of clicks is 50,000 times", "The average viewing duration is 1 minute and 20 seconds"]. The knowledge set of Task 2 can also be ["The number of clicks is 50,000 times", "The completion rate of playback is 60%"]. The above are only examples, and the composition of the knowledge set can have various different extraction and combination methods, which are not specifically limited in this application.
[0076] The above Task 1 or Task 2 is any one of multiple tasks. According to the same method, knowledge sets corresponding to multiple tasks can be obtained respectively.
[0077] A prompt template refers to a template or rule for defining the format of prompt information for a task. By adding task tags such as [Task 1] and [Task 2], this prompt template can indicate which task the current question belongs to, so as to prompt the pre-trained large language model to generate a prediction reply result that conforms to the current task. Exemplarily, Prompt Template A can be [Task x, question related to Task x, knowledge set of Task x], where x refers to the xth task and takes an integer greater than 1.
[0078] According to this Prompt Template A, at least one piece of prompt information for Task 1 can be obtained. Exemplarily, the prompt information 1 corresponding to Task 1 can be: ["Film and television drama data analysis", "What is the cumulative box office of Movie A", "[The box office of Movie A is x tens of millions, and the release date of Movie A is xx year xx month xx day]"], where "Film and television drama data analysis" corresponds to "Task x" in the above prompt template, "What is the cumulative box office of Movie A" corresponds to "question related to Task x" in the above prompt template, "[The box office of Movie A is x tens of millions, and the release date of Movie A is xx year xx month xx day]" corresponds to "knowledge set of Task x" in the above prompt template, and "The box office of Movie A is x tens of millions" is the labeled reply result for the question "What is the cumulative box office of Movie A".
[0079] Exemplarily, the prompt information 2 corresponding to task 1 can be: ["Data analysis of film and television dramas", "What is the cumulative box office of movie B", "[The box office of movie B is x billion, and the box office of movie A is x0 million]"], where "Data analysis of film and television dramas" corresponds to "Task x" in the above prompt template, "What is the cumulative box office of movie B" corresponds to "Question related to Task x" in the above prompt template, "[The box office of movie B is x billion, and the box office of movie A is x0 million]" corresponds to "Knowledge set of Task x" in the above prompt template, and "The box office of movie B is x billion" is the labeled reply result corresponding to the question "What is the cumulative box office of movie B".
[0080] In some embodiments, multiple tasks can share the same prompt template. Therefore, according to prompt template A for the above task 2, prompt information 1 can be obtained, such as ["Data analysis of videos", "What is the click-through rate of video A", "[The click-through rate is 50,000 times", "The average viewing duration is 1 minute and 20 seconds]"], where "The click-through rate is 50,000 times" is the labeled reply result of the question "What is the click-through rate of video A". According to the same method, the prompt information of multiple tasks can be obtained.
[0081] In some embodiments, multiple tasks can use different prompt templates. Exemplarily, prompt template B can be [Task x, Question related to Task x, Knowledge set, Field involved in Task x], where x refers to the xth task and takes an integer greater than 1.
[0082] Exemplarily, task 1 can use prompt template A to obtain the corresponding prompt information. Task 2 can obtain at least one prompt information of task 2 according to prompt template B. Exemplarily, the prompt information 2 of task 2 can be ["Data analysis of videos", "What is the click-through rate of video A", "[The click-through rate is 50,000 times", "The average viewing duration is 1 minute and 20 seconds", "Video]"], where "Video" corresponds to "Field involved in Task x" of prompt template B. According to the same method, the prompt templates corresponding to multiple tasks can be defined respectively, and the prompt information corresponding to multiple tasks can be obtained according to the prompt templates.
[0083] Such as Figure 6As shown, after obtaining the prompt information corresponding to multiple tasks, the prompt information corresponding to multiple tasks is fused. This fusion means merging multiple prompt information and then scrambling them. Exemplarily, merging the prompt information 1 corresponding to task 1, the prompt information 2 corresponding to task 1, and the prompt information 1 of task 2 means concatenating the three, such as [prompt information 1 corresponding to task 1, prompt information 2 corresponding to task 1, prompt information of task 2]. Scrambling means shuffling the order of the three, such as [prompt information 1 corresponding to task 1, prompt information 1 of task 2, prompt information 2 corresponding to task 1]. The final result can be {["Film and TV drama data analysis", "What is the cumulative box office of Movie A", "[The box office of Movie A is x ten million, and the release time of Movie A is xx year xx month xx day]"], ["Video data analysis", "What is the click-through rate of Video A", "[The click-through rate is 50,000 times", "The average viewing duration is 1 minute and 20 seconds]"], ["Film and TV drama data analysis", "What is the cumulative box office of Movie B", "[The box office of Movie B is x billion, and the box office of Movie A is x ten million]"]}.
[0084] In some embodiments, the prompt information of the task is input into a pre-trained large language model, and the pre-trained large language model outputs the predicted reply result corresponding to the question related to the task; according to the labeled reply result and the predicted reply result corresponding to each question related to multiple tasks respectively, the pre-trained large language model is adjusted, and the adjusted large language model is used as a knowledge question and answer model. Exemplarily, the prompt information 1 corresponding to task 1 is input into the pre-trained large language model. The model can select the predicted reply result corresponding to the question "What is the cumulative box office of Movie A" from the labeled reply results of "The box office of Movie A is x ten million" and "The release time of Movie A is xx year xx month xx day", calculate the value of the loss function through the difference between the predicted reply result and the labeled reply result, and aim to minimize the value of the loss function to adjust the parameters of the pre-trained large language model. When the value of the loss function is less than the set threshold, the adjusted large language model can be obtained and used as a knowledge question and answer model.
[0085] In some embodiments, after obtaining the knowledge question and answer model, it can be sent to a model using device, such as Figure 1 the model using device 120 in, and the knowledge question and answer model generates the predicted reply result corresponding to the to-be-replied question according to the to-be-replied question.
[0086] The technical solution provided by the embodiments of the present application adjusts a pre-trained large language model by using a training data set enhanced with multiple tasks, solving the problem of insufficient professionalism and complexity of existing large language models in dealing with specific domain problems. On the one hand, the knowledge-answering model can generate more accurate predicted answer results according to the question to be answered. On the other hand, through the enhanced training set integrating multiple tasks, the knowledge-answering model has the multi-task understanding ability and can generate corresponding predicted answer results for different tasks.
[0087] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0088] Please refer to Figure 7 , which shows a block diagram of a training device for a knowledge-answering model provided by an embodiment of the present application. The device has the function of implementing the above-mentioned training method of the knowledge-answering model, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be a computer device or can be set in a computer device. The device 700 may include: an acquisition module 710, a obtaining module 720, and an adjustment module 730.
[0089] The acquisition module 710 is configured to, for each task among multiple tasks, acquire the training data set of the task, where the training data set of the task includes multiple questions related to the task and the annotated answer results corresponding to the multiple questions respectively.
[0090] The obtaining module 720 is configured to perform knowledge enhancement on the training data set of the task to obtain the training data set after the task enhancement, where the knowledge enhancement is used to increase the number of the questions included in the training data set of the task.
[0091] The adjustment module 730 is configured to adjust a pre-trained large language model by using the training data sets after enhancement of the multiple tasks respectively, and use the adjusted large language model as a knowledge-answering model, and the knowledge-answering model is configured to generate a predicted answer result corresponding to the question to be answered according to the question to be answered.
[0092] In some embodiments, the obtaining module 720 includes: a selection unit, a first generation unit, a determination unit, and an addition unit ( Figure 7 not shown in the figure).
[0093] The selection unit is configured to select at least one target question from the multiple questions related to the task included in the training data set of the task.
[0094] A first generation unit for generating a similar question corresponding to the target question, where the similar question has the same or similar semantics as the target question but different expressions.
[0095] A determination unit for determining the labeled reply result corresponding to the target question as the labeled reply result corresponding to the similar question.
[0096] An addition unit for adding the similar question and the labeled reply result corresponding to the similar question to the training data set of the task to obtain the enhanced training data set of the task.
[0097] In some embodiments, the first generation unit is configured to generate a similar question corresponding to the target question through a pre-trained first large language model.
[0098] In some embodiments, the adjustment module 730 includes: a second generation unit, an input unit, and an adjustment unit ( Figure 7 not shown).
[0099] The second generation unit is configured to, for each of the multiple tasks, generate prompt information for the task according to the enhanced training data set of the task, where the prompt information is used to indicate selecting a reply result corresponding to a question related to the task from the knowledge set of the task, and the knowledge set of the task includes labeled reply results corresponding to multiple questions related to the task.
[0100] The input unit is configured to input the prompt information of the task into the pre-trained large language model, and output a predicted reply result corresponding to a question related to the task through the pre-trained large language model.
[0101] The adjustment unit is configured to adjust the pre-trained large language model according to the labeled reply results and predicted reply results corresponding to the questions respectively related to the multiple tasks, and use the adjusted large language model as the knowledge question and answer model.
[0102] In some embodiments, the second generation unit is configured to: generate the knowledge set of the task according to the labeled reply results corresponding to multiple questions related to the task; generate the prompt information of the task according to a prompt template, the questions included in the enhanced training data set of the task, and the knowledge set of the task, where the multiple tasks share the same prompt template, and the prompt template is used to define the format of the prompt information.
[0103] In some embodiments, the obtaining module 710 is configured to: collect at least one first question related to the task from the online log data related to the task, and generate an annotated reply result corresponding to each of the at least one first question through a second large language model; construct at least one second question related to the task according to predefined rules, and generate an annotated reply result corresponding to each of the at least one second question through a third large language model, where the predefined rules are used to specify the conditions satisfied by the questions related to the task; and obtain a training data set for the task according to the at least one first question and the annotated reply result corresponding to each of the at least one first question, and the at least one second question and the annotated reply result corresponding to each of the at least one second question.
[0104] The technical solution provided by the embodiments of the present application solves the problems of insufficient professionalism and complexity of existing large language models in dealing with specific domain problems by adjusting a pre-trained large language model with a training data set enhanced by multiple tasks. On the one hand, the knowledge-based question answering model can generate a more accurate predicted reply result according to the question to be replied. On the other hand, through the enhanced training set integrating multiple tasks, the knowledge-based question answering model has the ability to understand multiple tasks and can generate corresponding predicted reply results for different tasks.
[0105] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0106] Please refer to Figure 8 , which shows a structural block diagram of a computer device 800 provided by an embodiment of the present application.
[0107] Generally, the computer device 800 includes a processor 810 and a memory 820.
[0108] The processor 810 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 810 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 810 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 810 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 810 may further include an AI processor, which is used to process computational operations related to machine learning.
[0109] The memory 820 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 820 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 820 is used to store a computer program, and the computer program is configured to be executed by one or more processors to implement the training method of the above knowledge-answering model.
[0110] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the computer device 800, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0111] In some embodiments, a computer-readable storage medium is further provided. A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the training method of the above knowledge-answering model.
[0112] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0113] In some embodiments, a computer program product is also provided. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor reads and executes the computer program to implement the training method of the above knowledge Q&A model.
[0114] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0115] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for a knowledge Q&A model, characterized in that, The method includes: For each of multiple tasks, obtaining the training data set of the task, where the training data set of the task includes multiple questions related to the task, and the labeled reply results corresponding to each of the multiple questions; Performing knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task, where the knowledge enhancement is used to increase the number of questions included in the training data set of the task; Using the enhanced training data sets of the multiple tasks respectively to adjust a pre-trained large language model, and using the adjusted large language model as a knowledge Q&A model, where the knowledge Q&A model is used to generate a predicted reply result corresponding to a question to be replied according to the question to be replied; 2. The method according to claim 1, wherein The performing knowledge enhancement on the training data set of the task to obtain the enhanced training data set of the task includes: Selecting at least one target question from the multiple questions related to the task included in the training data set of the task; Generating a similar question corresponding to the target question, where the similar question has the same or similar semantics as the target question but different expressions; Determining the labeled reply result corresponding to the target question as the labeled reply result corresponding to the similar question; Adding the similar question and the labeled reply result corresponding to the similar question to the training data set of the task to obtain the enhanced training data set of the task.
3. The method according to claim 2, wherein The generating a similar question corresponding to the target question includes: Generating a similar question corresponding to the target question through a pre-trained first large language model.
4. The method according to claim 1, characterized in that, The using the enhanced training data sets of the multiple tasks respectively to adjust a pre-trained large language model and using the adjusted large language model as a knowledge Q&A model includes: For each of the multiple tasks, generating prompt information for the task according to the enhanced training data set of the task, where the prompt information is used to indicate selecting a reply result corresponding to a question related to the task from the knowledge set of the task, and the knowledge set of the task includes the labeled reply results corresponding to the multiple questions related to the task; Inputting the prompt information of the task into the pre-trained large language model, and outputting a predicted reply result corresponding to a question related to the task through the pre-trained large language model; Adjusting the pre-trained large language model according to the labeled reply results and predicted reply results corresponding to the questions respectively related to the multiple tasks, and using the adjusted large language model as the knowledge Q&A model.
5. The method according to claim 4, wherein The generating prompt information for the task according to the enhanced training data set of the task includes: Generating the knowledge set of the task according to the labeled reply results corresponding to the multiple questions related to the task; Generating the prompt information for the task according to a prompt template, the questions included in the enhanced training data set of the task, and the knowledge set of the task, where the multiple tasks share the same prompt template, and the prompt template is used to define the format of the prompt information.
6. The method according to claim 1, characterized in that, Obtaining the training data set for the task includes: Collecting at least one first question related to the task from the online log data related to the task, and generating an annotated reply result corresponding to each of the at least one first question through a second large language model; Constructing at least one second question related to the task according to predefined rules, and generating an annotated reply result corresponding to each of the at least one second question through a third large language model, where the predefined rules are used to specify the conditions satisfied by the questions related to the task; Obtaining the training data set for the task according to the at least one first question and the annotated reply result corresponding to each of the at least one first question, and the at least one second question and the annotated reply result corresponding to each of the at least one second question.
7. A training device for a knowledge Q&A model, characterized in that, The device includes: An obtaining module, configured to obtain the training data set for each of multiple tasks, where the training data set for the task includes multiple questions related to the task, and the annotated reply results corresponding to each of the multiple questions; A obtaining module, configured to perform knowledge enhancement on the training data set for the task to obtain the enhanced training data set for the task, where the knowledge enhancement is used to increase the number of questions included in the training data set for the task; An adjustment module, configured to adjust a pre-trained large language model by using the enhanced training data sets for the multiple tasks respectively, and use the adjusted large language model as a knowledge question and answer model, where the knowledge question and answer model is used to generate a predicted reply result corresponding to the question to be replied according to the question to be replied.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and the processor reads and executes the computer program from the computer-readable storage medium to implement the method according to any one of claims 1 to 6.
Citation Information
Cited By
Large language model data set protection method and device based on pollution lexical elements, electronic equipment, readable storage medium and computer program product
CN120523965A