Question and answer model training method, question processing method, question and answer model training system and related products
By adopting the training method of the question-and-answer model in the intelligent customer service system, using vector library and user feedback optimization model, the shortcomings of the existing intelligent customer service system in semantic understanding and intent recognition are solved, and a more accurate and personalized customer service response is achieved.
Patent Information
- Application Number
- CN202510112453.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
When handling customer service requests, the existing intelligent customer service system has limited semantic understanding ability and cannot deeply understand the customer's actual intentions. Especially when the customer's expression is not clear enough or uses synonyms, the system is prone to misunderstanding the customer's intentions and cannot correctly respond to customer's needs.
The training method of the question-answer model is adopted to obtain sample questions and answer corpus in the same knowledge field as the user consulting questions, build a vector library, and use these samples to train the initial big model to obtain a vertical question-and-answer model. The model generates reply content based on user consultation questions and related embedding vectors, and retrains the model through user feedback to optimize the accuracy of generated answers.
It improves the model's ability to understand specific domain knowledge and the accuracy of generating answers, reduces the fuzzy matching errors that traditional keyword matching may cause, and achieves a more personalized and accurate answer service.
Smart Images

Figure CN120011517A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method for a question-answering model, a problem-solving method, a system, and related products. Background Art
[0002] With the development of information technology, intelligent customer service is widely used in all walks of life and is loved by the public. When processing customer service requests, existing intelligent customer service mainly relies on predefined rules and keyword libraries, and uses simple text matching algorithms to identify user input and provide corresponding answers.
[0003] However, the intelligent customer service system that uses this keyword matching mode has limited semantic understanding capabilities and is unable to deeply understand the customer's actual intentions. Especially when the customer's expression is not clear enough or synonyms are used, the system is prone to misunderstand the customer's intentions and cannot correctly respond to the customer's needs. Summary of the invention
[0004] The embodiments of the present application provide a question-answering model training method, a question processing method, a system and related products for better understanding the user's input intention and providing the user with more personalized answering services and support.
[0005] A first aspect of an embodiment of the present application provides a method for training a question-answering model, comprising:
[0006] Obtaining various sample questions and answer corpora corresponding to the sample questions; the sample questions, the answer corpora and the user consultation questions have the same knowledge domain, and the embedded vectors obtained by encoding the sample questions and the answer corpora are used to construct a vector library of the knowledge domain;
[0007] Using each of the sample questions and each of the answer corpus to train the initial large model, to obtain a vertical domain question-answering model for the knowledge domain;
[0008] The vertical question-answering model is used to generate a reply content regarding the user consultation question based on the user consultation question and the relevant embedding vector; the relevant embedding vector is at least one embedding vector in the vector library that has a high similarity to the original embedding vector, and the original embedding vector is used to correspond to and represent the user consultation question.
[0009] Optionally, after obtaining the vertical domain question-answering model for the knowledge domain, the training method further includes:
[0010] receiving feedback information from the user on the reply content;
[0011] The feedback information and the reply content are combined into prompt words and input into the vertical domain question and answer model, so that the vertical domain question and answer model is retrained according to each of the sample questions, each of the reply corpus and the prompt words, and an updated vertical domain question and answer model is obtained for generating new reply content.
[0012] Optionally, the process of retraining the vertical domain question-answering model according to each of the sample questions, each of the answer corpus and the prompt word includes:
[0013] If the feedback information is positive feedback information, increase the usage weight of the reply content in the retraining process;
[0014] If the feedback information is negative feedback information, the usage weight of the reply content in the retraining process is reduced.
[0015] Optionally, the process of retraining the vertical domain question-answering model according to each of the sample questions, each of the answer corpus and the prompt word includes:
[0016] Constructing a policy network and a value network based on the vertical question-answering model; the policy network is used to reflect the decision-making performance of the vertical question-answering model on the question answering task, and the value network is used to evaluate the value obtained by the vertical question-answering model in executing the question answering task according to the decision;
[0017] The sample questions, the answer corpora and the prompt words are used to update the network parameters of the strategy network and the value network, and the updated strategy network and value network are used to construct a new vertical domain question-answering model.
[0018] A second aspect of the embodiment of the present application provides a problem handling method, including:
[0019] Obtaining the original embedding vector corresponding to the user's consultation question;
[0020] Performing a similarity comparison between the original embedding vector and each embedding vector in the vector library, so as to query from the vector library at least one embedding vector having a high similarity to the original embedding vector as a relevant embedding vector;
[0021] The relevant embedding vector and the user consultation question are input into a vertical domain question and answer model trained for the knowledge domain of the user consultation question to generate a reply content to the user consultation question; the vertical domain question and answer model is trained according to the training method described in any one of the first aspects above.
[0022] Optionally, if the user provides modification suggestions for the reply content, the problem handling method further includes:
[0023] The modification suggestions are input into the vertical domain question-answering model as prompt words and the user consultation questions to regenerate new reply content.
[0024] A third aspect of an embodiment of the present application provides a question-answering model training system, including: a first acquisition unit, a first processing unit;
[0025] The first acquisition unit is used to acquire various sample questions and answer corpora corresponding to the sample questions; the sample questions, the answer corpora and the user consultation questions have the same knowledge domain, and the embedding vectors obtained by encoding the sample questions and the answer corpora are used to construct a vector library of the knowledge domain;
[0026] The first processing unit is used to train the initial large model using each of the sample questions and each of the answer corpus to obtain a vertical domain question-answering model for the knowledge domain;
[0027] The vertical question-answering model is used to generate a reply content regarding the user consultation question based on the user consultation question and the relevant embedding vector; the relevant embedding vector is at least one embedding vector in the vector library that has a high similarity to the original embedding vector, and the original embedding vector is used to correspond to and represent the user consultation question.
[0028] A fourth aspect of the present application provides a specific embodiment of a problem handling system, the system comprising: a second acquisition unit, a second processing unit;
[0029] The second acquisition unit is used to obtain an original embedding vector corresponding to the user's consultation question;
[0030] The second processing unit is used to perform similarity comparison between the original embedding vector and each embedding vector in the vector library, so as to query from the vector library at least one embedding vector having a high similarity with the original embedding vector as a relevant embedding vector;
[0031] The second processing unit is also used to input relevant embedding vectors and user consultation questions into a vertical domain question and answer model trained for the knowledge domain of the user consultation questions to generate reply content to the user consultation questions; the vertical domain question and answer model is trained according to the training method described in the aforementioned first aspect or any specific method embodiment of the first aspect.
[0032] A fifth aspect of the embodiments of the present application provides an electronic device, including: a processor and a memory;
[0033] The processor is configured to communicate with the memory and execute instructions in the memory to implement the method described in the first aspect of the embodiment of the present application or any specific implementation of the first aspect.
[0034] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, they implement the method described in the first aspect of the embodiment of the present application or any specific implementation of the first aspect.
[0035] A seventh aspect of an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, they implement the method described in the first aspect of the embodiment of the present application or any specific implementation of the first aspect.
[0036] It can be seen from the above technical solutions that the embodiments of the present application have at least the following advantages:
[0037] The embodiment of the present application uses sample questions and answer corpus of the same knowledge domain as the user's consultation questions as training samples for the model, which can improve the model's ability to understand the knowledge in this field and enhance the model's generation ability. Using relevant embedded vectors in the vector library as a private domain knowledge base to assist in generating replies can reduce fuzzy matching errors that may be caused by traditional keyword matching; combining the two methods of vector retrieval and generation in this way can realize the retrieval augmented generation (RAG) framework, which effectively guarantees the generation of reasonable and accurate replies without sample precedents and improves user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0039] It should be noted that, although the steps in the process diagrams (if any) involved in the embodiments are drawn in sequence as indicated by the arrows, unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0040] Figure 1 A flowchart of a method for training a question-answering model according to an embodiment of the present application;
[0041] Figure 2 A flowchart of a method for solving a problem in an embodiment of the present application;
[0042] Figure 3 Another flowchart of the problem solving method of the embodiment of the present application;
[0043] Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0045] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0046] In the following description, similar expressions such as "a specific implementation" or "a specific example" are involved, which describe a subset of all possible embodiments, but it can be understood that "a specific implementation" or "a specific example" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. In the following description, the term multiple refers to at least two. In some specific examples, the numerical value mentioned in this application reaches the threshold value (if any), which may include the case where the former is greater than the latter of the threshold value; if similar expressions such as "any" or "at least one" are mentioned, it can specifically refer to any one of the listed examples or any combination of these examples.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0048] See also Figure 1 In a first aspect, the present application provides a specific embodiment of a method for training a question-answering model, and the embodiment includes the following steps:
[0049] Step S11, obtaining various sample questions and the answer corpus corresponding to each sample question;
[0050] In order to ensure that the model can better understand the knowledge domain content of the user's consultation questions, enable the model to better adapt to the needs of specific tasks, and improve the accuracy of the model's responses, the sample questions and response corpus can be designed to have the same knowledge domain as the user's consultation questions. This response corpus can be the response content corresponding to the question (i.e. the answer sentence itself), or it can be a reference material for data items such as documents, images, audio, and video referenced by any of the question and answer pairs. In some examples, the sample questions and response corpus can be cleaned and divided into blocks to obtain fine-grained, more useful and effective sample knowledge to optimize the initial large model. For details, see the following instructions, which are not limited here.
[0051] The embedding vectors obtained by encoding each sample question and each answer corpus are used to construct a vector library of knowledge domains, where the embedding vectors can express the essence of data items such as sample questions and their answer corpus. This vector library can be called a private domain knowledge base corresponding to the knowledge domain to which the user's consultation question belongs. In some examples, this knowledge domain can be the financial field, the medical field, the network security field, etc.
[0052] Step S12: Use various sample questions and answer corpora to train the initial large model to obtain a vertical domain question-answering model for the knowledge domain;
[0053] In actual situations, the initial large model can be pre-trained or fine-tuned (i.e., SFT) using QA question-answer pairs and document segmentation data to obtain a vertical domain question-answering model (also called a vertical domain large model) for a specific question-answering task. Specifically, some frequently asked questions and answers (FQA, Frequently Questions Answer) pairs in the above knowledge domain can be used as training samples for the model.
[0054] In some examples, the trained vertical question-answering model can be used to generate answers to user consultation questions based on user consultation questions and related embedding vectors. The details can be seen in the following description and will not be repeated here. This related embedding vector is at least one embedding vector in the vector library that has a high similarity with the original embedding vector, and the original embedding vector is used to represent the user consultation question; the process of finding the related embedding vector from the vector library can be called the recall reordering process.
[0055] In summary, the embodiment of the present application uses sample questions and answer corpus of the same knowledge domain as the user's consultation questions as training samples for the model, which can improve the model's ability to understand the knowledge in this field and enhance the model's generation ability. Using relevant embedded vectors in the vector library as a private domain knowledge base to assist in generating replies can reduce fuzzy matching errors that may be caused by traditional keyword matching; combining the two methods of vector retrieval and generation in this way can realize the Retrieval Augmented Generation (RAG) framework, which effectively guarantees the generation of reasonable and accurate replies without sample precedents and improves user experience.
[0056] Based on the above example descriptions, the method of the present application will be further described in detail below, and some specific possible implementation examples will be provided. In actual applications, the implementation contents between these examples can be combined or implemented separately as needed according to the corresponding functional principles and application logic. If implemented in combination, the execution order between the combined examples can be determined according to their respective processing logics, which may be determined by the actual scenario.
[0057] In some specific examples, after step S12, the training method of the embodiment of the present application may also include (enhancing the model generation capability through a customer feedback mechanism): receiving user feedback information on the reply content; combining the feedback information and the reply content into prompt words and inputting them into the vertical domain question and answer model, so that the vertical domain question and answer model is retrained based on each sample question, each reply corpus and the prompt words, and an updated vertical domain question and answer model is obtained for generating new reply content.
[0058] Exemplarily, the feedback information may be at least one of whether the user clicks on the reply content, the click frequency, the dwell time on the reply content, the satisfaction, modification suggestions, etc.
[0059] Generally, you can use the Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) strategy to enhance the model generation capability. Specifically, you can determine whether to use the RLHF strategy or the DPO strategy to further train the vertical domain question-answering model based on actual conditions.
[0060] RLHF refers to a method of training models through feedback provided by humans. This method usually involves the process of interacting with humans. By letting humans evaluate and score the output generated by the model, the model training and optimization are guided. Unlike supervised learning and unsupervised learning, RLHF pays more attention to human participation and feedback, so that the model can better adapt to human intentions and needs. In the RLHF stage, some reinforcement learning algorithms and techniques such as Q-learning and SARSA are usually used to optimize the performance of the model and improve the user experience. The training process of RLHF usually includes three core steps: generating samples and collecting human feedback with multiple strategies; training the reward model; training the reinforcement learning strategy and fine-tuning the language model. In RLHF, the model generates some candidate texts and obtains feedback through interaction with human users. These feedbacks are converted into numerical rewards by the reward model as signals for model training; in some examples, negative numerical rewards can be regarded as penalties. By continuously iterating this process, the model can gradually improve the quality of its generated texts to better meet the needs of human users.
[0061] DPO is a direct optimization strategy based on human preferences, which is used to train language models to better meet user needs or preferences. The core idea is to directly learn human preference data by minimizing the loss function without relying on a complex reinforcement learning framework. DPO uses a human preference dataset to construct training samples, which contains "preferred" and "non-preferred" comparison items. By directly optimizing the preference loss function, the policy model is guided to learn the preference distribution. In contrast, DPO simplifies the RLHF (i.e., reinforcement learning from human feedback) process, making the optimization process relatively simple and avoiding the complexity of reward model training; DPO directly searches in the policy space to find the optimal policy, avoiding the value function estimation and policy iteration process in traditional reinforcement learning. In general, DPO can use optimization algorithms such as gradient descent to search for the optimal parameter configuration in the policy space, thereby gradually improving the performance of the model on specific tasks.
[0062] In some specific examples, further, a reward mechanism and a penalty mechanism can be introduced to enhance the model generation capability: if the feedback information is positive feedback information, the weight of the reply content in the retraining process is increased; if the feedback information is negative feedback information, the weight of the reply content in the retraining process is reduced. In other words, a reward and penalty mechanism can be constructed to optimize and update the model. Specifically, the weight of the reply can be adjusted according to the degree of match between the reply content and the user's expected answer, and then these weighted reply contents and question and answer corpus are input into the model for retraining to continuously optimize the model performance.
[0063] The above-mentioned positive and negative feedback may refer to any positive reflection such as the user clicking on the reply content, browsing the reply content for a long time, being satisfied with the reply content, etc. The positive and negative feedback information may refer to any one of the click frequency, user ID, browsing time, satisfaction score, etc.; negative feedback may be the opposite of positive feedback.
[0064] For example, when the reply content is accurate and reasonable, a positive reward value can be given. This reward value can be quantified according to the accuracy of the reply. For example, a reward mechanism function can be designed to quantify the quality of the recommendation behavior, such as the user's satisfaction and / or dwell time after clicking the reply. For another example, indicators such as accuracy and F1 score can be used to measure the degree of match between the reply content and the standard answer or manual evaluation. For particularly accurate or innovative replies, additional reward values can be added to encourage the model to generate higher quality replies.
[0065] When the response deviates significantly from the expected answer, a negative penalty can be given. This penalty can also be quantified based on the degree of deviation to ensure that the model can gradually correct its errors during the training process. For obviously wrong or misleading responses, a larger penalty can be set to prevent the model from generating similar wrong responses in the future.
[0066] Furthermore, the weight of the reply content can be adjusted according to the reward value or penalty value obtained by the reply content. Replies with higher reward values are given greater weights, while replies with higher penalty values are given smaller weights. The weighted reply content and the corresponding question and answer corpus are combined to construct a new training data set, which will be used to retrain the question and answer system model. This training process may involve adjusting model parameters, optimizing model structure, etc., to improve the accuracy and generalization ability of the model.
[0067] In some specific examples, the process of retraining the above-mentioned vertical question and answer model based on various sample questions, various response corpora and prompt words may specifically include (model retraining based on the PPO strategy): constructing a policy network and a value network based on the vertical question and answer model; the policy network is used to reflect the decision-making performance of the vertical question and answer model in the question answering task, and the value network is used to evaluate the value obtained by the vertical question and answer model in executing the question answering task according to the decision; using various sample questions, various response corpora and the above-mentioned prompt words to update the respective network parameters of the policy network and the value network, and using the updated policy network and value network to construct a new vertical question and answer model.
[0068] As described above, the implementation results of the embodiments of the present application can be called an intelligent customer service or user behavior real-time recommendation feedback algorithm based on the PPO strategy. For example, in each training cycle, the intelligent customer service system can interact with the user to collect user feedback (such as whether to click or not, the dwell time on the reply content) and other experience trajectory information, and use this information for further iterative training of the model. Among them, the network parameters such as the weights of the policy network and the value network can be updated by the PPO algorithm to maximize the long-term cumulative rewards of the vertical domain large model; this policy network and value network can be neural network models, the policy network can be used to predict the probability distribution of actions taken under a given state, and the value network can be used to estimate the value of the state after performing a certain action.
[0069] Specifically, the policy network and value network can be constructed on the basis of the vertical domain big model. These two networks are the core components of the PPO algorithm. They are closely connected with the vertical domain big model and together constitute the learning and decision-making capabilities of the vertical domain big model. After the parameter update is completed, the policy network and value network will jointly construct the components of the new vertical domain big model. In other words, the vertical domain big model can use the two networks of policy network and value network, through the training and optimization of the PPO algorithm, to have the ability to obtain higher rewards on specific tasks.
[0070] As can be seen from the above description, the training method of the embodiment of the present application can utilize the multimodal large model understanding ability, RAG retrieval enhancement and recall rearrangement technology, combined with the private domain knowledge base (i.e., vector library), to build a new intelligent customer service system (which can refer to the vertical domain question and answer model), which can achieve the following effects:
[0071] (1) Better understanding ability: Multimodal large models can process multiple types of data such as text, voice, and images, allowing the system to better understand the user's intentions and context.
[0072] (2) Better generation capability: RAG combines the recall rearrangement technology framework to enhance generation capability, which can provide more accurate and relevant responses and reduce the fuzzy matching errors that may be caused by traditional keyword matching.
[0073] (3) Strong adaptability: By combining multiple methods such as retrieval generation and user feedback, the model can learn additional knowledge without sample precedents, that is, it can ensure that there is support from relevant data, so that reasonable and accurate responses can be provided. This breaks the current situation where the existing intelligent customer service system relies on fixed knowledge bases and rules, and is often unable to flexibly respond to customer requests when faced with new or unexpected problems.
[0074] (4) Natural and smooth interaction: Multimodal processing capabilities and natural language generation technology make human-computer dialogue more natural, and users can communicate in a variety of ways, not just limited to text input.
[0075] (5) Personalized services: Private domain knowledge bases can be customized according to the specific business of corporate users, thereby providing users with more personalized services and support.
[0076] See also Figure 2 The second aspect of the present application provides a specific embodiment of a problem handling method, which includes the following steps:
[0077] Step S21, obtaining an original embedding vector corresponding to the user's consultation question;
[0078] Step S22: perform a similarity comparison between the original embedding vector and each embedding vector in the vector library, so as to query from the vector library at least one embedding vector having a high similarity with the original embedding vector as a relevant embedding vector;
[0079] The above-mentioned vector library contains the embedding vectors obtained by encoding each sample question and its answer corpus respectively; an association relationship (or index) can be established between each sample question and each answer corpus and their respective embedding vectors. This index can be regarded as a key-value pair. Taking the sample question as an example, a sample question can correspond to at least one embedding vector, that is, multiple value values. When there are multiple embedding vectors, these embedding vectors are used to jointly represent this sample question.
[0080] Specifically, the relevant embedding vectors (or information fragments) can be recalled through the similarity algorithm, and then the recalled information fragments can be rearranged according to factors such as the similarity size, so that the information fragments with higher relevance to the user's consultation questions are preferred as the basis for generating the reply content, making the model's response more accurate.
[0081] Step S23: input the relevant embedding vectors and user consultation questions into the vertical domain question-answering model trained for the knowledge domain of the user consultation questions to generate response content for the user consultation questions.
[0082] This vertical domain question-answering model is trained according to the training method described in the first aspect or any specific implementation of the first aspect. For details, please refer to the above description and will not be repeated here.
[0083] In some examples, if the user provides feedback on the reply content, the problem handling method of the embodiment of the present application may also include: inputting the modification opinions as prompt words and the user's inquiry question into the vertical domain question-answering model to regenerate new reply content. Quoting the user's feedback on the reply content can help the model understand the user's actual intention more deeply and improve the accuracy of the reply.
[0084] See also Figure 3 The following example illustrates the application process of this application. This process can be divided into: Step 1, organize the knowledge base and pre-train the model; Step 2, retrieve, recall and rearrange; Step 3, use the RLHF strategy and reasoning application. For details, see the following description.
[0085] 1. Clean and segment the original policy FQA and policy document materials to obtain QA question-answer pairs and document segmentation data.
[0086] The cleaning process may include at least one of removing invalid text content (wrong characters or redundant characters), filling in missing content (such as numerical units, etc.), and standardizing text content (such as unifying numerical types). The document segmentation process may include at least one of the following segmentation methods:
[0087] Segment by punctuation; Fixed length segmentation: segment the text into multiple blocks according to the number of characters or words in the text, for example, 500 characters per block; Sliding window segmentation: create an overlapping sliding window, for example, the window size is 500 and the step length is 100; Segmentation based on themes: segment by identifying the change points of the article topic; Segmentation based on semantic similarity: use a model to evaluate the semantic similarity between texts, and segment when the similarity drops below a certain threshold; Segmentation by document structure: if a segmentation tool is used, segment according to the document structure (such as abstract, paragraph, conclusion, etc.). Specifically, any of the above cleaning operations and segmentation methods can be selected according to the actual situation, and there is no limit here.
[0088] 2. Encode the QA question-answer pairs and document chunk data through the embedding model to obtain their respective embedding vectors;
[0089] 3. Create an index by combining the embedding vector with the QA question-answer pair and the document block data, and store it in the vector database (which can be called the enterprise private domain knowledge base);
[0090] 4. Use QA question-answer pairs and document segmentation data to pre-train or fine-tune the initial large model to obtain a vertical domain large model; this initial large model can be any large language model such as LLaMa, LaMDA, GPT, etc., which can be selected according to needs.
[0091] 5. Input the user consultation question into the embedding model to obtain its embedding vector, that is, the original embedding vector.
[0092] 6. Perform similarity queries between the embedding vectors of the user's inquiry question and the embedding vectors in the vector library, so as to recall n embedding vectors from the vector library as relevant fragments;
[0093] 7. After re-ranking the n recalled relevant segments, the top k relevant segments (i.e., relevant embedding vectors) with the highest ranking (e.g., the highest similarity) are obtained;
[0094] 8. Input the rearranged top k relevant segments and the original user inquiry question into the vertical domain big model to generate the answer content;
[0095] 9. Provide feedback on the content of the reply (such as click or not, dwell time, satisfaction, etc.), reward accurate and reasonable replies, and punish replies with large deviations, so as to continuously improve the accuracy of the reply content.
[0096] As described above, the embodiments of the present application have the following points and effects.
[0097] (1) Multimodal big model: Based on the open source big model, the big model is fine-tuned and trained using multiple forms of data such as private domain images, text, and voice, which can better understand and handle complex service scenarios.
[0098] (2) RAG framework: Combining retrieval technology with generation models, it helps generate more accurate and contextual answers by retrieving relevant information from the enterprise's private domain knowledge base.
[0099] (3) Recall and re-rank: Use the similarity model to recall information fragments, and then re-rank the recalled information fragments so that the knowledge that is more relevant to the customer's inquiry question is ranked at the front, making the model's response more accurate.
[0100] (4) Reinforcement learning strategy: Through customer click feedback data, use the reinforcement learning DPO and PPO strategies to continuously learn and optimize the model to improve the accuracy and efficiency of the system.
[0101] In the embodiment of the present application, the operation performed by the problem handling method is similar to the operation described in the first aspect or any specific method embodiment of the first aspect, and will not be described in detail here. Of course, the specific implementation process of each operation in the first aspect of the present application can also refer to the relevant description of the second aspect.
[0102] A third aspect of the present application provides a specific embodiment of a training system for a question-answering model, the system comprising: a first acquisition unit, a first processing unit;
[0103] The first acquisition unit is used to acquire various sample questions and the answer corpus corresponding to each sample question; the knowledge domains of the sample questions, the answer corpus and the user's consultation questions are the same, and the embedding vectors obtained by encoding each sample question and each answer corpus are used to construct a vector library of the knowledge domain;
[0104] The first processing unit is used to train the initial large model using various sample questions and various answer corpora to obtain a vertical domain question-answering model for the knowledge domain;
[0105] The vertical question-answering model is used to generate reply content about user consultation questions based on user consultation questions and relevant embedding vectors; the relevant embedding vector is at least one embedding vector in the vector library that has a high similarity with the original embedding vector, and the original embedding vector is used to correspond to the user consultation question.
[0106] In some examples, the first processing unit is further configured to:
[0107] Receive user feedback on the reply content;
[0108] The feedback information and reply content are combined into prompt words and input into the vertical domain question and answer model, so that the vertical domain question and answer model is retrained according to each sample question, each reply corpus and prompt words, and the updated vertical domain question and answer model is obtained to generate new reply content.
[0109] In some examples, the first processing unit is specifically configured to:
[0110] If the feedback information is positive feedback information, increase the weight of the reply content in the retraining process;
[0111] If the feedback information is negative feedback information, reduce the weight of the reply content in the retraining process.
[0112] In some examples, the first processing unit is specifically configured to:
[0113] Construct a policy network and a value network based on the vertical question-answering model; the policy network is used to reflect the decision-making performance of the vertical question-answering model in the question answering task, and the value network is used to evaluate the value obtained by the vertical question-answering model in executing the question answering task according to the decision;
[0114] Use various sample questions, answer corpora and prompt words to update the network parameters of the strategy network and the value network, and use the updated strategy network and value network to build a new vertical domain question and answer model.
[0115] In the embodiment of the present application, the operations performed by each unit of the training system of the question-answering model are similar to the operations described in the aforementioned first aspect or any specific method embodiment of the first aspect, and the details will not be repeated here.
[0116] A fourth aspect of the present application provides a specific embodiment of a problem handling system, the system comprising: a second acquisition unit, a second processing unit;
[0117] The second acquisition unit is used to obtain an original embedding vector corresponding to the user's consultation question;
[0118] The second processing unit is used to perform similarity comparison between the original embedding vector and each embedding vector in the vector library, so as to query from the vector library at least one embedding vector having a high similarity with the original embedding vector as a relevant embedding vector;
[0119] The second processing unit is also used to input relevant embedding vectors and user consultation questions into a vertical domain question and answer model trained for the knowledge domain of the user consultation questions to generate reply content to the user consultation questions; the vertical domain question and answer model is trained according to the training method described in the aforementioned first aspect or any specific method embodiment of the first aspect.
[0120] In some examples, the second processing unit is further configured to:
[0121] If the user provides modification suggestions for the reply content, the modification suggestions will be input into the vertical domain question-answering model as prompt words and the user's consultation questions to regenerate new reply content.
[0122] In the embodiment of the present application, the operations performed by each unit of the question and answer processing system are similar to the operations described in the aforementioned second aspect or any specific method embodiment of the second aspect, and will not be described in detail here.
[0123] See also Figure 4 The electronic device of the embodiment of the present application may include one or more processors (such as central processing units (CPU)) and a memory, in which one or more applications or data are stored.
[0124] The memory may be a volatile memory or a persistent memory. The program stored in the memory may include one or more modules, each of which may include a series of instruction operations in the electronic device. Furthermore, the processor may be configured to communicate with the memory and execute a series of instruction operations in the memory on the electronic device.
[0125] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input and output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.
[0126] The processor can execute the operations performed by the aforementioned first aspect or any specific method embodiment of the first aspect, and the details will not be repeated here.
[0127] The present application provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enable the computer to execute the method described in the first aspect or any specific implementation of the first aspect.
[0128] The present application provides a computer program product comprising instructions or a computer program. When the computer program product is run on a computer, the computer is enabled to execute the method described in the first aspect or any specific implementation of the first aspect.
[0129] It is understood that in various embodiments of the present application, the sequence number of each step does not mean the order of execution, and the execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The operation content added or refined in each example scheme of the above method, system or device (if any) does not necessarily have to be executed in the specific implementation. If more than two operations are added, these operations can be implemented in combination or separately, depending on the actual scenario.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system (if any) and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0131] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system or device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0132] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product (or computer program product) is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a business server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A training method for a question-answering model, characterized in that: include: Obtaining various sample questions and answer corpora corresponding to the sample questions; The sample questions, the answer corpus and the user consultation questions have the same knowledge domain, and the embedding vectors obtained by encoding each of the sample questions and each of the answer corpus are used to construct a vector library of the knowledge domain; Using each of the sample questions and each of the answer corpus to train the initial large model, to obtain a vertical domain question-answering model for the knowledge domain; The vertical question-answering model is used to generate a reply content regarding the user consultation question based on the user consultation question and the relevant embedding vector; the relevant embedding vector is at least one embedding vector in the vector library that has a high similarity to the original embedding vector, and the original embedding vector is used to correspond to and represent the user consultation question.
2. The training method of the question-answering model according to claim 1, characterized in that: After obtaining the vertical domain question-answering model for the knowledge domain, the training method further includes: receiving feedback information from the user on the reply content; The feedback information and the reply content are combined into prompt words and input into the vertical domain question and answer model, so that the vertical domain question and answer model is retrained according to each of the sample questions, each of the reply corpus and the prompt words, and an updated vertical domain question and answer model is obtained for generating new reply content.
3. The training method of the question-answering model according to claim 2, characterized in that: The process of retraining the vertical domain question answering model according to each of the sample questions, each of the answer corpus and the prompt word includes: If the feedback information is positive feedback information, increase the usage weight of the reply content in the retraining process; If the feedback information is negative feedback information, the usage weight of the reply content in the retraining process is reduced.
4. The training method of the question-answering model according to claim 2 or 3, characterized in that: The process of retraining the vertical domain question answering model according to each of the sample questions, each of the answer corpus and the prompt word includes: Constructing a policy network and a value network based on the vertical question-answering model; the policy network is used to reflect the decision-making performance of the vertical question-answering model on the question answering task, and the value network is used to evaluate the value obtained by the vertical question-answering model in executing the question answering task according to the decision; The sample questions, the answer corpora and the prompt words are used to update the network parameters of the strategy network and the value network, and the updated strategy network and value network are used to construct a new vertical domain question-answering model.
5. A problem solving method, characterized in that: include: Obtaining the original embedding vector corresponding to the user's consultation question; Performing a similarity comparison between the original embedding vector and each embedding vector in the vector library, so as to query from the vector library at least one embedding vector having a high similarity to the original embedding vector as a relevant embedding vector; Inputting the relevant embedding vector and the user consultation question into a vertical domain question answering model trained on the knowledge domain of the user consultation question to generate a reply content about the user consultation question; The vertical question-answering model is trained according to the training method described in any one of claims 1 to 4.
6. The problem solving method according to claim 5, characterized in that: If the user provides modification suggestions for the reply content, the problem handling method further includes: The modification suggestions are input into the vertical domain question-answering model as prompt words and the user consultation questions to regenerate new reply content.
7. A question-answering model training system, characterized in that: include: A first acquisition unit and a first processing unit; The first acquisition unit is used to acquire each sample question and the answer corpus corresponding to each sample question; The sample questions, the answer corpus and the user consultation questions have the same knowledge domain, and the embedding vectors obtained by encoding each of the sample questions and each of the answer corpus are used to construct a vector library of the knowledge domain; The first processing unit is used to train the initial large model using each of the sample questions and each of the answer corpus to obtain a vertical domain question-answering model for the knowledge domain; The vertical question-answering model is used to generate a reply content regarding the user consultation question based on the user consultation question and the relevant embedding vector; the relevant embedding vector is at least one embedding vector in the vector library that has a high similarity to the original embedding vector, and the original embedding vector is used to correspond to and represent the user consultation question.
8. An electronic device, characterized in that: include: Processor and memory; The processor is configured to communicate with the memory and execute instructions in the memory to implement the method of any one of claims 1 to 4 or 5 to 6.
9. A readable storage medium, characterized in that: The readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 4 or 5 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises computer instructions, which implement the method according to any one of claims 1 to 4 or 5 to 6 when executed by a processor.