Model training method and device, equipment, medium and program product
By text segmentation and data generation of knowledge documents, training the second search model and search library, the problem of poor answer retrieval performance in smart question-and-answer is solved, and higher answer retrieval accuracy and data quality are achieved.
Patent Information
- Application Number
- CN202311529043.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
AI Technical Summary
The existing search models have poor answer retrieval performance during the intelligent question and answer process, and the data quality of the search library is poor, resulting in the search model being unable to provide users with accurate and detailed answers.
By obtaining knowledge documents, text segmentation processing is performed to generate a collection of text blocks. Each text block consists of reference text blocks and general description information. Based on these text blocks, the set is generated, the first search model is trained to obtain the second search model and the search library with good data quality.
Improves the reliability and accuracy of smart Q&A, ensuring that the search model can provide users with accurate and meticulous answers.
Smart Images

Figure CN120045645A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, particularly to the field of artificial intelligence, and specifically to a model training method, a model training apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Intelligent question answering refers to the process of answering questions raised by users in accurate and concise natural language; it belongs to a direction in the field of Natural Language Processing (NLP) that has received much attention and has broad development prospects.
[0003] Currently, it is supported to use a retrieval model to retrieve answers from a retrieval library for user questions during the intelligent question answering process. Through practice, it is found that the answer retrieval performance of existing retrieval models is not good, and the data quality of the retrieval library is poor; this makes it that during the intelligent question answering process, the answers retrieved by the retrieval model are often hallucinations of the retrieval model, resulting in the retrieval model not being able to provide accurate and detailed answers to user questions. Summary of the Invention
[0004] Embodiments of this application provide a model training method, apparatus, device, medium, and program product, which can construct a second retrieval model with better retrieval performance and a retrieval library with better data quality, and improve the reliability and accuracy of intelligent question answering.
[0005] On the one hand, embodiments of this application provide a model training method, which includes:
[0006] Obtain knowledge documents, and perform text segmentation processing on the knowledge documents to obtain a set of text blocks; the set of text blocks includes multiple text blocks, each text block has a theme respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by generalizing the text semantics of the corresponding reference text block;
[0007] Based on the set of text blocks and the summary information included in each text block, generate a set of data pairs; the set of data pairs includes multiple data pairs, each data pair is composed of a matching question text and a text block; the question text in each data pair is written as a question for the text block that matches it, and the text block in each data pair serves as the answer source for the question text that matches it;
[0008] The first retrieval model is trained using a set of data pairs to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors of each text block in the text block set; wherein, the second retrieval model and the retrieval library are used to generate a target answer for the target question text in an interactive dialogue scenario.
[0009] On the other hand, an embodiment of the present application provides a model training device, and the device includes:
[0010] An acquisition unit, configured to acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set; the text block set includes multiple text blocks, each text block has a theme respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by generalizing the text semantics of the corresponding reference text block;
[0011] A processing unit, configured to generate a set of data pairs based on the text block set and the summary information included in each text block; the set of data pairs includes multiple data pairs, and each data pair is composed of a matching question text and a text block; the question text in each data pair is written as a question for the text block it matches, and the text block in each data pair serves as the answer source for the question text it matches;
[0012] The processing unit is further configured to train the first retrieval model using the set of data pairs to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors of each text block in the text block set; wherein, the second retrieval model and the retrieval library are used to generate a target answer for the target question text in an interactive dialogue scenario.
[0013] In one implementation, the number of knowledge documents is at least one; when the processing unit performs text segmentation processing on the knowledge document to obtain a text block set, it is specifically configured to:
[0014] Obtain the semantic information of the knowledge document, and perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document;
[0015] By performing semantic generalization on the text semantics of each reference text block in one or more reference text blocks, obtain the summary information corresponding to each reference text block;
[0016] Add the summary information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block form a text block set.
[0017] In one implementation, the semantic information of the knowledge document includes the theme to which the knowledge document belongs, and the knowledge document is composed of one or more characters; the processing unit is specifically used for performing segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document, and specifically includes:
[0018] Judge whether the theme to which the knowledge document belongs is single;
[0019] If the theme to which the knowledge document belongs is single, count the number of characters included in the knowledge document;
[0020] When the number of characters included in the knowledge document is less than the character number threshold, use the knowledge document as a reference text block; or,
[0021] When the number of characters included in the knowledge document is greater than or equal to the character number threshold, perform paragraph cutting processing on the knowledge document according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
[0022] In one implementation, the processing unit is further used for:
[0023] If the knowledge document includes at least two themes, perform hierarchical expression segmentation on the knowledge document to obtain at least two initial text blocks; each of the at least two initial text blocks has one theme;
[0024] Count the number of characters included in each of the at least two initial text blocks, and use the initial text block with the number of characters less than the character number threshold as a reference text block; and,
[0025] Perform paragraph cutting processing on the initial text block with the number of characters greater than or equal to the character number threshold according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
[0026] In one implementation, the text logical relationships are respectively the general-part relationship and the parallel relationship; the general-part relationship indicates that the content structure of the text content is the general-part structure, and the parallel relationship indicates that the content structure of the text content is the parallel structure; the knowledge document or the initial text block to which the paragraph cutting processing is performed is represented as the text content; the paragraph cutting processing includes:
[0027] Perform paragraph cutting processing on the text content according to the text logical relationship of the text content to obtain the reference text block corresponding to the knowledge document;
[0028] Among them, when the text logical relationship of the text content is a general - specific relationship, the text content is cut into a general description text block and one or more sub - description text blocks; the general description text block is the content with a generalizing effect in the text content, and the general information corresponding to the general description text block is used to indicate the overall semantics expressed by the text content; the sub - description text block is the content that explains the general description text block in the text content, and the general information corresponding to the sub - description text block is used to indicate the sub - semantics expressed by the sub - description text block and the overall semantics expressed by the text content;
[0029] When the text logical relationship is a parallel relationship, the text content is cut into at least two sub - description text blocks.
[0030] In one implementation, the processing unit is further configured to:
[0031] Extract metadata for each text block in the text block set to obtain the metadata of each text block; the metadata of the text block is used to describe the data attributes of the text block, and the data attributes include: the retrieval domain to which the text block belongs, the source of the text block, and the theme to which the text block belongs;
[0032] Among them, the metadata of each text block is used to update the retrieval library and perform index retrieval for the text block; the update includes: when the data attributes of the text block in the retrieval library are updated, modifying the attribute value of the corresponding data attribute in the metadata of the text block;
[0033] The index retrieval includes at least: in an interactive dialogue scenario, according to the retrieval requirements indicated by the target question text, retrieving text blocks whose metadata meets the retrieval requirements from the metadata to generate an answer for the target question text; and, when any text block has been retrieved in the interactive dialogue scenario, retrieving text blocks whose metadata is the same as the metadata of any text block from the metadata to generate an answer for the target question text.
[0034] In one implementation, when the processing unit is used to generate a data pair set based on the text block set and the general information included in each text block, it is specifically configured to:
[0035] Obtain the scenario information corresponding to the interactive dialogue scenario, where the scenario information includes the dialogue style, interaction mode, and interaction domain; and,
[0036] Based on the scenario information, the text block set, and the general information included in each text block, determine the question style and the number of questions suitable for each text block;
[0037] According to the question style and the number of questions suitable for each text block, write questions for each text block to obtain one or more question texts corresponding to each text block;
[0038] Combine one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block;
[0039] Generate a set of data pairs based on one or more data pairs corresponding to each text block.
[0040] In one implementation, when the processing unit is used to determine the question style and the number of questions adapted to each text block based on the scenario information, the text block set, and the general description information included in each text block, it is specifically used for:
[0041] Mine key elements for each text block according to the text block set and the general description information included in each text block to obtain the key elements of each text block;
[0042] Determine the question style and the number of questions adapted to each text block based on the scenario information and the key elements of each text block;
[0043] Among them, the question style adapted to each text block is one or more, and each question style belongs to one of multiple dialogue styles; the dialogue styles include at least one of the following: factual style, explanatory style, reasoning style, evaluative style, and hypothetical style.
[0044] In one implementation, any text block in the text block set is represented as a target text block, and any data pair corresponding to the target text block is represented as a target data pair. The target data pair includes a target question text that matches the target text block. When the processing unit is used to generate a set of data pairs based on one or more data pairs corresponding to each text block, it is specifically used for:
[0045] Use a key-value verification model to extract, from the target text block, a first answer that matches the target question text based on the target question text, and the position information of the first answer in the target text block;
[0046] Extract a second answer indicated by the position information from the target text block according to the position information of the first answer in the target text block;
[0047] Compare the first answer and the second answer, and add the target data pair to the set of data pairs based on the comparison result;
[0048] Among them, if the comparison result indicates that the first answer is the same as the second answer, the target text block is added to the set of data pairs; if the comparison result indicates that the first answer is different from the second answer, the target text block is not added to the set of data pairs.
[0049] In one implementation, the set of text blocks includes candidate text blocks. A candidate text block refers to a text block whose format conforms to the Q&A structure. The candidate text block includes a question part and an answer part. The processing unit is further configured to:
[0050] Add the candidate text block to the set of data pairs; and,
[0051] Perform question generalization processing on the question part included in the candidate text block to generate a generalized question text, and form a new data pair corresponding to the candidate text block with the generalized question text and the answer part included in the candidate text block, and add the new data pair to the set of data pairs.
[0052] In one implementation, when the processing unit is configured to train the first retrieval model using the set of data pairs to obtain the second retrieval model, it is specifically configured to:
[0053] Perform data allocation processing on the set of data pairs according to the data allocation strategy to obtain a first target set and a second target set; the data allocation strategy includes an answer stratification strategy and a word embedding classification strategy; the first target set is the validation set, and the second target set is the training set; or, the first target set is the training set, and the second target set is the validation set;
[0054] Train the first retrieval model using the training set to obtain a trained first retrieval model;
[0055] Test the trained first retrieval model using the validation set to obtain the second retrieval model.
[0056] In one implementation, the data allocation strategy is the answer stratification strategy; any text block in the set of text blocks is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero. When the processing unit is configured to perform data allocation processing on the set of data pairs according to the data allocation strategy to obtain a first target set and a second target set, it is specifically configured to:
[0057] Determine the position information of the answer corresponding to the question text in each target data pair in the target text block;
[0058] Perform stratification processing on the target text block according to the position information of the answer corresponding to the question text in each target data pair in the target text block to obtain multiple text sub-layers corresponding to the target text block; one text sub-layer corresponds to at least one of the Q target data pairs;
[0059] Select reference data pairs from the target data pairs corresponding to each text sub-layer among the multiple text sub-layers and add them to the first target set, and add the target data pairs other than the selected reference data pairs among the multiple text sub-layers to the second target set.
[0060] In one implementation, the data allocation strategy is a word embedding classification strategy; any text block in the text block set is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero, and the question text in the target data pair consists of one or more characters; the processing unit is specifically used for performing data allocation processing on the data pair set according to the data allocation strategy to obtain a first target set and a second target set, and specifically includes:
[0061] Perform word vector representation on the Q question texts in the Q target data pairs respectively to obtain the word vectors corresponding to each question text in the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts;
[0062] Perform clustering processing on the word vectors corresponding to the Q question texts to obtain one or more clustering groups; one clustering group includes the target data pairs corresponding to one or more question texts whose vector distances meet the distance requirements;
[0063] Select reference data pairs from one or more clustering groups and add them to the first target set, and add the target data pairs in one or more clustering groups except the selected reference data pairs to the second target set.
[0064] In one implementation, the retrieval library includes the feature vectors of P text blocks, where P is a positive integer and P is less than Q; the processing unit is further used for:
[0065] Receive the target question text input by the object in the interactive dialogue scenario, and perform vector embedding processing on the target question text using the second retrieval model to obtain the feature vector of the target question text; the feature vector of the target question text is used to represent the semantic information of the target question text;
[0066] Perform similarity calculation processing on the feature vector of the target question text and the feature vectors of the P text blocks in the retrieval library to obtain the matching scores between each of the P text blocks and the target question text; the matching score is used to indicate the credibility of the text block as the answer source of the target question text;
[0067] Sort the P matching scores in descending order, and perform gradient operation on each matching score in the sorted score sequence to obtain a plurality of gradient information;
[0068] Based on the plurality of gradient information, dynamically screen one or more matching scores from the score sequence whose gradient information meets the gradient descent condition, and use the text blocks corresponding to the one or more matching scores as the answer source of the target question text;
[0069] Generate a target answer for the target question text based on the text blocks corresponding to the one or more matching scores.
[0070] In one implementation, when the processing unit generates a target answer for the target question text based on text blocks corresponding to one or more matching scores, it is specifically configured to:
[0071] Perform task classification processing on the target question text to obtain a classification result, where the classification result is used to indicate the difficulty of answering the target question text;
[0072] According to the difficulty of answering indicated by the classification result, use an answer generation model that matches the difficulty of the answer to generate an initial answer to the target question text based on text blocks corresponding to one or more matching scores;
[0073] Perform answer review processing on the initial answer to obtain the target answer to the target question text; the target answer is the initial answer or the answer after the initial answer is fine-tuned.
[0074] In one implementation, the interactive dialogue scenario includes multiple rounds of interactive dialogue. Any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogue is represented as the target round of interactive dialogue; the processing unit is further configured to:
[0075] Obtain the target question text input by the object in the target round of interactive dialogue, the historical dialogue data in the historical round of interactive dialogue, and the historical object data about the object; the historical object data is generated based on the historical dialogue data;
[0076] Rewrite the target text question based on the target question text, the historical dialogue data, and the historical object data to obtain a new target question text;
[0077] When the processing unit performs vector embedding processing on the target question text using the second retrieval model, it is specifically configured to:
[0078] Perform vector embedding processing on the new target question text using the second retrieval model.
[0079] On the other hand, an embodiment of the present application provides a computer device, which includes:
[0080] A processor for loading and executing a computer program;
[0081] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, the above-mentioned model training method is implemented.
[0082] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded and executed by the processor to implement the above-mentioned model training method.
[0083] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions that, when executed by a processor, implement the above-mentioned model training method.
[0084] In an embodiment of the present application, after obtaining a knowledge document in a vertical field, first, the knowledge document can be subjected to text segmentation processing to obtain a text block set including a plurality of text blocks; where each text block has a theme respectively, so as to ensure that the theme of each document block is clear, and each text block not only includes a reference text block belonging to the knowledge document, but also includes summary information for summarizing the text semantics of the reference text block, so that each text block has the advantage of clear logic. Then, it is supported to generate a data pair set including a plurality of data pairs based on the text block set, specifically, the text blocks in the text block set; where each data pair is composed of a question text and a text block, and the question text in each data pair is written for the text block that matches it, greatly improving the question productivity, and the text block of each data pair is used as the answer source for the question text that matches it, ensuring the authenticity and availability of the data pair. Finally, the first retrieval model can be trained using the data pair set to obtain a second retrieval model and a retrieval library (including the feature vectors obtained by performing vector embedding processing on each text block using the second retrieval model); during the training process, it is necessary to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold, so as to ensure that the features of the question text and the text block in the same data pair are similar, which is beneficial to mapping the features of the target question text to the retrieval library with better data quality during subsequent answer retrieval, and being able to find a text block with similar features from the retrieval library to generate a target answer for the target question text, improving the accuracy and professionalism of answer retrieval in the intelligent question and answer process. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0086] Figure 1a It is a schematic diagram of an intelligent question and answer scenario provided by an exemplary embodiment of the present application;
[0087] Figure 1b It is a schematic diagram of another intelligent question and answer scenario provided by an exemplary embodiment of the present application;
[0088] Figure 1c It is a schematic diagram of yet another intelligent question and answer scenario provided by an exemplary embodiment of the present application;
[0089] Figure 2 It is a schematic flow chart of model training provided by an exemplary embodiment of the present application;
[0090] Figure 3 It is a schematic flow chart of a model training method provided by an exemplary embodiment of the present application;
[0091] Figure 4 It is a schematic flow chart of text segmentation processing provided by an exemplary embodiment of the present application;
[0092] Figure 5 It is a schematic diagram of the effect of text segmentation provided by an exemplary embodiment of the present application;
[0093] Figure 6 It is a schematic diagram of key-value pair verification provided by an exemplary embodiment of the present application;
[0094] Figure 7 It is a schematic flow chart of text segmentation, key-value pair generation and verification provided by an exemplary embodiment of the present application;
[0095] Figure 8a It is a schematic diagram of text allocation provided by an exemplary embodiment of the present application;
[0096] Figure 8b It is another schematic diagram of text allocation provided by an exemplary embodiment of the present application;
[0097] Figure 9 It is a schematic flow chart of model application provided by an exemplary embodiment of the present application;
[0098] Figure 10 It is a schematic flow chart of another model training method provided by an exemplary embodiment of the present application;
[0099] Figure 11 It is a schematic diagram of intelligent question-and-answer interaction provided by an exemplary embodiment of the present application;
[0100] Figure 12 It is a schematic diagram of matching the feature vector of the target question text with the retrieval library during the model application provided by an exemplary embodiment of the present application;
[0101] Figure 13 It is a schematic structural diagram of a model training device provided by an exemplary embodiment of the present application;
[0102] Figure 14 It is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0103] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0104] In the embodiments of the present application, a model training scheme is proposed, specifically providing a model training and application scheme for a retrieval model based on artificial intelligence and intelligent question answering. The following briefly introduces the technical terms and related concepts involved in the model training scheme provided by the embodiments of the present application, where:
[0105] I. Artificial Intelligence (AI).
[0106] Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. That is to say, artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0107] The model training scheme provided by the embodiments of the present application mainly involves the directions of natural language processing technology and machine learning under artificial intelligence. Among them:
[0108] (1) Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics, and is an important technology for model training in the field of artificial intelligence. Among them, the pre-trained model is developed from the large language model (LLM) in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0109] (2) Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The emerging technology - the pre-trained model - is the latest development result of deep learning, integrating the above technologies.
[0110] Furthermore, the model training solution provided by the embodiments of this application mainly involves large language models (LLMs), Dense Passage Retrieval (DPR), and Generative Pre-Trained Transformer (GPT) in machine learning, etc. Among them:
[0111] ① The large language model LLM, also known as the large language model, is a model based on machine learning and natural language processing technologies. It learns the ability to serve human language understanding and generation by training on a large amount of text data. Specifically, the large language model aims to understand and generate human language; by training on a large amount of text data, it can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and so on. The characteristic of the LLM is its huge scale, containing billions of parameters, which helps them learn complex patterns in language data.
[0112] ② The DPR model, also known as the dense passage retrieval algorithm / model, is an open-domain question answering model. When learning dense vector representations, it requires a large number of labeled data pairs (or simply data pairs): question texts and relevant texts. Its main principle is to map all text passages into a low-dimensional continuous space, enabling it to effectively retrieve the top-k (i.e., the first k) text passages to answer the input question.
[0113] ③ The GPT model is a series of neural network models using the Transformer architecture and represents a key advancement in artificial intelligence for supporting generative AI applications. The GPT model enables applications to create human-like text and content (such as images, music, etc.) and answer user questions in a conversational manner. Organizations across various industries are using the GPT model and generative AI for chatbots, text summarization, content generation, and search. Among them, GPT (Generative Pre-trained Transformer), as an emerging large-scale natural language processing (NLP) model based on the Transformer architecture, has several significant advantages over traditional NLP methods:
[0114] 1. Context Understanding: The GPT model can better capture the context information in the text, thus understanding and generating natural language more accurately. In contrast, traditional NLP methods (such as rule-based or statistical methods) may perform poorly when dealing with long-distance dependencies and complex semantic relationships. 2. Pre-training and Transfer Learning: The GPT model adopts the method of pre-training and fine-tuning. It can be pre-trained on a large amount of unlabeled text data and then fine-tuned on specific tasks. This enables GPT to make full use of the information in the massive text data and improve the model performance. Traditional NLP methods usually require a large amount of labeled data for training, resulting in higher costs. 3. End-to-End Training: The GPT model uses an end-to-end training method, which can directly map from the input text to the output text without the need for complex feature engineering. In contrast, traditional NLP methods usually require manual design of features and construction of processing flows, which may lead to lower efficiency and flexibility. 4. Multi-task Learning: The GPT model can be jointly trained on multiple tasks to achieve knowledge sharing and generalization. This enables GPT to solve multiple NLP tasks, such as text classification, named entity recognition, question answering, etc., in one model. Traditional NLP methods usually require separate training models for each task, which may result in higher computational and storage costs. 5. Powerful Generation Ability: The GPT model has a powerful text generation ability and can generate coherent and natural language. This gives GPT an advantage in text generation tasks (such as text summarization, dialogue generation, etc.). Traditional NLP methods may perform poorly in generation tasks, and the generated text may not be smooth and natural enough. 6. Model Scalability: The GPT model has good scalability and can improve performance by increasing the model parameters and the scale of training data. This enables GPT to adapt to the ever-evolving NLP tasks and application requirements. The scalability of traditional NLP methods may be limited and it is difficult to handle complex and large-scale tasks.
[0115] II. Intelligent Question Answering (QA).
[0116] Intelligent question answering can be called open-ended question answering or interactive dialogue, etc., belonging to the field of human-computer interaction; it is an advanced form of information retrieval system, mainly aiming to receive users' questions through computer devices and answer the questions raised by users in accurate and concise natural language. With the rapid development and application of artificial intelligence AI, intelligent question answering technology has become a direction that has attracted much attention and has broad development prospects in the field of NLP. An intelligent question answering system mainly consists of three parts: question understanding, knowledge retrieval, and answer generation. Among them, question understanding mainly includes technologies such as question classification and keyword extraction, aiming to enable the computer device to understand the semantics of the question to be answered input by the user; knowledge retrieval mainly includes structured and unstructured information retrieval, aiming to retrieve knowledge points matching the question to be answered input by the user from the database through the understood question; answer generation mainly includes answer extraction and answer verification, aiming to generate corresponding answers for the question to be answered from the knowledge points retrieved by knowledge retrieval. The core issue of intelligent question answering technology is to handle well the understanding of the question and the matching degree between the question and the answer.
[0117] In practical applications, there are problems in the intelligent question answering system such as low question-answer matching degree, resulting in poor retrieval performance and low accuracy of answers. For example, due to the lack of relevant vertical knowledge in a certain field, the intelligent question answering system cannot accurately record information from the corpus (i.e., language corpus), so in vertical applications, it cannot directly answer specific questions (such as asking whether an insurance product can cover certain diseases. The intelligent question answering system has no relevant field knowledge, and its answer is an illusion (i.e., an answer fabricated by the intelligent question answering system itself), not the correct answer). To improve the retrieval performance of the intelligent question answering system and the accuracy of answer retrieval, the model training solution provided in the embodiments of this application mainly aims at vertical scenarios, combines the powerful text understanding ability of GPT and the local retrieval method of the traditional DPR algorithm to form a new intelligent question answering system for retrieval library generation, training, recall, and multi-round question answering.
[0118] Among them, the core content of the model training solution provided in the embodiments of this application can include the creation, update, and application of the retrieval library (i.e., using the retrieval library to retrieve answers during the intelligent question answering process). Among them, the creation of the retrieval library can refer to the process of training a second retrieval model based on the language corpus in the vertical field for the first retrieval model and constructing the retrieval library based on the second retrieval model and the language corpus. The update of the retrieval library can refer to the design of the metadata structure of the corpus to facilitate operations such as data addition, deletion, and replacement on the retrieval library. The application of the retrieval library can include but is not limited to general problem processing based on the second retrieval model and the retrieval library (such as single-round question retrieval) and user-customized problem processing (such as multi-round question retrieval), etc. The following briefly introduces the main key technologies and innovative methods of the model training solution:
[0119] (1) Construction of the second retrieval model and the retrieval library.
[0120] Among them, the second retrieval model is a retrieval model obtained by training the first retrieval model with training data, and the first retrieval model is the model to be trained before training. The retrieval library can be simply understood as a database storing text blocks of answer sources in an intelligent question-answering system.
[0121] The construction of the second retrieval model and the retrieval library mainly includes: text segmentation, construction of data pairs, and model training. Specifically, first, obtain knowledge documents in the vertical field, which can be understood as a large amount of corpus belonging to the vertical field; in this way, the knowledge documents can be processed by using the full-text understanding ability and inductive summary ability of the GPT model (such as formal transformation, knowledge induction, document segmentation, and document summary, etc.) to obtain a set of text blocks (or called a set of search text blocks); the set of text blocks includes multiple text blocks, each text block has a single theme, and each text block consists of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by generalizing the text semantics of the corresponding reference text block. Then, by using the content summary ability and innovative writing ability of the GPT model, one or more data pairs (or called key-value pairs) corresponding to each text block are constructed according to the key elements of each text block in the set of text blocks; the language styles of different data pairs corresponding to the same text block can be different, and the question text in each data pair is written as a question for the text block that matches it, and the text block in each data pair is used as the answer source for the question text that matches it. Next, use the content screening ability of the LLM model, as well as the ability of dialogue and communication, etc. to check the quality of the data pairs corresponding to each text block, and obtain a set of data pairs composed of data pairs that meet the quality requirements, so as to ensure that the text blocks can cover the question answers, that is, to ensure that the text blocks in the data pairs are the answer sources of the corresponding question texts, and the position information of the answers to the question texts can be located in the text blocks. Finally, use the set of data pairs to train the first retrieval model to obtain the trained second retrieval model and the retrieval library.
[0122] In addition, when constructing a collection of text blocks, it also supports storing text blocks using a metadata structure to manage the retrieval library through the metadata structure, such as updating the retrieval library. Among them, metadata, also known as mediation data or relay data, is a type of data that describes data (data about data); it is mainly information that describes the properties of data and is used to support functions such as indicating storage locations, historical data, resource search, and file records. In this way, by designing the metadata structure for text blocks, operations such as association search, search priority, and convenient addition, deletion, replacement of the retrieval library can be achieved. It should be noted that the update of the retrieval library based on metadata and the association search of text blocks will be introduced in detail in the subsequent specific embodiments, and only a brief description is provided here.
[0123] (2) Handling of generality issues - Application of the retrieval library.
[0124] After constructing the second retrieval model and the retrieval library based on the above implementation method (1), in the actual retrieval process, if the object has a need for the answer to the retrieval question, then the object can input the target question text into the intelligent question-answering system, and the intelligent question-answering system calls the second retrieval model to perform vector embedding processing on the target question text to obtain the vector representation (or called the feature vector) of the target question text. In this way, the intelligent question-answering system can perform similarity matching between the vector representation of the target question text and the feature vectors of each of the multiple text blocks stored in the retrieval library to match multiple text blocks for the target question text, so that based on these multiple text blocks, a target question that matches the target question answer can be generated.
[0125] (3) Handling of user-customized questions - Application of the retrieval library.
[0126] In a multi-round question-answering scenario, the embodiments of the present application support constructing historical object data about the object based on the historical conversation data of the object. To a certain extent, this historical object data can represent the object's inquiry content, inquiry direction, inquiry style, etc. In this way, the GPT model is used to extract and fill slots (or simply referred to as slot filling) of the historical object data in each round of conversation to keep the object data refreshed, so as to maintain the accuracy and timeliness of the object data, match the object situation in real time, and is conducive to realizing personalized recommendation and customized question consultation for the object.
[0127] It can be seen that, on the one hand, the model training solution provided by the embodiments of the present application not only ensures good text block quality by ensuring the single theme and the presence of summary information in the text blocks during text segmentation. Moreover, it directly generates matching questions based on the text blocks to construct data pairs, effectively improving the authenticity and data quality of the data pairs. This data quality is reflected in that the answer to the question text can definitely be extracted from the text block. In this way, when training the retrieval model based on data pairs with good quality, training the retrieval model in the direction of ensuring the reduction of the feature difference of the same data pair enables the retrieval model to have good feature representation capabilities for both the question text and the text block, thereby obtaining a second retrieval model with better vector expression capabilities and a retrieval library with better data quality. On the other hand, the embodiments of the present application make full use of the rich functions of the GPT model, and realize the efficient processing of general questions (i.e., questions in single-round Q&A) and user-customized questions (i.e., questions in multi-round Q&A) by constructing a general retrieval library and a real-time object data capture and synchronization mechanism based on GPT. For general questions, the answer is retrieved by calculating the similarity between the question of the object and the feature vectors of the text blocks in the retrieval library (specifically, the text block is retrieved and the answer is generated based on this text block). For user-customized questions, the object data of the object is extracted by the GPT model, and conditional retrieval is performed based on this object data and the target question text of the object to achieve personalized question answering for different objects, improving the accuracy of the intelligent Q&A system and the experience of object question retrieval.
[0128] The intelligent question-answering system provided by the embodiments of the present application, as an automatic question-answering solution based on artificial intelligence technology, can understand, analyze, and answer questions raised by users; this enables the model training solution provided by the embodiments of the present application to be applicable to various interactive dialogue scenarios that require the use of an intelligent question-answering system to achieve inquiries. The interactive dialogue scenarios may include but are not limited to: ① Customer support: The intelligent question-answering AI system can be used as a customer support tool to answer common questions of users, reduce the workload of customer service staff, and improve customer satisfaction. ② Enterprise internal knowledge base: Enterprises can use the intelligent question-answering AI system to build an internal knowledge base to help enterprise employees quickly find the information they need and improve work efficiency. ③ Virtual assistant: The intelligent question-answering AI system can be used as a virtual assistant for individuals or enterprises, providing functions such as daily task management, schedule arrangement, and reminder services. ④ Online education: The intelligent question-answering AI system can be applied to the field of online education to provide personalized learning resources and real-time question-answering services for students. ⑤ E-commerce: The intelligent question-answering AI system can help users answer questions during the shopping process, provide shopping suggestions, and improve the shopping experience. ⑥ Financial services: The intelligent question-answering AI system can provide real-time consultation services for customers of financial institutions such as banks and insurance companies, answering questions about accounts, transactions, products, etc. ⑦ Medical consultation: The intelligent question-answering AI system can provide basic medical consultation services for patients, answering questions about diseases, treatments, medications, etc. ⑧ Travel consultation: The intelligent question-answering AI system can provide real-time travel information for tourists, answering questions about scenic spots, hotels, transportation, etc. ⑨ News and information retrieval: The AI question-answering system can help users quickly find the news and information they need and improve the efficiency of information retrieval.
[0129] It should be understood that the above description is only an exemplary product performance and interactive dialogue scenario given by the embodiments of the present application, and will not limit the product performance and interactive dialogue scenario of the model training solution provided by the embodiments of the present application. The intelligent question-answering system provided by the embodiments of the present application can provide efficient, accurate, and convenient question-answering services in various interactive dialogue scenarios, showing high value and practicality in various interactive dialogue scenarios, and helping to improve user experience and satisfaction.
[0130] To facilitate the understanding of the model training solution provided by the embodiments of the present application, the following combines Figure 1a the following scenario schematic diagram to briefly introduce the interactive dialogue scenarios involved in the embodiments of the present application; as Figure 1a shown, the system includes an object 101, a terminal 102, and a server 103. The embodiments of the present application do not limit the number and naming of the object 101, the terminal 102, and the server 103.
[0131] Among them, the terminal 102 may refer to a terminal device with an interactive dialogue function; the object 101 can have a single-round or multi-round interactive dialogue with the terminal 102 to obtain relevant knowledge. The terminal 102 may include, but is not limited to: smart phones (such as smart phones deployed with the Android system or smart phones deployed with the Internetworking Operating System (IOS)), tablet computers, portable personal computers, Mobile Internet Devices (MIDs), vehicle-mounted devices, head-mounted devices, intelligent chatbots, and aircraft. The embodiments of the present application do not limit the type of the terminal device, which is hereby explained. The server 103 is the server corresponding to the terminal 102, and is used to perform data interaction with the terminal 102 to provide computing and application service support for the terminal 102. The server 103 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 102 and the server 103 may be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0132] In a specific implementation, the server 103 is supported to enable the model training object to obtain knowledge documents in vertical scenarios / domains (such as education, medical care, insurance, entertainment, banking, and real estate, etc.) from different platforms or systems through the network. Then, the server 103 may perform operations such as document segmentation processing, constructing a data pair set, training a model, and constructing a retrieval library on the knowledge document in sequence according to the model training solution provided by the embodiments of the present application to obtain a trained second retrieval model and a retrieval library.
[0133] Furthermore, the trained second retrieval model and the retrieval library can be directly deployed in the server 103. In this way, when the object initiates an interactive conversation through the terminal 102, the terminal 102 can send the target question text of the object to the server 103 where the trained second retrieval model is deployed. Thus, the server 103 uses the trained second retrieval model to perform vector embedding processing on the target question text to obtain the vector representation of the target question text. Furthermore, the server 103 can retrieve one or more text blocks that match the vector representation of the target question text from the retrieval library based on the vector representation of the target question text, and thus generate a corresponding target answer for the target question text based on the one or more text blocks. Finally, the server 103 sends the target answer to the terminal 102 so that the object can obtain the target answer to the target question text through the terminal 102.
[0134] Of course, the trained second retrieval model can also be deployed in the terminal 102. Then, according to the different deployment locations of the retrieval library, the process of the interactive conversation is different. Optionally, as Figure 1b shown, when the retrieval library is maintained by the server 103, after the terminal 102 receives the target question text of the object and calls the deployed second retrieval model to perform vector embedding processing on the target question text to obtain the vector representation of the target question text, it can send the vector representation to the server 103. The server 103 retrieves one or more text blocks that match the vector representation of the target question text from the retrieval library based on the vector representation of the target question text; then, the server 103 can return the one or more text blocks to the terminal 102, and the terminal 102 generates a corresponding target answer for the target question text based on the one or more text blocks and outputs it. It should be noted that the process of generating a corresponding target answer for the target question text based on the one or more text blocks described above can also be executed by the server 103. In this case, the server 103 directly outputs the target answer to the terminal 102 for display, which can relieve the burden on the terminal 102 to a certain extent. Optionally, as Figure 1c shown, when the retrieval library is maintained by the terminal 102, after the terminal 102 receives the target question text of the object and calls the deployed second retrieval model to perform vector embedding processing on the target question text to obtain the vector representation of the target question text, it can directly retrieve one or more text blocks that match the vector representation of the target question text from the retrieval library based on the vector representation of the target question text, and generate a corresponding target answer for the target question text based on the one or more text blocks and output it.
[0135] Further, the trained second retrieval model can be deployed in the terminal 102 in the form of a plug-in or an application. For example, if the trained second retrieval model is deployed in the terminal 102 as a system-level plug-in, any application deployed in the terminal 102 can call the plug-in to implement an interactive dialogue and provide a target answer for the target question text of the object. Another example is that if the trained second retrieval model is deployed in a certain application, after the terminal 102 starts the certain application, the trained second retrieval model can be called in the certain application to implement an interactive dialogue. Herein, an application program may refer to a computer program for completing one or more specific tasks; classifying application programs according to different dimensions (such as the running mode, function, etc. of the application program), the types of the same application program under different dimensions can be obtained. For example: classified according to the running mode of the application program, the application program may include, but is not limited to: a client installed in the terminal, a small program that can be used without downloading and installing (as a subprogram of the client), a web (World Wide Web) application program opened through a browser, and so on. Another example: classified according to the function type of the application program, the application program may include, but is not limited to: an IM (Instant Messaging) application program, a content interaction application program, and so on. Among them, an instant messaging application program refers to an application program for instant communication messages and social interaction based on the Internet. The instant messaging application program may include, but is not limited to: a social application program with a communication function, a map application program with a social interaction function, a game application program, and so on. A content interaction application program refers to an application program that can implement content interaction. For example, it may be an online banking application program, a sharing platform, a personal space, a news application program, etc. The embodiments of the present application do not limit the specific type of the application program running in the terminal 102 and deployed with the trained second retrieval model, which is hereby explained.
[0136] It should be noted that the above Figure 1a 、 Figure 1b and Figure 1c are only schematic diagrams of the exemplary scenario architecture provided by the embodiments of the present application; in actual applications, this architecture may change adaptively.
[0137] It should also be noted that in the embodiments of the present application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations. Obtaining personal information requires the informed consent of the personal subject (or having a legal basis for information acquisition), and subsequent data use and processing behaviors should be carried out within the scope authorized by laws, regulations and the personal information subject. For example, when the embodiments of the present application are applied to specific products or technologies, such as when obtaining the target question text of the object, the permission or consent of the user needs to be obtained, and the collection, use and processing of relevant data (such as the collection and release of the barrage published by the object, etc.) need to comply with the relevant laws, regulations and standards of the relevant region.
[0138] Based on the above-described model training solution, the embodiments of the present application propose a more detailed model training method. The following will introduce the model training method proposed by the embodiments of the present application in detail with reference to the accompanying drawings. As can be known from the foregoing related descriptions, the model training method provided by the embodiments of the present application mainly includes two parts: model training and model application for the retrieval model. Among them, the model training part includes the construction of the retrieval library. For the convenience of understanding, different embodiments will be used later to introduce the specific implementation processes of model training and model application respectively.
[0139] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a model training method provided by an exemplary embodiment of the present application; this schematic flowchart mainly gives the overall training process of model training from the perspective of model training. As Figure 2 shown, the general process of the model training method may include:
[0140] First, determine the scenario information of the interactive dialogue scenario according to the type of the intelligent Q&A system (such as intelligent robots in shopping malls, food delivery robots in the catering industry, inquiry robots in insurance, etc.); then, collect and organize knowledge documents related to the scenario information of the interactive dialogue scenario as the corpus for model training and the retrieval library to overcome the shortcomings that the GPT model lacks relevant knowledge in vertical scenarios and cannot achieve accurate information recording and recommendation. Among them, the scenario information of the interactive dialogue scenario can include but is not limited to: ① The functional positioning (or called the interaction field) of the intelligent Q&A system, mainly including product positioning and user groups, so as to determine the question purpose and questioning angle of the intelligent Q&A system according to the functional positioning; for example: for the intelligent insurance consulting AI Q&A product, the user group with inquiry needs needs to be positioned as users who need to purchase insurance and have questions about insurance products. ② Interaction style, or called dialogue style, mainly refers to the type or style of questions during human-computer interaction; the dialogue style can include at least one of the following: factual style (or called factual questions, which means that the intelligent Q&A system needs to answer the questions raised by users according to the actual facts), explanatory style (or called explanatory questions, which means that the intelligent Q&A system needs to explain the questions raised by users), reasoning style (or called reasoning questions, which means that the intelligent Q&A system needs to have a certain reasoning ability to answer the questions raised by users), evaluative style (or called evaluative questions, which means that the intelligent Q&A system needs to give an evaluative answer to a certain point in the questions raised by users), and hypothetical style (or called hypothetical questions, which means that the intelligent Q&A system needs to answer the questions raised by users in a hypothetical way). ③ Interaction method, which refers to the dialogue method adopted during human-computer interaction; exemplarily, the interaction method between the user and the intelligent Q&A system can include but is not limited to: text-to-speech interaction method, text-to-text interaction method, voice-to-voice interaction method, and voice-to-text interaction method, etc. For example, the text-to-speech interaction method indicates that the user can enter their questions in the form of input text in the intelligent Q&A system, and when the intelligent Q&A system answers the questions raised by the user, the answer is played in the form of voice broadcast.
[0141] Then, the GPT model is used to judge the document accuracy and Q&A integrity of the knowledge document. Specifically, operations such as document segmentation (or splitting), rewriting, and induction are performed on the knowledge document to split the complete knowledge document into manageable segments that can be processed individually (which can be called text blocks in the embodiments of the present application); this is beneficial for processing the text blocks with limited length individually and greatly facilitates the management of the text blocks. Among them, the logic for judging the document accuracy and Q&A integrity of the knowledge document provided in the embodiments of the present application generally includes: ① Theme consistency, that is, each text block is extracted according to the theme; that is to say, it is necessary to ensure that each text block after segmentation has a theme to ensure the singularity of the theme of the text block, so that each text block has the advantage of a clear theme. ② Word count limit. Considering that the length of the input text that the GPT model can accept is limited and it cannot accept a too long or too large knowledge base, it is impossible to directly answer specific questions in the vertical application; therefore, the present application needs to ensure that the number of characters included in each text block obtained by segmentation is less than the number threshold, and the number threshold is specifically determined by the number of words allowed by the GPT model. ③ Logical relationship. The embodiments of the present application support summarizing the text semantics of each reference text block (that is, directly segmented from the knowledge document) according to the text logical relationship of the knowledge document (such as the general-subordinate relationship or the parallel relationship, etc.) to obtain the general description information of each reference text block, and adding the general description information to the corresponding reference text block to obtain the text block; in this way, the semantics of each text block in the knowledge document can be clarified, and the clarity of the overall logical relationship of the knowledge document can be improved. ④ Q&A structure. The embodiments of the present application support directly storing the text block originally structured as a Q&A structure in the form of a Q&A structure; subsequently, the text block in the Q&A structure can be directly used for question generalization to generate question variants corresponding to the text block in the Q&A structure, improving the richness of the data pair, and thus gradually improving the retrieval accuracy. It can be seen that the embodiments of the present application first utilize the language ability of the GPT model to design text blocks for the knowledge document related to the vertical category, and give a new text segmentation method, making the text blocks have the advantages of the same important knowledge points, complete information, clear theme, and clear logic. It not only ensures the coherence, integrity, and relevance among the text blocks, but also greatly improves the data quality of the retrieval library.
[0142] Secondly, after splitting the knowledge document into multiple text blocks based on the above steps, the embodiments of the present application support storing the metadata structure for each document block, so as to facilitate the subsequent update of the retrieval library based on the metadata. In addition, it also supports generating questions for each text block; specifically, using the text block as the answer source to generate one or more question texts that match the text block, so as to combine each question text and the text block to obtain a data pair, and thus obtain one or more data pairs corresponding to each text block; a data pair includes a question text and a text block, the question text is generated based on the text block, and the text block is the answer source of the question text. Further, after obtaining one or more data pairs corresponding to each text block, in order to ensure that each data pair is not an illusion (i.e., falsely constructed by the GPT model), the embodiments of the present application also need to perform key-value pair verification on each data pair to ensure that each data pair obtained by verification actually exists, thereby improving the authenticity and reliability of the model training and the corpus of the retrieval library.
[0143] Finally, use the multiple processed data pairs (belonging to the data pair set) to train the first retrieval model to obtain the trained first retrieval model (referred to as the second retrieval model in the embodiments of the present application) and the retrieval library. Among them, the determination method of the retrieval library can specifically include: Optionally, obtain the feature vectors of the text blocks obtained by performing vector embedding processing on the text blocks in each data pair during the last iteration training in the model training process, and after removing duplicates from the feature vectors of the text blocks, add the feature vectors of the text blocks to the retrieval library. Optionally, use the second retrieval model to perform vector embedding processing on each text block after the text segmentation processing, and add the obtained feature vectors of the text blocks to the retrieval library.
[0144] Based on the above Figure 2 Brief introduction to the overall process of model training, the following combines Figure 3 Introduce the specific implementation process of model training in detail; Figure 3 Shows a schematic flowchart of a model method provided by an exemplary embodiment of the present application; Figure 3 The model training method process shown is mainly about the model training of the retrieval model and the construction process of the retrieval library. This model training method can be executed by a computer device, and the computer device can be the aforementioned Figure 1a Server 103 shown. The model training method may include but is not limited to steps S301 - S303:
[0145] S301: Obtain a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set.
[0146] As described above, knowledge documents can be understood as a large amount of corpus belonging to a vertical category; the vertical category here, or called a vertical category scenario, refers to a specific field / scenario, so the knowledge document belonging to the vertical category can refer to the language corpus (referred to as corpus) belonging to the specific field. For example, if the vertical category scenario is an insurance scenario, then the knowledge documents belonging to the vertical category scenario may include insurance policies, insurance contracts, and insurance clause texts related to insurance. For another example, if the vertical category scenario is a financial scenario, then the knowledge documents belonging to the vertical category scenario may include deposit rules related to funds or finance, financial product descriptions, etc. It should be understood that the embodiments of the present application do not limit the number (such as the number of knowledge documents is at least one, such as one or more insurance policies on medical insurance, etc.) and types of knowledge documents belonging to the vertical category scenario; for the sake of convenience, the vertical category scenario will be introduced as an insurance scenario, and the knowledge document will be a medical insurance policy belonging to the insurance scenario as an example, which is specially explained here. Furthermore, the embodiments of the present application do not limit the method for acquiring knowledge documents; illustratively, the method for acquiring knowledge documents may include but is not limited to: a computer device acquires them from a platform or system related to a vertical scenario through a network, or a computer device directly acquires them from content publicly available on the Internet, etc.
[0147] After obtaining the knowledge documents in vertical scenarios, considering that knowledge documents often have long content, such as medical insurance policies often have dozens or even hundreds of pages, and knowledge documents with too long content often include many knowledge points (such as insurance clauses about different types of diseases, etc.); if the entire knowledge document is directly used for model training and retrieval library construction, the model performance of the retrieval model and the data quality of the retrieval library will be poor due to factors such as the complex subject matter and rich content of the knowledge document.
[0148] To this end, the embodiments of the present application support the use of a text segmenter to perform text segmentation processing on knowledge documents, so as to split knowledge documents with longer content into small blocks or fragments with less content (i.e., the text blocks mentioned above); in this way, the use of text blocks with less content is more conducive to the realization of model training and the construction of a retrieval library. Among them: ① The text blocks involved in the embodiments of the present application can be represented as Chunks; in the fields of information retrieval and natural language processing, Chunks refer to a smaller fragment or part of a text; in a retrieval library, Chunks refer to the division of larger documents or data sets into smaller, easier to process and analyze parts; in this way, dividing longer content into Chunks and then processing them can improve the efficiency of retrieval and analysis, while Chunks with less content are easier to extract valuable information. ② The text segmenter involved in the embodiments of the present application is an algorithm or method for splitting large text into smaller blocks or fragments; its goal is to create manageable fragments (i.e., text blocks) that can be processed individually, which is usually necessary when processing large documents or data sets.
[0149] In a specific implementation, a schematic diagram of the principle of performing text segmentation processing on a knowledge document by using a text splitter can be referred to Figure 2 the relevant content of the text block design part shown, and the specific implementation process of the text segmentation processing can be as Figure 4 shown, including but not limited to steps S11 - S13, where:
[0150] S11. Obtain the semantic information of the knowledge document, and perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document. Among them, the knowledge document is text content composed of one or more characters, and the text content may also include images. The semantic information of the knowledge document can be used to indicate the semantics expressed by the knowledge document, such as the theme to which the knowledge document belongs, the specific content elaborated by the knowledge document, the text type to which the knowledge document belongs, and so on. The embodiment of the present application does not limit the manner of obtaining the semantic information of the knowledge document; for example: if the text splitter provided by the embodiment of the present application has a semantic extraction function, then the semantic information of the knowledge document can be directly obtained by using the text splitter; for another example, some existing semantic extraction networks (such as the bag - of - words model, etc.) or tools can be used to extract the semantic information of the knowledge document.
[0151] Considering that knowledge documents are often long texts, and overly long content will affect subsequent retrieval and summarization processes. To refine knowledge points and meet the data quality requirements of subsequent retrieval processes, ensure that the themes of the segmented text blocks are clear, and improve the document accuracy of the text blocks (that is, the content expressed by each text block is single); the embodiment of the present application needs to ensure that each text block has only a single theme. The so - called text block having a single theme can be simply understood as that the semantics expressed by the text block is unique and not mixed. For example, if the content of the text block is "Go hiking tomorrow. The weather is nice today", this text block includes two themes, namely the theme of "Go hiking tomorrow" and the theme of "The weather is nice today". Therefore, after obtaining the semantic information of the knowledge document, the text splitter will judge whether the theme to which the knowledge document belongs is single based on the semantic information of the knowledge document.
[0152] On the one hand, if the knowledge document belongs to a single theme, indicating that the knowledge document only expresses the same theme, then judge the number of characters included in the knowledge document to ensure that the number of characters included in the text blocks obtained by splitting meets the requirements of the text splitter, avoiding that the content of the text block is too long for the text splitter to accept and unable to meet the word count requirements of the retrieval library (specifically, the length requirement of the feature vector embedding of the text block stored in the retrieval library). When the number of characters included in the knowledge document is less than the character number threshold, the knowledge document is used as a reference text block; or, when the number of characters included in the knowledge document is greater than or equal to the character number threshold, the knowledge document is processed by paragraph cutting according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document. Among them, the specific value of the character number threshold is related to the type of the text splitter. For example, when the text splitter is a GPT model, most of the time it is necessary to ensure that the word count of the text block is within 512 tokens (a token can be understood as a character unit, and a character unit can include one or more characters).
[0153] On the other hand, if the knowledge document includes at least two themes, indicating that the theme of the knowledge document is not clear, then the knowledge document is segmented hierarchically according to the theme type to obtain at least two initial text blocks, and each of the at least two initial text blocks has a single theme. Then, count the number of characters included in each of the at least two initial text blocks after hierarchical expression segmentation, and use the initial text blocks with the number of characters less than the character number threshold as reference text blocks; and, process the initial text blocks with the number of characters greater than or equal to the character number threshold by paragraph cutting according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document. That is to say, when it is judged that the theme of the knowledge document is not single, the knowledge document can be segmented hierarchically according to the theme type to obtain multiple initial text blocks with a single theme; and for each initial text block, count the number of characters, and use the initial text blocks with the number of characters less than the character number threshold as reference text blocks, and vice versa, process the initial text blocks with the number of characters greater than or equal to the character number threshold by paragraph cutting to obtain multiple reference text blocks corresponding to the knowledge document.
[0154] Among them, it can be seen from the above two aspects that both the knowledge document and the initial text block may be processed by paragraph cutting according to the text logical relationship; for the convenience of description, the knowledge document or the initial text block to be subjected to paragraph cutting in the embodiments of the present application is represented as text content. Among them, the text logical relationship of the text content is a logical relationship used to characterize the overall structure of the text content, and may include: the general-part relationship and the parallel relationship. Among them: the general-part relationship indicates that the content structure of the text content is a general-part structure; that is, the beginning of the text content is often a paragraph summarizing the full text, and the paragraphs after the beginning of the text content are often some explanatory paragraphs for the paragraph summarizing the full text. The parallel relationship indicates that the content structure of the text content is a parallel structure; that is, each part included in the text content has a parallel logical relationship. Then, according to the different text logical relationships of the text content, the methods for paragraph cutting processing of the text content are different; among them: ① when the text logical relationship of the text content is the general-part relationship, the text splitter may split the text content into a general description text block and one or more sub-description text blocks according to the general-part structure, that is, the text content is cut into a general description text block and one or more sub-description text blocks. Among them, the general description text block is the content with a generalizing effect in the text content, that is, a part of the text content with a summarizing effect in the text content; the sub-description text block is the content with an explanatory effect on the general description text block in the text content. For example, a certain sub-description text block mainly elaborates on a knowledge point under the general description text block. ② When the text logical relationship is the parallel relationship, the text splitter may split the text content into at least two sub-description text blocks according to the parallel structure, that is, the text content is cut into at least two sub-description text blocks; there is an independent relationship (or independence) between the at least two sub-description text blocks. The independent relationship here is reflected in that the content / views / facts expressed by the at least two sub-description text blocks are different, so as to clearly identify the independence between the sub-description text blocks during subsequent analysis and avoid theme and logic confusion.
[0155] In summary, it can be seen that the embodiments of the present application mainly perform multiple splits on the knowledge document from the dimensions of theme consistency, word count requirements, and text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document; this makes each reference text block obtained by splitting have the characteristics of being the same important knowledge point (such as a single theme), having a word count less than the character quantity threshold, and being logically clear. In this way, when constructing subsequent data pairs and retrieval libraries based on multiple reference text blocks, the authenticity and rationality of the data pairs can be greatly improved, thereby improving the data quality of the retrieval library.
[0156] S12. By performing semantic generalization on the text semantics of each reference text block in one or more reference text blocks, the general description information corresponding to each reference text block is obtained.
[0157] s13. Add the general information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block form a text block set.
[0158] In steps s12 - s13, it should be noted that the text segmentation process provided in the embodiments of the present application not only performs segmentation on the text content (such as knowledge documents or initial text blocks), but also adds a small amount of summary and rewritten content (referred to as general information in the embodiments of the present application) to the corresponding reference text blocks based on the original text of the reference text blocks. This summary and rewritten content mainly describes the logical status of the reference text blocks in the overall document, so that the information of each text block is complete enough. Taking the general text block and the sub - description text block mentioned in step s11 as the reference text blocks, the summary and rewritten content of the reference text blocks will be introduced. Among them: the general information obtained by the text segmenter through semantic generalization of the general text block is used to indicate the overall semantics expressed by the text content corresponding to the general text block, so that the generality of the general text block can be clearly identified during subsequent analysis. Similarly, the general information obtained by the text segmenter through semantic generalization of the sub - description text block is used to indicate the sub - description semantics, logical structure (that is, the sub - description text block has an explanatory effect on the general text block), and the overall semantics expressed by the text content to which the sub - description text block belongs, so as to ensure the integrity of the information of each finally cut text block chunk.
[0159] Based on the above introduction of the general information for the reference text blocks, the text segmenter in the embodiments of the present application performs semantic generalization on each of the multiple reference text blocks corresponding to the knowledge document, and then the text block corresponding to each reference text block can be obtained, thereby obtaining a text block set. Among them, the text block set includes multiple text blocks, each text block has a theme respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding general information (that is, the above - mentioned summary and rewritten content), and the general information is obtained by generalizing the text semantics of the corresponding reference text block. In addition, the embodiments of the present application support adding the general information of the reference text block to the target position in the reference text block to obtain the text block corresponding to the reference text block; the target position here can include the beginning, middle, or end of the reference text block, and there is no limitation on this.
[0160] It should be noted that the text splitter for performing the text splitting process shown in the above steps S11 - S13 can be a generative pre - trained model. Among them, from the relevant description of the aforementioned generative pre - trained model - the GPT model, it can be seen that the GPT model has extremely strong full - text understanding ability and summarization ability, can adapt to various languages and document formats, and has a large amount of knowledge background at the bottom layer. In the embodiments of the present application, the GPT model can be used as a text splitter at the semantic level to implement the text splitting process for knowledge documents. In practical applications, the embodiments of the present application support using the GPT model to implement the above - mentioned text splitting process through the method of Prompt guidance and fine - tuning. Among them: ① Prompt can be called a prompt word. In large AI models, the main role of Prompt is to prompt the context of the input information and the parameter information of the input model for the AI model (such as the GPT model involved in the present application), so that the AI model can implement the corresponding functions under the guidance / prompt of Prompt; in the embodiments of the present application, the GPT model is mainly prompted through Prompt, so that the GPT model can execute the above - mentioned text splitting process according to the Prompt prompt. ② Fine - tuning can refer to the process of pre - training the GPT model by using training data to enable the GPT model to have certain specific capabilities; for example, in the embodiments of the present application, the GPT model can be fine - tuned so that the GPT model has the ability to extract semantic information from knowledge documents.
[0161] The following combines Figure 5 and takes the GPT model as an example of the text splitter to introduce an exemplary process of using the GPT model to perform text splitting on knowledge documents. As Figure 5 shown, assuming that the text logical relationship of the knowledge document is a general - to - specific relationship, then according to the processing logic of the text splitting process shown in the above Figure 4 the splitting effect of the GPT model on the knowledge document through Prompt guidance is as follows: Through semantic understanding and splitting based on the general - to - specific relationship, the first paragraph of the knowledge document is divided into the first text block Chunk, and this first text block Chunk is the general description text block, and the corresponding general description information is used to represent the logical structure (such as the summary part) and semantic information of the first text block Chunk in the knowledge document. Each paragraph after the first paragraph in the knowledge document is divided into a text block Chunk; of course, if the themes of two adjacent paragraphs are the same and the number of characters is less than the character number threshold, then the same text block Chunk can also include at least two paragraphs, and the embodiments of the present application do not limit the number of paragraphs included in the text block Chunk. And, each text block Chunk after the first text block Chunk in the knowledge document has corresponding general description information, which is used to represent the semantics of the corresponding text block and the overall semantics of the entire knowledge document.
[0162] It should be understood that Figure 5 taking adding the general description information to the front position of the reference text block (such as after the first few characters of the reference text block) as an example, the addition of the general description information to the reference text block to generate the corresponding text block is described. In practical applications, the general description information can also be added to the beginning or end of the reference text block, and there is no limitation on this.
[0163] Taking the text segmentation method provided by the embodiments of the present application and the traditional text segmentation as an example, the advantages of the embodiments of the present application are described as follows: The traditional text segmentation methods may include: LangChain algorithm, nltk algorithm, spacy algorithm, etc. Among them, the traditional LangChain takes a list of characters as a parameter and tries to put all paragraphs together as much as possible, resulting in the summary paragraphs in the general - sub structure of the same theme being divided into multiple Chunks, leading to overly fine segmentation of the text blocks and causing information loss. The sentence - breaking method in the traditional nltk algorithm mainly divides the text into sentences based on punctuation marks (such as full stops, question marks, exclamation marks, etc.); for texts with complex structures and language features in the original document, it cannot be segmented. The sentence - breaking method in the traditional spacy algorithm is mainly based on rules and statistical methods, and realizes text tokenization and sentence - breaking by identifying spaces, punctuation marks, dependency relationships, and grammar rules of specific languages, and does not segment by referring to the semantics of the document itself, resulting in problems such as unclear logic and information loss.
[0164] However, to ensure the information integrity and coherence of each reference text block under the knowledge document, the embodiments of the present application use the GPT model to reasonably summarize and generalize each reference text block to obtain summary information, and add the summary information to the corresponding reference text block to achieve a short explanation of the corresponding reference text block. In this way, each text block Chunk has complete information and a reasonable knowledge density, and when the feature vector corresponding to the text block is put into the retrieval library, it shows coherent information, which is more conducive to the data management and retrieval of the retrieval library. That is to say, compared with some existing text segmentation methods that mainly rely on document formats (such as line breaks, the number of markdowns, punctuation marks, etc.) or use traditional NLP methods for segmentation, the segmentation processing method provided by the embodiments of the present application can effectively overcome information loss (such as in the process of text segmentation, when cutting the document using rigid criteria or traditional algorithms, it essentially does not understand the meaning of the document, especially when concepts in the text span multiple parts, some important information may be lost), loss of context relationship (such as the text fragments after segmentation may lose the original context relationship, resulting in difficulty in understanding), and strong language domain dependence (such as traditional text segmentation methods may be effective for documents in a specific language or domain, but may not work well in other cases. For documents with complex structures and formats, segmentation may become more difficult), etc. defects, so that the text blocks are all of the same important knowledge points, with complete information, clear themes, and clear logic, greatly improving the data quality of the retrieval library.
[0165] S302: Generate a set of data pairs based on the set of text blocks and the summary information included in each text block.
[0166] After obtaining the set of text blocks obtained by segmenting the knowledge document based on the foregoing step S301, the embodiments of the present application support using a key pair generator to generate questions for each text block in the set of text blocks, so as to form a data pair by the text block and the question text generated based on the text block. Thus, a set of data pairs can include multiple data pairs, each data pair consisting of a matching question text and a text block; the question text in each data pair is written for the text block it matches, and the text block in each data pair serves as the answer source for the question text it matches.
[0167] It should be noted that in the embodiments of this application, considering that the data format of the input data of the retrieval model is a data pair, the step of generating a data pair set based on the text block set is proposed. Specifically, considering traditional retrieval models such as the TF-IDF or BM25 algorithm, etc., they all represent questions and context as a sparse high-dimensional space vector by efficiently matching keywords. This way of retrieving only based on word matching does not consider the semantic relevance between text blocks and has great limitations. The DPR model / algorithm represents through dense space vectors that can contain semantic information. By optimizing the maximum inner product of the feature vectors of the question text and the relevant text blocks, the purpose is to compare the similarity of the feature vectors between the question text and the corresponding text blocks in all data pairs in the data pair set (that is, a batch of data pairs), where the data pairs are question-passage, or can be expressed as k-v key-value pairs (abbreviated as k-v pairs), k is the abbreviation of key, which can be expressed as the question text in the embodiments of this application, and v is the abbreviation of value, which can be expressed as the text block including the answer source of the matching question text in the embodiments of this application. This seemingly simple method has a high retrieval accuracy. For example, the article retrieval accuracy of the top-20 articles is 9%-19% higher than that of Lucene-BM25. Based on this, this application uses the DPR model as the retrieval model to implement answer retrieval in an interactive dialogue scenario, specifically the retrieval of the text block containing the answer to the question. The most core part of applying the DPR algorithm as the retrieval model is the construction of the k-v key-value pairs of question-retrieval answer (specifically the text block containing the retrieval answer).
[0168] Furthermore, the embodiments of this application support using a key pair generator to generate a matching question text for each text block in the text block set, and generating the data pair corresponding to the text block based on the text block and its matching question text, so as to obtain a data pair set. Among them, considering that the GPT model has extremely strong text understanding ability and dialogue interaction ability, the embodiments of this application support using the GPT model as the key pair generator to write questions for text blocks through Prompt guidance (for the relevant content of Prompt, refer to the foregoing relevant description and will not be elaborated here).
[0169] Among them, the process of using the GPT model to generate a data pair set based on the text block set and the summary information included in each text block may include:
[0170] (1) The GPT model obtains the scenario information corresponding to the interactive dialogue scenario, and the scenario information includes the dialogue style, interaction method, and interaction field. For the relevant content of the scenario information in the interactive dialogue scenario, refer to the foregoingFigure 2 The relevant content of the embodiments will not be elaborated here. In addition, based on the scenario information, the set of text blocks, and the summary information included in each text block, the GPT model determines the question style and the number of questions (or the number of problem statements) suitable for each text block.
[0171] Specifically, the GPT model mines the key elements of each text block according to the set of text blocks and the summary information included in each text block, and obtains the key elements of each text block; the key elements of the text block here may include, but are not limited to: keywords, context relationships, and logical relationships, etc. Then, based on the scenario information and the key elements of each text block, the GPT model determines the question style and the number of questions suitable for each text block; among them, the question style suitable for each text block is one or more, and each question style belongs to one of multiple conversation styles, and the conversation styles include at least one of the following: factual style, explanatory style, reasoning style, evaluative style, and hypothetical style. In addition, according to the differences in the interactive conversation scenarios, the weight values of the number of questions of each question style corresponding to the text block may be different. Here, the greater the weight value of the number of questions, the more the number of questions; conversely, the smaller the weight value of the number of questions, the fewer the number of questions. For example, in the interactive conversation scenario of insurance query, the number of questions in the explanatory style is more than that in the reasoning style.
[0172] (2) According to the question style and the number of questions suitable for each text block, questions are written for each text block to obtain one or more question texts corresponding to each text block; in the embodiments of the present application, the GPT model is mainly used to implement semantic extraction and other processing of the text block to generate one or more question texts for the text block. Exemplarily, 3 question texts generated for the text block by using the GPT model are as shown in Table 1 below:
[0173] Table 1
[0174]
[0175] It should be understood that Table 1 is introduced by taking the product positioning as medical insurance as an example, and will not limit the products applicable to the embodiments of the present application and the generated questions. This is specifically stated here.
[0176] In addition, referring to the foregoing Figure 2As can be seen from the relevant description of text segmentation in the embodiments shown, there may be some text blocks whose structure is a question-and-answer structure, that is, such text blocks can be used as a data pair by themselves; for the convenience of description, taking the text block set including candidate text blocks, and the candidate text blocks refer to text blocks whose format conforms to the question-and-answer structure, and the candidate text blocks include a question part and an answer part as an example, then for such candidate text blocks with a question-and-answer structure, the embodiments of the present application support constructing new data pairs through question generalization; this way of constructing data pairs through question generalization helps to generate key-value pairs of various dialogue styles and type variants to a certain extent, improves the richness of key-value pairs, and effectively reduces the workload compared with using a model to generate question texts. Specifically, when it is detected that the text block set includes candidate text blocks, the candidate text blocks can be added to the data pair set; and, perform question generalization processing on the question part included in the candidate text blocks to generate a generalized question text, and use the generalized question text and the answer part included in the candidate text blocks to form a new data pair corresponding to the candidate text block, and add the new data pair to the data pair set.
[0177] It should be noted that the above-mentioned question generalization processing method can not only be applied to candidate text blocks with a question-and-answer structure in the text block set, but also be applied to the DPR model iteration stage. Specifically, use the data pairs with recall failure to perform few shots guidance on the questions, realize question generalization for the few shots, and add the new data pairs after question generalization to the training set, which helps to gradually improve the vector expression ability of the DPR model using failure cases, and then improve the model retrieval accuracy.
[0178] (3) Combine each of the one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block; and generate a data pair set based on the one or more data pairs corresponding to each text block.
[0179] It should be noted that, in order to minimize the generation of hallucinations by the GPT model during the process of generating question texts based on text chunks, that is, the generated question texts are false, the embodiments of the present application enforce that the answers corresponding to the generated question texts must originate from the corresponding text chunks, and the position information of the answers to the question texts can be located in the corresponding text chunks, further improving the pertinence and usability of the questions. Therefore, after the key-value pair generator generates a sufficient amount of data pairs through the above description, the embodiments of the present application also support using a key-value pair validator to perform key-value pair verification on each data pair, and only the data pairs that meet the verification requirements can be added to the data pair set to ensure that the answers to the question texts included in the data pairs in the data pair set can be extracted from the corresponding texts, that is, to determine that the data pairs in the data pair set are training data that can be used for subsequent training and verification. In the embodiments of the present application, it is supported to use the GPT model as the key-value pair validator to implement the key-value pair verification of the data pairs corresponding to each text chunk; by utilizing the powerful text understanding ability of the GPT model, etc., to perform secondary verification on the data pairs to ensure the usability of the data pairs and the rationality of question generation.
[0180] Exemplarily, assume that any text chunk in the text chunk set is represented as a target text chunk, and any data pair corresponding to the target text chunk is represented as a target data pair, and the target data pair includes a target question text that matches the target text chunk; then the specific process of using the GPT model to perform secondary verification on the data pair and generating the data pair combination can be referred to Figure 6 . As Figure 6 shown, first, use the key-value pair verification model (i.e., the aforementioned key-value pair validator, which can be the GPT model in the embodiments of the present application) to extract the first answer that matches the target question text from the target text chunk based on the target question text, and mark the position information of the first answer in the target text chunk. The position information of the first answer in the target text chunk can be generated based on the positions of the start character and the end character that make up the first answer in the target text chunk; for example, when marking the characters included in the target text chunk in order from left to right, Figure 6 as shown, the position information of the start character "B" of the first answer "Disease B" in the target text chunk is 23, and the position information of the end character "disease" in the target text chunk is 25.
[0181] Then, according to the position information of the first answer in the target text block, the second answer indicated by the position information is extracted from the target text block; for example, the second answer extracted from the target text block according to the position information 23 and the position information 25 is "B disease". Secondly, the first answer is compared with the second answer extracted based on the position information to obtain a comparison result; the comparison between the first answer and the second answer here may include but is not limited to: determining whether the character string composed of the characters included in the first answer and the character string composed of the characters included in the second answer are exactly the same; and keyword extraction technology (such as Ner technology) can be used to extract keywords from the first answer and the second answer respectively, and compare whether the keywords of the two are matched (i.e., the same).
[0182] Finally, the target data pair is added to the data pair set according to the comparison result; if the comparison result indicates that the first answer is the same as the second answer, which means that the target answer of the target question text in the target data pair can be extracted from the target text block in the target data pair, then the target text block is added to the data pair set; conversely, if the comparison result indicates that the first answer is different from the second answer, which means that the target answer of the target question text in the target data pair cannot be extracted from the target text block in the target data pair, then the target text block is not added to the data pair set, that is, the target data pair is an illusion of the GPT model, and the target data pair is discarded at this time.
[0183] It can be seen that the embodiment of the present application determines the logic and rules for generating questions by setting the prompt, and uses the GPT model for the first time to write questions for text blocks without question text. Compared with the traditional method of extracting question and answer records from historical data (such as the common customer service question and answer manual, using the answers as the retrieval library and the user questions as the corresponding questions to form key-value pairs) or manually writing questions based on text blocks, it greatly improves the productivity of questions and saves time and resources for question generation. In addition, in order to avoid the problem of hallucinations in the process of generating question texts for text blocks by the GPT model as much as possible, the embodiment of the present application also constructs a key pair verifier to perform secondary verification on the generated data pairs. Only data pairs that have been successfully verified (that is, the text blocks in the data pairs are indeed the source of the answers to the corresponding question texts) can be added to the data pair set, which can effectively ensure the availability and data quality of the key pairs, thereby improving the model performance of the second retrieval model trained using the data pair set and the data quality of the retrieval library.
[0184] To facilitate understanding of the specific process shown in the above steps S301-S302, Figure 7 The complete process of text segmentation, key-value pair generation and key-value pair verification is given again. Figure 7As shown in the figure, first, after receiving a knowledge document, the text splitter can perform text splitting on the knowledge document to obtain one or more text chunks, which form a text chunk set. Then, a key-value pair generator is used to generate questions for each text chunk in the text set to obtain one or more data pairs corresponding to each text chunk; and if the text chunk itself is in a question-and-answer structure, the question part of the text chunk is directly generalized to generate data pairs. Finally, a key-value pair validator is used to perform secondary verification on the data pairs corresponding to each text chunk, and the data pairs that are not hallucinated by the key-value pair generator are retained to form a data pair set. The text splitter, key-value pair generator, and key-value pair validator in the above process can all be GPT models. However, during the execution of the above processes by the intelligent question-and-answer system, for different processes and functions, data fine-tuning and Prompt adjustments will be made to GPT. By using the GPT model to complete the functions of each link, the computing and storage costs are greatly reduced, and at the same time, the model accuracy is significantly improved, and the ability boundary of the AI system (i.e., the intelligent question-and-answer system) is broadened.
[0185] S303: Train the first retrieval model using the data pair set to obtain the second retrieval model.
[0186] It should be understood that the training data for model training often includes a training set for model training and a validation set for model testing / verification. Therefore, after constructing the data pair set based on the foregoing steps, a text allocator needs to perform text allocation processing on the data pair set to facilitate extracting some data pairs from the data pair set as the validation set of the first retrieval model, and the remaining data pairs in the data pair set except the extracted data pairs as the training set of the first retrieval model; thus, the first retrieval model is trained based on the training set and the validation set to obtain the trained first retrieval model (i.e., the second retrieval model).
[0187] In specific implementation: (1) The text allocator can perform data allocation processing on the data pair set according to the data allocation strategy to obtain a first target set and a second target set; when the first target set here is the validation set, the second target set is the training set complementary to the first target set, and vice versa. When the first target set is the training set, the second target set is the validation set complementary to the first target set. The types of the first target set and the second target set are not limited in the embodiments of the present application.
[0188] Among them, the data allocation strategy in the above description may include: an answer stratification strategy and a word embedding classification strategy. Among them: ① The answer stratification strategy mainly stratifies the text block based on the position of the answer in the text block, and extracts a certain proportion of data pairs from each layer to form the target set. Specifically, as can be seen from the foregoing description, the answers corresponding to the question texts in multiple data pairs generated from the same text block may be in different positions in the text block. Then, the text block can be stratified according to the position information (composed of the start position and the end position) of different answers in the text block, and then a certain proportion of the data pairs corresponding to the question texts are randomly selected from each layer as the first target set, and the data pairs not selected in each layer are used as the second target set. In this way, it can be ensured that the answer knowledge points involved in the question texts in the target set (such as the first target set or the second target set) are similar to those in the overall question set. ② The word embedding classification strategy mainly extracts a certain proportion of data pairs to form the target set by clustering the word vectors of the question texts in the data pairs. Specifically, according to the representation form of the text (or characters, strings), the word embeddings technology can be used; word embeddings are used to represent words or phrases as fixed-size vectors, and these vectors can capture features such as the similarity and context relationship between words. First, use the word embedding method (such as Word2Vec, GloVe, FastText, BERT, etc.) to convert each character in the string that makes up the question text into a vector, and perform clustering (such as K-means or hierarchical clustering) on these character vectors, etc. The clustering result is used as the division basis, and the strings with similar representation forms are divided into the same group; thus, extraction can be performed from these groups, so as to retain the representation forms of different questions to the greatest extent and increase the diversity and balance of the target set.
[0189] Taking any text block in the text block set as the target text block below, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero, and the question text in the target data pair is composed of one or more characters as an example, the specific implementation of the above two data allocation strategies will be introduced in detail.
[0190] In one implementation, the data distribution strategy is the answer stratification strategy. In this implementation, the text allocator can determine the position information of the answer corresponding to the question text in each of the Q target data pairs corresponding to the target text block within the target text block. Then, based on the position information of the answer corresponding to the question text in each target data pair within the target text block, the target text block is stratified to obtain multiple text sub-layers corresponding to the target text block; one text sub-layer corresponds to at least one of the Q target data pairs. Finally, from the target data pairs corresponding to each text sub-layer among the multiple text sub-layers, reference data pairs are selected and added to the first target set, and the target data pairs other than the selected reference data pairs among the multiple text sub-layers are added to the second target set. Thus, it can be seen that when constructing a target set (such as the first target set or the second target set) through the answer stratification strategy, it can ensure that the target data pairs in the target set come from different layers in the target text block, thereby ensuring that the answer knowledge points involved in the question texts in the target set are located in each layer of the target text block, and further improving the similarity between the answer knowledge points involved in the question texts in the target set and the overall question set.
[0191] Such as Figure 8aAs shown, assume that the target text block includes a total of 10 characters, and there are 4 target data pairs corresponding to this target text block (i.e., Q = 4), namely target data pair 1, target data pair 2, target data pair 3, and target data pair 4. Then, the position information of the answer corresponding to the question text in each of the 4 target data pairs in the target text block can be expressed as follows: The position information of the answer to the question text in target data pair 1 in the target text block is the starting position 1 (1 is the order of the first character of the answer in the target text block) and the ending position 4 (4 is the order of the last character of the answer in the target text block); the position information of the answer to the question text in target data pair 1 in the target text block is the starting position 1 and the ending position 4; the position information of the answer to the question text in target data pair 3 in the target text block is the starting position 5 and the ending position 9; and, the position information of the answer to the question text in target data pair 4 in the target text block is the starting position 5 and the ending position 9. Further, according to the position information of the answers corresponding to the question texts in the 4 target data pairs in the target text block, the target text block is hierarchically processed, and two text sub-layers can be roughly obtained; among them, the position information of text sub-layer 1 is the starting position 1 and the ending position 4, and this text sub-layer 1 corresponds to target data pair 1 and target data pair 2. Similarly, the position information of text sub-layer 2 is the starting position 5 and the ending position 9, and this text sub-layer 2 corresponds to target data pair 3 and target data pair 4. Even further, a certain proportion (such as 50%) of the target data pairs can be extracted from target data pair 1 and target data pair 2 corresponding to text sub-layer 1 and added to the first target set, and the target data pairs not extracted in this text sub-layer 1 are added to the second target set; similarly, a certain proportion (such as 100%) of the target data pairs can be extracted from target data pair 3 and target data pair 4 corresponding to text sub-layer 2 and added to the first target set, and the target data pairs not extracted in this text sub-layer 2 are added to the second target set.
[0192] In other implementation manners, the data allocation strategy is a word embedding classification strategy. In this implementation manner, the text allocator can respectively perform word vector representation on the Q question texts in the Q target data pairs corresponding to the target text block, to obtain the word vectors corresponding to each of the Q question texts; the vector distances between the word vectors corresponding to different question texts are used to indicate the similarities between different question texts. Specifically, the closer the vector distances between at least two word vectors are, the higher the similarities between the at least two question texts corresponding to the at least two word vectors are, indicating that the questions that the at least two question texts want to inquire about may be closer or the same. Then, clustering processing is performed on the word vectors corresponding to the Q question texts, to obtain one or more clustering groups; one clustering group includes the target data pairs corresponding to one or more question texts whose vector distances meet the distance requirement. Here, the vector distances meeting the distance requirement may refer to the vector distances being less than the distance threshold; in other words, it is supported to divide the target data pairs corresponding to the question texts with vector distances less than the distance threshold into one clustering group, which makes the question texts included in the target data pairs in each clustering group similar. Finally, reference data pairs are selected from one or more of the clustering groups and added to the first target set, and the target data pairs in one or more clustering groups except the selected reference data pairs are added to the second target set. Thus, in the case where the same clustering group includes multiple similar target data pairs, by extracting target data pairs from different clustering groups to form the target set, the manifestation forms of different questions can be retained to a large extent, and the diversity and balance of the target set are increased.
[0193] As Figure 8b shown, assume that the target text block totally includes 10 characters, and there are 4 target data pairs corresponding to the target text block (i.e., Q = 4), namely target data pair 1, target data pair 2, target data pair 3, and target data pair 4. Then, after performing word vector representation on the question texts in the 4 target data pairs to obtain the word vectors corresponding to the 4 question texts, clustering processing can be performed on the word vectors corresponding to the 4 question texts to obtain one or more clustering groups. Assume that the vector distance between the word vector of the question text in target data pair 1 and the word vector of the question text in target data pair 3 is less than the distance threshold, then it is determined that target data pair 1 and target data pair 3 are divided into the same clustering group (such as clustering group 1); similarly, assume that the vector distance between the word vector of the question text in target data pair 2 and the word vector of the question text in target data pair 4 is less than the distance threshold, then it is determined that target data pair 2 and target data pair 4 are divided into the same clustering group (such as clustering group 2). In this way, a certain proportion of target data pairs can be extracted from clustering group 1 and clustering group 2 respectively and added to the first target set, and the remaining target data pairs that are not extracted are added to the second target set.
[0194] It should be noted that the above Figure 8aand Figure 8b Both take a single text block as an example to introduce the process of text assignment for one or more data pairs corresponding to the single text block in the data pair set. However, in actual applications, the text assignment method for the data pairs corresponding to each text block in the data pair set is the same as the process shown above Figure 8a and Figure 8b and will not be elaborated here.
[0195] (2) The first retrieval model is trained using the training set to obtain the trained first retrieval model. Among them, the training process of the first retrieval model using the training set can include multiple rounds of iterative training, and the first retrieval model after the last round of iterative training is used as the second retrieval model with better model prediction performance. The second retrieval model is used for problem retrieval in subsequent model applications.
[0196] Among them, when the first retrieval model can be a two-tower model, such as the DPR model, any round of iterative training process for the DPR model can include: taking the data pairs included in the training set as the input data of the DPR model. At this time, according to the two-tower model structure of the DPR model (that is, including sub-model 1 for vector expression of questions and sub-model 2 for vector expression of text blocks), sub-model 1 in the DPR model can perform vector embedding processing on the question text in the data pair to obtain the feature vector of the question text; at the same time, sub-model 2 in the DPR model also performs vector embedding processing on the text block in the data pair to obtain the feature vector of the text block. Then, model training is used to make the feature difference between the matching question text and text block in each data pair less than the preset threshold, that is, optimize the model parameters of the DPR model in the direction of reducing the difference between the feature vectors of the question text and text block in the same data pair; repeat the above optimization process until the DPR model has good vector expression ability for question text and text block, specifically reflected in the similarity of the vector expressions of the question text and text block in the same data pair.
[0197] (3) The trained first retrieval model is tested using a validation set to obtain a second retrieval model. After obtaining the second retrieval model through the above model training process, the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors obtained by performing vector embedding processing on each text block in the text block set by the second retrieval model. In this way, the second retrieval model and the retrieval library can be jointly used to generate a target answer for the target question text input by the object in an interactive dialogue scenario. The construction method of the retrieval library corresponding to the second retrieval model may include: Optionally, the feature vectors of the text blocks included in the retrieval library are obtained during the model training process. Specifically, considering that during the last iteration training of the first retrieval model, the first retrieval model has a better vector representation of the text blocks included in the data pairs in the training set. Therefore, this embodiment of the present application supports adding the feature vectors of the text blocks included in each data pair in the training set during the last round of iteration training to the retrieval library to construct the retrieval library. It should be noted that since the same text block often corresponds to multiple data pairs, before adding the feature vectors of the text blocks included in each data pair in the training set to the retrieval library, it is also necessary to perform deduplication processing on the feature vectors of the text blocks included in each data pair. Specifically, if the same text block is included in different data pairs, only the feature vector of the text block included in one data pair needs to be added to the retrieval library. Optionally, the feature vectors of the text blocks in the retrieval library can also be obtained by performing vector embedding processing on the text blocks in the text block set using the second retrieval model. That is to say, considering that each text block in the text block set is independent and unique, after obtaining the second retrieval model, the second retrieval model can be directly used to perform vector embedding processing on the text blocks in the text block set to obtain the feature vectors, without the need to perform deduplication operations.
[0198] In summary, on the one hand, this embodiment of the present application ensures that the text blocks have a single theme and summary information during text segmentation, ensuring better text block quality. On the other hand, by directly generating matching questions based on the text blocks to construct data pairs, the authenticity and data quality of the data pairs are effectively improved. On the other hand, when training the retrieval model based on data pairs with better quality, the retrieval model is trained in a direction that ensures the reduction of the feature difference of the same data pair, so that the retrieval model has better feature representation capabilities for both question texts and text blocks, thereby obtaining a second retrieval model with better vector expression capabilities and a retrieval library with better data quality.
[0199] The foregoing Figure 2 and Figure 3 The embodiments shown mainly introduce the construction of data pairs, model training, and the construction of the retrieval library. Next, the specific content of model application will be introduced in combination with Figure 9 and Figure 10 embodiments. As Figure 9As shown, during the model application process, after the intelligent question-answering system receives the target question text of the object, it can use the second retrieval model to perform vector embedding processing on the target question text to obtain the feature vector of the target question text, and calculate the similarity between the feature vector of the target question text and the feature vectors of the text blocks included in the retrieval library to determine one or more text blocks that match the target question text; it should be noted that during the retrieval and filtering stage, in addition to using the similarity calculation method described above, other matching methods can also be used, such as extracting key information from the target question text and performing keyword matching on the retrieval library to facilitate determining the uniqueness of the keywords and the necessity of recalling text blocks based on the number of matching keywords.
[0200] Then, the intelligent question-answering system will perform task classification on the target question text to determine the difficulty level of the target question text; it should be noted that in addition to classifying tasks according to the question difficulty, the target question text can also be classified according to other dimensions, such as the urgency of the task, the text theme of the target question text, etc. Finally, one or more text blocks that match the target question text are sent to the corresponding answer model according to the difficulty level of the target question text, so as to generate the corresponding target answer for the target question text based on the one or more text blocks. For example, the answer generation models corresponding to ordinary customer service, professional customer service, and question retrieval, and these answer generation models can correspond to different GPT models; and ordinary customer service and professional customer service are mainly used for simple judgments that require yes / no judgments based on user information or known core instructions, while specific questions (such as the guarantee details of insurance products) or difficult questions can be answered using question retrieval. Thus, considering the performance and memory reasons of a single GPT model, it cannot carry an extremely complex intelligent question-answering system. Therefore, the embodiments of this application support using multiple GPTs for different role-playing and function implementation (such as different GPT models for handling questions of different difficulty levels (such as ordinary customer service, professional customer service, etc.), another example is the GPT model for text segmentation, and another example is the GPT model for key pair generation and the GPT model for key pair verification, etc.) to jointly complete the overall intelligent AI question-answering system.
[0201] In addition, as described above, the intelligent question-answering system provided by the embodiments of this application also supports multi-round question-answering, specifically by combining the historical object data of the same object during the multi-round interactive dialogue process to achieve a customized response for the object. During the model application process, it is reflected that after the intelligent question-answering system obtains the target question text of the object, it can rewrite the target question text based on the historical object data of the object, so that the rewritten question text is more matched with the object data of the object, thereby achieving better providing personalized and customized target answers for the object.
[0202] Based on the above Figure 9 general introduction to the application of the model, the following will combine Figure 10 to introduce in detail the complete implementation process of model training and model application; Figure 10 The flowchart shows the schematic flowchart of a model method provided by an exemplary embodiment of the present application; Figure 10 The model training method process shown is mainly about the process of model application of the second retrieval model. This model training method can be executed by a computer device, specifically by an intelligent question answering system installed in the computer device. The model training method may include but is not limited to steps S1001 - S1008:
[0203] S1001: Obtain knowledge documents, and perform text segmentation processing on the knowledge documents to obtain a set of text blocks.
[0204] S1002: Generate a set of data pairs based on the set of text blocks and the summary information included in each text block.
[0205] S1003: Use the set of data pairs to train the first retrieval model to obtain the second retrieval model.
[0206] It should be noted that for the specific implementation processes shown in steps S1001 - S1003, reference can be made to the relevant descriptions of the specific implementation processes shown in steps S301 - S303 in the foregoing Figure 3 shown embodiments, and details will not be elaborated here.
[0207] It should also be noted that in traditional retrieval schemes, the processing methods and application status of all data are the same. However, in fact, the information reliability and application popularity of data from different data sources are different, and to ensure the information granularity, long texts belonging to the same topic are segmented, so that multiple text blocks after segmentation come from the same paragraph of text, that is, the topics of multiple text blocks may be the same, but their meanings and correlations may be different. To make full use of these differences of different text blocks in the retrieval library, in the embodiments of the present application, after performing text segmentation processing on the knowledge documents to obtain a set of text blocks, a new index design is also set for each text block in the set of text blocks, so that it is possible to perform searches on specified text blocks (such as according to the source or topic, etc.) based on metadata, and it is helpful to manage the retrieval library according to metadata (such as updating, adding, or deleting the feature vectors of text blocks in the retrieval library).
[0208] In a specific implementation, after obtaining the text block set, metadata extraction can be performed on each text block in the text block set to obtain the metadata of each text block; the metadata of any text block can be used to describe the data attributes of the any text block, and the data attributes can at least include: the retrieval domain to which the text block belongs, the source of the text block, and the theme to which the text block belongs, etc. That is to say, three concepts are introduced for the text block in the index design mechanism set in the embodiments of the present application, namely: retrieval domain (or simply referred to as "domain"), source, and theme. Among them: ① The "domain" is the overall scope of the retrieval library during retrieval. For example, in the case where the text block feature vectors of multiple insurance products are included in the retrieval library, each insurance product can be a "domain"; it is feasible to use data pairs within multiple domains during the model training process to train the first retrieval model. ② The "source" can refer to the source information or source location of the text block; for example, the source of the text block can be the problem summary obtained during the customer service Q&A process, or the source of the text block can be the information obtained from the insurance terms, etc. ③ The "theme" is the main knowledge content of the text block Chunk; the relevance between the text block and the question can be confirmed through the theme.
[0209] In the embodiments of the present application, it is supported to update the retrieval library by using the metadata of each text block and perform index retrieval on the text block. Among them: ① The update of the retrieval library may include: when the data attribute of the text block in the retrieval library is updated, modify the attribute value of the corresponding data attribute in the metadata of the text block to implement the update of the text block. For example, when the retrieval library is continuously updated, the new version number or version time (i.e., version update time) of the text block can be appended to the data attribute "source" of the text block, so that the retrieval library can be managed, increased, decreased, and replaced. ② The index retrieval of the text block at least includes: 1. In an interactive dialogue scenario, according to the retrieval requirements indicated by the target question text, retrieve the text blocks whose metadata meets the retrieval requirements from the metadata to generate an answer for the target question text. For example, in the process of model application, if the object specifies a query for a certain insurance product, the intelligent question and answer system can only retrieve from the text blocks within the insurance product, rather than from the text blocks within other insurance products. At this time, it is necessary to filter the text blocks that meet the specified insurance product according to the data attribute "domain" as the text blocks that match the target question text of the object. 2. When any text block has been retrieved in an interactive dialogue scenario, retrieve the text blocks whose metadata is the same as the metadata of any text block from the metadata to generate an answer for the target question text. For example, based on the relevant content of the foregoing text segmentation processing, for a long text belonging to the same theme, the multiple text blocks after the long text is segmented all adopt the same theme. Then, in the process of model application, if one text block belonging to the same theme is recalled, other text blocks belonging to the same theme can be found according to the data attribute "theme" as auxiliary information for retrieval (such as assisting in generating an answer for the retrieved text block).
[0210] It should be noted that in the actual process of using metadata to implement index retrieval and retrieval library update, considering that the data attribute "domain" and the data attribute "source" are strongly related to the retrieval service and are explicit information (that is, directly marked for the text block or added to the text block), they can be directly used. If the data attribute "theme" exists in the text block (such as the title of the text block), it can also be directly used. However, if the data attribute "theme" does not exist in the text block, that is, it is implicit information, the GPT model mentioned above can be used to perform theme induction in real time.
[0211] In summary, the embodiments of the present application construct the above-described index mechanism for the text blocks in the retrieval library, enabling the index retrieval of text blocks and the update of the retrieval library to be realized by using the metadata structure, thereby making the retrieval library clearer and improving the retrieval speed and efficiency during the retrieval process, and facilitating the management of the retrieval library.
[0212] S1004: Receive the target question text input by the object in an interactive dialogue scenario, and perform vector embedding processing on the target question text using the second retrieval model to obtain the feature vector of the target question text.
[0213] As described above, the second retrieval model involved in the embodiments of the present application can be a two-tower model. For example, in the DPR model, it includes a sub-model for vector representation of questions and a sub-model for vector representation of text chunks. Therefore, in the interactive dialogue scenario of model application, after the intelligent question-answering system receives the target question text input by the object, it can use the sub-model for vector representation of questions in the second retrieval model to perform vector embedding processing on the target question text to obtain the feature vector of the target question text, and this feature vector is used to represent the semantic information of the target question text.
[0214] It should be noted that the embodiments of the present application do not limit the interaction method between the intelligent question-answering system and the object, that is, it does not limit the way the intelligent question-answering system receives the target question text of the object. For example: Figure 11 As shown in the first attached figure, the object can directly input the target question text in text form on the display screen of the intelligent question-answering system. As Figure 11 shown in the second attached figure, the object can ask questions by voice. At this time, the intelligent question-answering system can collect the voice signal of the object and convert the voice signal into the target question text in text form. As Figure 11 shown in the third attached figure, the intelligent question-answering system can also actively output one or more candidate questions in semantic or text form, so that the object selects one candidate question from the one or more candidate questions displayed on the display screen of the intelligent question-answering system as the target question text to be consulted; and so on.
[0215] S1005: Perform similarity calculation processing on the feature vector of the target question text and the feature vectors of P text chunks in the retrieval library to obtain the matching scores between each of the P text chunks and the target question text.
[0216] S1006: Sort the P matching scores in descending order, and perform gradient operations on each matching score in the sorted score sequence to obtain multiple gradient information.
[0217] S1007: Based on the multiple gradient information, dynamically screen one or more matching scores in the score sequence whose gradient information meets the gradient descent condition, and use the text chunks corresponding to the one or more matching scores as the answer sources of the target question text.
[0218] In steps S1005 - S1007, considering that the goal of training the second retrieval model in the embodiments of the present application is to make the feature differences between the question text and the text block in the same data pair smaller, that is, the feature vectors of the question text and the text block in the same data pair are more similar, so as to ensure that the question text in the same data pair can definitely find the corresponding answer from the text block. Therefore, in the intelligent question - answering system in the embodiments of the present application, after receiving and invoking the second retrieval model to perform vector embedding processing on the target question text to obtain the feature vector of the target question text, the feature vector of the target question text can be mapped to the retrieval library; specifically, the similarity calculation process is performed between the feature vector of the target question text and the feature vectors of each of the P text blocks in the retrieval library to obtain the matching scores between the P text blocks and the target question text respectively; the matching score corresponding to each text block is used to indicate the credibility of the corresponding text block as the answer source of the target question text, or to indicate the similarity degree between the feature vector of the corresponding text block and the feature vector of the target question text; in this way, the higher the matching score corresponding to the text block, the higher the possibility that the text block is used as the answer source of the target question text.
[0219] Considering that the general retrieval method in the industry is to extract a specified number of text blocks from the retrieval library and hand them over to the intelligent question - answering system for answering; however, directly recalling a specified number of text blocks may, due to the specified quantity, result in too much invalid information in the recalled text blocks, leading to problems such as overly verbose or even self - contradictory answers. To improve the flexibility of recalling the number of text blocks, the embodiments of the present application design a method of dynamically recalling text blocks based on matching scores (or a new split - score dynamic selection form), ensuring that all the recalled text blocks are as concise and clear as possible around the target question text, thereby improving the effectiveness and accuracy of the target answer generated for the target question text and further improving the retrieval accuracy.
[0220] Among them, the general principle of the dynamic allocation method for the recall number of text chunks Chunk according to the matching score score may include: for the target question text proposed by the object, after calculating the matching scores between the target question text and P text chunks in the retrieval library using the foregoing description, the P matching scores can be arranged in descending order (or from large to small) to obtain a score sequence, and the score sequence includes P matching scores with values from large to small. Then, calculate the descending gradient one by one for this score sequence. Specifically, subtract the previous matching score from the latter matching score among two adjacent matching scores in the score sequence to obtain P - 1 gradient information, and perform normal fitting on these gradient information and find one or more significant descending gradients outside [mean - 3 * standard deviation, mean + 3 * standard deviation]. These significant descending gradients indicate a significant decrease in the similarity between the answer and the question. Then, find one or more matching scores corresponding to the first significant descending gradient among the one or more significant descending extractions from the score sequence (i.e., one or more matching scores that meet the gradient descent condition), and determine that the text chunks corresponding to this matching score and the text chunks before this text chunk in the score sequence are more likely to be the answer sources of the target question text.
[0221] It should be noted that the above is introduced by directly sorting the P matching scores corresponding to the P text chunks in the retrieval rate; in practical applications, in order to save computational overhead, the embodiments of the present application support specifying to output the top-k text chunks (such as k = 10) and the matching scores corresponding to the top-k text chunks to achieve dynamic recall. For example, as Figure 12 shown, assume that after calculating the similarity between the feature vector of the target question text and the feature vectors of P text chunks in the retrieval library, the top-5 text chunks corresponding to the matching scores and the corresponding matching scores are screened, and the text chunks are sorted according to the matching scores as follows: text chunk 1 → text chunk 2 → text chunk 3 → text chunk 4 → text chunk 5, and the score sequence obtained by sorting the matching scores corresponding to each text chunk is: score 1 → score 2 → score 3 → score 4 → score 5. Then calculate the gradient information as: gradient 1, gradient 2, gradient 3, and gradient 4, and perform normal fitting on these gradient information to obtain [mean - 3 * standard deviation, mean + 3 * standard deviation]. If it is determined that the significant descending gradients outside [mean - 3 * standard deviation, mean + 3 * standard deviation] are gradient 3 and gradient 4. Further, determine that the first significant descending gradient among the two significant descending gradients is gradient 3, then find the matching score corresponding to the first significant descending gradient "gradient 3" in the score sequence as score 4, and determine that the text chunks 3, 2, and 1 before the text chunk 4 corresponding to this score 4 in the score sequence are more likely to be the answer sources of the target question text, then determine the recalled text chunks as: text chunk 1, text chunk 2, and text chunk 3.
[0222] It can be seen that the text block recall method provided in the embodiments of the present application does not simply limit the recall of the first few text blocks from the retrieval library, but dynamically recalls text blocks according to the descending gradient; moreover, when the answer to the target question text is relatively clear, the matching score between the target question text and the text block including the answer will increase significantly, so one or more of the most relevant text blocks can be returned, further improving the recall accuracy and reducing the dilution of effective information by redundant answers.
[0223] S1008: Generate a target answer for the target question text based on the text block corresponding to one or more matching scores.
[0224] As described above Figure 9 In the related description of the foregoing embodiments, the intelligent question-answering system provided in the embodiments of the present application is also configured with answer generation models corresponding to different difficulty levels; then, after obtaining one or more matching scores (specifically, the matching scores before the first significant descending gradient) that meet the gradient descent condition based on the foregoing steps S1004-S1007, one or more matching scores can be assigned to the answer generation models corresponding to the difficulty levels according to the difficulty level of the target question text to generate the target answer. In a specific implementation, the intelligent question-answering system performs task classification processing on the target question text to obtain a classification result, which is used to indicate the reply difficulty of the target question text (i.e., the difficulty level mentioned above); then, the intelligent question-answering system, according to the reply difficulty indicated by the classification result, uses an answer generation model that matches the answer difficulty to generate an initial answer to the target question text based on the text block corresponding to one or more matching scores.
[0225] Furthermore, in order to ensure the accuracy of the initial answer generated by the answer generation model, the embodiments of the present application also support answer quality inspection / review of the initial answer, and use the reviewed initial answer as the target answer that matches the target question text. Specifically, the intelligent question-answering system supports answer review processing of the initial answer to obtain the target answer to the target question text; in the actual review process, if the initial answer is a suitable answer (such as the style is suitable for the target question text), then the initial answer is the target answer, and if the initial answer is not a suitable answer, then the answer after the initial answer is fine-tuned is determined as the target answer. In detail, the intelligent question-answering system may include a quality inspection model with a quality inspection function, so the intelligent question-answering system can call the quality inspection model to perform quality inspection on the initial answer. Here, the quality inspection is mainly to input the target question text to the quality inspection model for checking the question style, and rewrite the initial answer to obtain the target answer when the form or style of the initial answer is inappropriate.
[0226] Furthermore, traditional Q&A mainly focuses on single-round Q&A. When faced with multi-round Q&A, it usually splits it into single-round forms or uses manual strategies for forced intervention. This results in the loss of a lot of key information between multi-round Q&As, and the relevance between multi-round Q&As is not strong, and the response is rigid. To improve the flexibility and relevance of multi-round Q&A, the embodiments of the present application support leveraging the long-term memory and extraction capabilities of the GPT model, adopting different GPT models for role-playing, and combining the object data of the object to play different advantages in different links of multi-round Q&A, so as to achieve customized answers for different objects.
[0227] Taking the example of an interactive dialogue scenario including multi-round interactive dialogues, and any round of interactive dialogue except the first round of interactive dialogue in the multi-round interactive dialogue is represented as the target round interactive dialogue, the operations performed by the intelligent Q&A system in the target round interactive dialogue of the multi-round interactive dialogue may include: First, obtain the target question text input by the object in the target round interactive dialogue, the historical dialogue data (such as the historical question text and the answer to the historical question text) in the historical round interactive dialogue (such as the interactive dialogue located between the target round interactive dialogue in the multi-round interactive dialogue), and the historical object data about the object (such as some personalized data about the object extracted during the historical round interactive dialogue, such as the object's inquiry focus, the object's basic information (such as age and gender, etc.)), and this historical object data is generated based on the historical dialogue data. Then, based on the target question text, historical dialogue data, and historical object data, rewrite the target text question to obtain a new target question text. In this way, the second retrieval model can be used to perform vector embedding processing on the new target question text, calculate the similarity between the feature vector of the new target question text and the feature vectors of the text blocks in the retrieval library to obtain multiple matching scores, and perform operations such as screening the text blocks that match the target question text based on the multiple matching scores (that is, the specific processes shown in steps S1004 - S1007 above). It should be noted that the target round interactive dialogue described above is any round of interactive dialogue except the first round of interactive dialogue in the multi-round interactive dialogue. For the first round of interactive dialogue in the multi-round interactive dialogue, it can be regarded as a single-round interactive dialogue, and there is no object data that can be introduced.
[0228] As described above, the embodiments of the present application implement different processes of multi-round Q&A by using multiple GPT models. The following introduces the general process of the multi-round interactive dialogue of the intelligent Q&A system from the perspective of different GPT models:
[0229] First, after receiving the target question text, the intelligent question-answering system can use GPT model A as a process controller and input the current question text, historical conversation data (if this is the first round of the interactive conversation, the historical conversation data is empty), and historical object data (if this is the first round of the interactive conversation, the historical object data is empty) proposed by the object in the current round of the interactive conversation process into GPT model A. In this way, GPT model A performs the following operations: extracting new object data based on the current question text, historical conversation data, and historical object data in the current round of the interactive conversation process; determining whether the historical object data needs to be updated or added based on the current interactive conversation; and rewriting the current question text according to the new object data and the current question text to generate a new, simple and understandable question text.
[0230] Then, the intelligent question-answering system calls the second retrieval model to perform vector embedding processing on the new question text and executes the specific steps shown in the aforementioned steps S1004 - S1007 to obtain one or more text blocks that match the new question text.
[0231] Finally, the intelligent question-answering system uses GPT model B as a logical question distributor to determine whether, for the rewritten question (i.e., the new question text), an answer can be directly given based on one or more text blocks that match the new question text. If so, it indicates that the difficulty of answering the new question text is relatively low, such as questions about consulting prices or policy amounts. In this case, GPT model C (such as the ordinary customer service shown in Figure 9 is used as a general answerer to directly generate a target answer based on one or more text blocks that match the new question text for answering. Conversely, if not, it indicates that the new question text has a certain degree of difficulty in answering, such as determining whether a certain industry can be insured. In this case, it is necessary to determine whether more professional logical reasoning is required. If not, it indicates that although the new question text has a certain degree of difficulty in answering, the difficulty level is not high, and it can be handed over to GPT model D (such as the professional customer service shown in Figure 9 to generate a target answer based on one or more text blocks that match the new question text for answering. If so, it indicates that the difficulty level of answering the new question text is relatively high, and GPT model E (such as the question retrieval model shown in Figure 9 is used to perform logical reasoning and answer output in combination with object data.
[0232] In the embodiments of the present application, from the user experience level, the intelligent question answering system can support multi-round interactive conversations and user-customized recommendations, greatly improving the fluency, naturalness, and accuracy of answers. From the underlying technology level, the embodiments of the present application have at least the following advantages: 1. A new text segmentation technology is proposed from the perspective of text segmentation. This text segmentation technology can ensure that the segmented text blocks have a single theme and carry summary information, making the text blocks have the advantages of clear themes and clear logic, thereby improving the data quality of the retrieval library constructed based on the text blocks. 2. A new data pair generation scheme is proposed from the perspective of data pair generation. Specifically, the functions and advantages of the emerging technology GPT are utilized to construct more real and reliable data pairs based on the text blocks, that is, to ensure that the text blocks in the same data pair must be the answer sources of the question text. 3. From the perspective of model training, when training the first retrieval model with high-quality data pairs, it also supports training the model in the direction of reducing the feature differences between the question text and the text blocks in the same data pair, so that the trained first retrieval model (i.e., the second retrieval model) has good vector expression capabilities for both the question text and the text blocks, and can make the vector expression of the question text closer to the vector expression of the text block that is the answer source of the question text; thus, during the model application process, a better and more accurate text block can be matched for the target question text. 4. From the perspective of text block recall, it supports the method of dynamic recall numbers based on matching scores, which can dynamically recall different numbers of text blocks for different questions and ensure the matching of the recalled text blocks and the corresponding questions, thereby improving the effective information concentration of the recall. 5. Starting from the retrieval library, it not only supports storing the metadata structure for the text blocks in the retrieval library, which is beneficial for index retrieval and retrieval library update (such as iterative update) based on the metadata of the text blocks, but also the retrieval library supports various forms, including forms, pictures, videos, and other forms. In summary, the embodiments of the present application support using multiple GPT models to play different roles for link control, function implementation, and reply quality inspection, etc., and are committed to improving the retrieval ability, link control ability, and continuous iteration ability (such as multi-round question answering combined with object data) of the intelligent question answering system from each link, and generating sustainable iterative solutions.
[0233] The method of the embodiments of the present application is described in detail above. To facilitate the better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below. In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.
[0234] Figure 13 shows a schematic structural diagram of a model training device provided by an exemplary embodiment of the present application; the model training device can be used to execute Figure 3 or Figure 10 part or all of the steps in the method embodiments shown. Please refer to Figure 13 , the device includes the following units:
[0235] An acquisition unit 1301, configured to acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a set of text blocks; the set of text blocks includes multiple text blocks, each text block has a theme respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by generalizing the text semantics of the corresponding reference text block;
[0236] A processing unit 1302, configured to generate a set of data pairs based on the set of text blocks and the summary information included in each text block; the set of data pairs includes multiple data pairs, and each data pair is composed of a matching question text and a text block; the question text in each data pair is written as a question for the text block that matches it, and the text block in each data pair serves as the answer source for the question text that matches it;
[0237] The processing unit 1302 is further configured to train a first retrieval model using the set of data pairs to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes feature vectors of each text block in the set of text blocks; wherein, the second retrieval model and the retrieval library are used to generate a target answer for the target question text in an interactive dialogue scenario.
[0238] In one implementation, the number of knowledge documents is at least one; when the processing unit 1302 is configured to perform text segmentation processing on the knowledge document to obtain a set of text blocks, it is specifically configured to:
[0239] Acquire the semantic information of the knowledge document, and perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document;
[0240] By generalizing the text semantics of each reference text block in the one or more reference text blocks, obtain the summary information corresponding to each reference text block;
[0241] Add the summary information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block form a set of text blocks.
[0242] In one implementation, the semantic information of the knowledge document includes the theme to which the knowledge document belongs, and the knowledge document is composed of one or more characters; when the processing unit 1302 is used to perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document, it is specifically used for:
[0243] Determine whether the theme to which the knowledge document belongs is single;
[0244] If the theme to which the knowledge document belongs is single, count the number of characters included in the knowledge document;
[0245] When the number of characters included in the knowledge document is less than the character number threshold, use the knowledge document as a reference text block; or,
[0246] When the number of characters included in the knowledge document is greater than or equal to the character number threshold, perform paragraph cutting processing on the knowledge document according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
[0247] In one implementation, the processing unit 1302 is further used for:
[0248] If the knowledge document includes at least two themes, perform hierarchical expression segmentation on the knowledge document to obtain at least two initial text blocks; each of the at least two initial text blocks has one theme;
[0249] Count the number of characters included in each of the at least two initial text blocks, and use the initial text block with the number of characters less than the character number threshold as a reference text block; and,
[0250] Perform paragraph cutting processing on the initial text block with the number of characters greater than or equal to the character number threshold according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
[0251] In one implementation, the text logical relationships are respectively the general - part relationship and the parallel relationship; the general - part relationship indicates that the content structure of the text content is a general - part structure, and the parallel relationship indicates that the content structure of the text content is a parallel structure; the knowledge document or the initial text block to which the paragraph cutting processing is performed is represented as the text content; the paragraph cutting processing includes:
[0252] Perform paragraph cutting processing on the text content according to the text logical relationship of the text content to obtain the reference text blocks corresponding to the knowledge document;
[0253] Among them, when the text logical relationship of the text content is a general - specific relationship, the text content is cut into a general description text block and one or more sub - description text blocks; the general description text block is the content with a generalizing effect in the text content, and the general information corresponding to the general description text block is used to indicate the overall semantics expressed by the text content; the sub - description text block is the content that explains the general description text block in the text content, and the general information corresponding to the sub - description text block is used to indicate the sub - semantics expressed by the sub - description text block and the overall semantics expressed by the text content;
[0254] When the text logical relationship is a parallel relationship, the text content is cut into at least two sub - description text blocks.
[0255] In one implementation, the processing unit 1302 is further configured to:
[0256] Extract metadata for each text block in the text block set to obtain the metadata of each text block; the metadata of the text block is used to describe the data attributes of the text block, and the data attributes include: the retrieval domain to which the text block belongs, the source of the text block, and the topic to which the text block belongs;
[0257] Among them, the metadata of each text block is used to update the retrieval library and perform index retrieval for the text block; the update includes: when the data attributes of the text block in the retrieval library are updated, modifying the attribute value of the corresponding data attribute in the metadata of the text block;
[0258] The index retrieval at least includes: in an interactive dialogue scenario, according to the retrieval requirements indicated by the target question text, retrieving text blocks whose metadata meets the retrieval requirements from the metadata to generate an answer for the target question text; and, when any text block has been retrieved in the interactive dialogue scenario, retrieving text blocks whose metadata is the same as the metadata of any text block from the metadata to generate an answer for the target question text.
[0259] In one implementation, when the processing unit 1302 is used to generate a data pair set based on the text block set and the general information included in each text block, it is specifically configured to:
[0260] Obtain the scenario information corresponding to the interactive dialogue scenario, where the scenario information includes the dialogue style, interaction method, and interaction field; and,
[0261] Based on the scenario information, the text block set, and the general information included in each text block, determine the question style and the number of questions adapted to each text block;
[0262] According to the question style and the number of questions adapted to each text block, write questions for each text block to obtain one or more question texts corresponding to each text block;
[0263] Combine one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block;
[0264] Generate a set of data pairs based on one or more data pairs corresponding to each text block.
[0265] In one implementation, when the processing unit 1302 is used to determine the question style and the number of questions adapted to each text block based on the scenario information, the text block set, and the general description information included in each text block, it is specifically used for:
[0266] Mine key elements for each text block according to the text block set and the general description information included in each text block to obtain the key elements of each text block;
[0267] Determine the question style and the number of questions adapted to each text block based on the scenario information and the key elements of each text block;
[0268] Among them, the question style adapted to each text block is one or more, and each question style belongs to one of multiple dialogue styles; the dialogue styles include at least one of the following: factual style, explanatory style, reasoning style, evaluative style, and hypothetical style.
[0269] In one implementation, any text block in the text block set is represented as a target text block, and any data pair corresponding to the target text block is represented as a target data pair. The target data pair includes a target question text that matches the target text block. When the processing unit 1302 is used to generate a set of data pairs based on one or more data pairs corresponding to each text block, it is specifically used for:
[0270] Use a key-value verification model to extract, from the target text block, a first answer that matches the target question text and the position information of the first answer in the target text block based on the target question text;
[0271] Extract a second answer indicated by the position information from the target text block according to the position information of the first answer in the target text block;
[0272] Compare the first answer and the second answer, and add the target data pair to the set of data pairs based on the comparison result;
[0273] Among them, if the comparison result indicates that the first answer is the same as the second answer, the target text block is added to the set of data pairs; if the comparison result indicates that the first answer is different from the second answer, the target text block is not added to the set of data pairs.
[0274] In one implementation, the set of text blocks includes candidate text blocks. A candidate text block refers to a text block whose format conforms to the Q&A structure. The candidate text block includes a question part and an answer part. The processing unit 1302 is further configured to:
[0275] Add the candidate text block to the set of data pairs; and,
[0276] Perform question generalization processing on the question part included in the candidate text block to generate a generalized question text, and form a new data pair corresponding to the candidate text block with the generalized question text and the answer part included in the candidate text block, and add the new data pair to the set of data pairs.
[0277] In one implementation, when the processing unit 1302 is configured to train the first retrieval model using the set of data pairs to obtain the second retrieval model, it is specifically configured to:
[0278] Perform data allocation processing on the set of data pairs according to the data allocation strategy to obtain a first target set and a second target set; the data allocation strategy includes an answer stratification strategy and a word embedding classification strategy; the first target set is the validation set, and the second target set is the training set; or, the first target set is the training set, and the second target set is the validation set;
[0279] Train the first retrieval model using the training set to obtain the trained first retrieval model;
[0280] Test the trained first retrieval model using the validation set to obtain the second retrieval model.
[0281] In one implementation, the data allocation strategy is the answer stratification strategy; any text block in the set of text blocks is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero. When the processing unit 1302 is configured to perform data allocation processing on the set of data pairs according to the data allocation strategy to obtain a first target set and a second target set, it is specifically configured to:
[0282] Determine the position information of the answer corresponding to the question text in each target data pair in the target text block;
[0283] Perform stratification processing on the target text block according to the position information of the answer corresponding to the question text in each target data pair in the target text block to obtain multiple text sub-layers corresponding to the target text block; one text sub-layer corresponds to at least one of the Q target data pairs;
[0284] Select reference data pairs from the target data pairs corresponding to each text sub-layer among the multiple text sub-layers and add them to the first target set, and add the target data pairs other than the selected reference data pairs among the multiple text sub-layers to the second target set.
[0285] In one implementation, the data distribution strategy is a word embedding classification strategy; any text block in the text block set is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero, and the question text in the target data pair consists of one or more characters; the processing unit 1302 is specifically configured to perform data distribution processing on the data pair set according to the data distribution strategy to obtain a first target set and a second target set, and specifically includes:
[0286] Perform word vector representation on the Q question texts in the Q target data pairs respectively to obtain word vectors corresponding to each question text in the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts;
[0287] Perform clustering processing on the word vectors corresponding to the Q question texts to obtain one or more clustering groups; one clustering group includes target data pairs corresponding to one or more question texts whose vector distances meet the distance requirements;
[0288] Select reference data pairs from the one or more clustering groups and add them to the first target set, and add the target data pairs in the one or more clustering groups except the selected reference data pairs to the second target set.
[0289] In one implementation, the retrieval library includes feature vectors of P text blocks, where P is a positive integer and P is less than Q; the processing unit 1302 is further configured to:
[0290] Receive the target question text input by the object in the interactive dialogue scenario, and perform vector embedding processing on the target question text using the second retrieval model to obtain the feature vector of the target question text; the feature vector of the target question text is used to characterize the semantic information of the target question text;
[0291] Perform similarity calculation processing on the feature vector of the target question text and the feature vectors of the P text blocks in the retrieval library to obtain the matching scores between the P text blocks and the target question text respectively; the matching score is used to indicate the credibility of the text block as the answer source of the target question text;
[0292] Sort the P matching scores in descending order, and perform gradient operation on each matching score in the sorted score sequence to obtain multiple gradient information;
[0293] Based on the multiple gradient information, dynamically screen one or more matching scores in the score sequence whose gradient information meets the gradient descent condition, and use the text blocks corresponding to the one or more matching scores as the answer source of the target question text;
[0294] Generate a target answer for the target question text based on one or more text blocks corresponding to matching scores.
[0295] In one implementation, when the processing unit 1302 is used to generate a target answer for the target question text based on one or more text blocks corresponding to matching scores, it is specifically used for:
[0296] Perform task classification processing on the target question text to obtain a classification result, where the classification result is used to indicate the difficulty of answering the target question text;
[0297] According to the difficulty of answering indicated by the classification result, use an answer generation model that matches the answer difficulty to generate an initial answer to the target question text based on one or more text blocks corresponding to matching scores;
[0298] Perform answer review processing on the initial answer to obtain the target answer to the target question text; the target answer is the initial answer or the answer after the initial answer is fine-tuned.
[0299] In one implementation, the interactive dialogue scenario includes multiple rounds of interactive dialogue. Any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogue is represented as the target round of interactive dialogue; the processing unit 1302 is further used for:
[0300] Obtain the target question text input by the object in the target round of interactive dialogue, the historical dialogue data in the historical round of interactive dialogue, and the historical object data about the object; the historical object data is generated based on the historical dialogue data;
[0301] Rewrite the target text question based on the target question text, historical dialogue data, and historical object data to obtain a new target question text;
[0302] When the processing unit 1302 is used to perform vector embedding processing on the target question text using the second retrieval model, it is specifically used for:
[0303] Perform vector embedding processing on the new target question text using the second retrieval model.
[0304] According to an embodiment of the present application, Figure 13Each unit in the model training device shown can be separately or wholly combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units are realized by one unit. In other embodiments of this application, the model training device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units. According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding method shown in Figure 3 and Figure 10 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), an access storage medium (RAM), and a read-only storage medium (ROM), to construct the model training device shown in Figure 13 and to implement the model training method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0305] In the embodiments of the present application, after obtaining the knowledge documents in the vertical field, first, the knowledge documents can be processed by text segmentation to obtain a text block set including a plurality of text blocks; wherein each text block has a theme respectively, so as to ensure that the theme of each document block is clear, and each text block not only includes the reference text blocks belonging to the knowledge document, but also includes the general information for summarizing the text semantics of the reference text blocks, so that each text block has the advantage of clear logic. Then, it is supported to generate a data pair set including a plurality of data pairs based on the text block set, specifically the text blocks in the text block set; wherein each data pair is composed of a question text and a text block, and the question text in each data pair is written for the text block that matches it, which greatly improves the question productivity, and the text block of each data pair serves as the answer source for the question text that matches it, ensuring the authenticity and availability of the data pair. Finally, the first retrieval model can be trained with the data pair set to obtain a second retrieval model and a retrieval library (including the feature vectors obtained by performing vector embedding processing on each text block using the second retrieval model); during the training process, it is necessary to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold, so as to ensure that the features of the question text and the text block in the same data pair are similar, which is beneficial to mapping the features of the target question text to the retrieval library with better data quality during subsequent answer retrieval, and being able to find a text block with similar features from the retrieval library to generate the target answer for the target question text, improving the accuracy and professionalism of answer retrieval in the intelligent question and answer process.
[0306] Figure 14 FIG. shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Please refer to Figure 14 The computer device includes a processor 1401, a communication interface 1402, and a computer-readable storage medium 1403. Among them, the processor 1401, the communication interface 1402, and the computer-readable storage medium 1403 can be connected through a bus or other means. Among them, the communication interface 1402 is used to receive and send data. The computer-readable storage medium 1403 can be stored in the memory of the computer device. The computer-readable storage medium 1403 is used to store a computer program, and the computer program includes program instructions. The processor 1401 is used to execute the program instructions stored in the computer-readable storage medium 1403. The processor 1401 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0307] The embodiments of the present application also provide a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device. And, in this storage space, there are also stored one or more instructions suitable for being loaded and executed by the processor 1401. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0308] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor 1401 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned method embodiments of model training; in a specific implementation, the one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1401 to perform the following steps:
[0309] Obtain a knowledge document, and perform text segmentation processing on the knowledge document to obtain a set of text blocks; the set of text blocks includes multiple text blocks, each text block has a topic respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding summary information. The summary information is obtained by generalizing the text semantics of the corresponding reference text block.
[0310] Based on the set of text blocks and the summary information included in each text block, generate a set of data pairs; the set of data pairs includes multiple data pairs, and each data pair is composed of a matching question text and a text block; the question text in each data pair is written for the text block it matches, and the text block in each data pair serves as the answer source for the question text it matches.
[0311] Use the set of data pairs to train a first retrieval model to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors of each text block in the set of text blocks; wherein, the second retrieval model and the retrieval library are used to generate a target answer for the target question text in an interactive dialogue scenario.
[0312] In one implementation, the number of knowledge documents is at least one; when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to perform text segmentation processing on the knowledge documents to obtain a set of text blocks, the following steps are specifically executed:
[0313] Obtain the semantic information of the knowledge document, and perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document;
[0314] By performing semantic generalization on the text semantics of each reference text block in one or more reference text blocks, obtain the summary information corresponding to each reference text block;
[0315] Add the summary information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block form a set of text blocks.
[0316] In one implementation, the semantic information of the knowledge document includes the theme to which the knowledge document belongs, and the knowledge document is composed of one or more characters; when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document, the following steps are specifically executed:
[0317] Judge whether the theme to which the knowledge document belongs is single;
[0318] If the theme to which the knowledge document belongs is single, count the number of characters included in the knowledge document;
[0319] When the number of characters included in the knowledge document is less than the character number threshold, use the knowledge document as one reference text block; or,
[0320] When the number of characters included in the knowledge document is greater than or equal to the character number threshold, perform paragraph cutting processing on the knowledge document according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
[0321] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and further execute the following steps:
[0322] If the knowledge document includes at least two themes, perform hierarchical expression segmentation on the knowledge document to obtain at least two initial text blocks; each of the at least two initial text blocks has one theme;
[0323] Count the number of characters included in each of at least two initial text blocks, and use the initial text blocks with the number of characters less than the character number threshold as reference text blocks; and,
[0324] Perform paragraph cutting processing on the initial text blocks with the number of characters greater than or equal to the character number threshold according to the text logical relationship, to obtain multiple reference text blocks corresponding to the knowledge document.
[0325] In one implementation, the text logical relationships are respectively the general-part relationship and the parallel relationship; the general-part relationship indicates that the content structure of the text content is a general-part structure, and the parallel relationship indicates that the content structure of the text content is a parallel structure; the knowledge document or the initial text block to which the paragraph cutting processing is performed is represented as the text content; the paragraph cutting processing includes:
[0326] Perform paragraph cutting processing on the text content according to the text logical relationship of the text content, to obtain reference text blocks corresponding to the knowledge document;
[0327] Among them, when the text logical relationship of the text content is the general-part relationship, the text content is cut into a general description text block and one or more sub-description text blocks; the general description text block is the content with a generalizing effect in the text content, and the general description information corresponding to the general description text block is used to indicate the overall semantics expressed by the text content; the sub-description text block is the content with an explanatory effect on the general description text block in the text content, and the general description information corresponding to the sub-description text block is used to indicate the sub-description semantics expressed by the sub-description text block and the overall semantics expressed by the text content;
[0328] When the text logical relationship is the parallel relationship, the text content is cut into at least two sub-description text blocks.
[0329] In one implementation, one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1401 to perform the following steps:
[0330] Extract metadata for each text block in the text block set to obtain the metadata of each text block; the metadata of the text block is used to describe the data attributes of the text block, and the data attributes include: the retrieval domain to which the text block belongs, the source of the text block, and the subject to which the text block belongs;
[0331] Among them, the metadata of each text block is used to update the retrieval library and perform index retrieval for the text block; the update includes: when the data attributes of the text block in the retrieval library are updated, modify the attribute values of the corresponding data attributes in the metadata of the text block;
[0332] The index retrieval at least includes: retrieving, according to the retrieval requirements indicated by the target question text, text blocks in the metadata that meet the retrieval requirements from the metadata to generate an answer for the target question text in an interactive dialogue scenario; and, when any text block has been retrieved in the interactive dialogue scenario, retrieving text blocks in the metadata that have the same metadata as any text block from the metadata to generate an answer for the target question text.
[0333] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and execute to generate a set of data pairs based on the set of text blocks and the summary information included in each text block, the following steps are specifically executed:
[0334] Obtain the scenario information corresponding to the interactive dialogue scenario, where the scenario information includes the dialogue style, interaction mode, and interaction field; and,
[0335] Based on the scenario information, the set of text blocks, and the summary information included in each text block, determine the question style and the number of questions adapted to each text block;
[0336] According to the question style and the number of questions adapted to each text block, write questions for each text block to obtain one or more question texts corresponding to each text block;
[0337] Combine one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block;
[0338] Generate a set of data pairs based on one or more data pairs corresponding to each text block.
[0339] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and execute to determine the question style and the number of questions adapted to each text block based on the scenario information, the set of text blocks, and the summary information included in each text block, the following steps are specifically executed:
[0340] Mine key elements for each text block according to the set of text blocks and the summary information included in each text block to obtain the key elements of each text block;
[0341] Based on the scenario information and the key elements of each text block, determine the question style and the number of questions adapted to each text block;
[0342] Among them, the question style adapted to each text block is one or more, and each question style belongs to one of multiple dialogue styles; the dialogue styles include at least one of the following: factual style, interpretive style, reasoning style, evaluative style, and hypothetical style.
[0343] In one implementation, any text block in the text block set is represented as a target text block, and any data pair corresponding to the target text block is represented as a target data pair. The target data pair includes target question text that matches the target text block. When one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to generate a data pair set based on one or more data pairs corresponding to each text block, the following steps are specifically executed:
[0344] Using a key pair verification model, based on the target question text, extract a first answer that matches the target question text from the target text block, and position information of the first answer in the target text block;
[0345] According to the position information of the first answer in the target text block, extract a second answer indicated by the position information from the target text block;
[0346] Compare the first answer and the second answer, and add the target data pair to the data pair set based on the comparison result;
[0347] Among them, if the comparison result indicates that the first answer is the same as the second answer, the target text block is added to the data pair set; if the comparison result indicates that the first answer is different from the second answer, the target text block is not added to the data pair set.
[0348] In one implementation, the text block set includes candidate text blocks. A candidate text block refers to a text block whose format conforms to the Q&A structure. The candidate text block includes a question part and an answer part. When one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and further executed, the following steps are also performed:
[0349] Add the candidate text block to the data pair set; and,
[0350] Perform question generalization processing on the question part included in the candidate text block to generate generalized question text, and form a new data pair corresponding to the candidate text block with the generalized question text and the answer part included in the candidate text block, and add the new data pair to the data pair set.
[0351] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to train a first retrieval model using the data pair set to obtain a second retrieval model, the following steps are specifically executed:
[0352] Perform data allocation processing on the data pair set according to a data allocation strategy to obtain a first target set and a second target set; the data allocation strategy includes an answer stratification strategy and a word embedding classification strategy; the first target set is a validation set, and the second target set is a training set; or, the first target set is a training set, and the second target set is a validation set;
[0353] The first retrieval model is trained using a training set to obtain a trained first retrieval model;
[0354] The trained first retrieval model is tested using a validation set to obtain a second retrieval model.
[0355] In one implementation, the data allocation strategy is an answer stratification strategy; any text block in the text block set is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero; when one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to perform data allocation processing on the data pair set according to the data allocation strategy to obtain a first target set and a second target set, the following steps are specifically executed:
[0356] Determine the position information of the answer corresponding to the question text in each of the Q target data pairs in the target text block;
[0357] According to the position information of the answer corresponding to the question text in each target data pair in the target text block, perform stratification processing on the target text block to obtain multiple text sub-layers corresponding to the target text block; one text sub-layer corresponds to at least one of the Q target data pairs;
[0358] Select reference data pairs from the target data pairs corresponding to each text sub-layer among the multiple text sub-layers and add them to the first target set, and add the target data pairs other than the selected reference data pairs among the multiple text sub-layers to the second target set.
[0359] In one implementation, the data allocation strategy is a word embedding classification strategy; any text block in the text block set is represented as a target text block, the target text block corresponds to Q target data pairs, Q is an integer greater than zero, and the question text in the target data pair consists of one or more characters; the processing unit 1302, when performing data allocation processing on the data pair set according to the data allocation strategy to obtain a first target set and a second target set, is specifically used for:
[0360] Perform word vector representation on the Q question texts in the Q target data pairs respectively to obtain word vectors corresponding to each of the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts;
[0361] Perform clustering processing on the word vectors corresponding to the Q question texts to obtain one or more clustering groups; one clustering group includes target data pairs corresponding to one or more question texts whose vector distances meet the distance requirements;
[0362] Select reference data pairs from one or more clustering groups and add them to the first target set, and add target data pairs in the one or more clustering groups other than the selected reference data pairs to the second target set.
[0363] In one implementation, the retrieval library includes feature vectors of P text blocks, where P is a positive integer and P is less than Q; one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and further perform the following steps:
[0364] Receive the target question text input by the object in an interactive dialogue scenario, and perform vector embedding processing on the target question text using the second retrieval model to obtain the feature vector of the target question text; the feature vector of the target question text is used to represent the semantic information of the target question text;
[0365] Perform similarity calculation processing on the feature vector of the target question text and the feature vectors of the P text blocks in the retrieval library to obtain the matching scores between the P text blocks and the target question text respectively; the matching scores are used to indicate the credibility of the text blocks as the answer sources of the target question text;
[0366] Sort the P matching scores in descending order, and perform gradient operations on each matching score in the sorted score sequence to obtain multiple gradient information;
[0367] Based on the multiple gradient information, dynamically filter one or more matching scores in the score sequence whose gradient information meets the gradient descent condition, and use the text blocks corresponding to the one or more matching scores as the answer sources of the target question text;
[0368] Generate a target answer for the target question text based on the text blocks corresponding to the one or more matching scores.
[0369] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and when generating a target answer for the target question text based on the text blocks corresponding to the one or more matching scores, specifically perform the following steps:
[0370] Perform task classification processing on the target question text to obtain a classification result, and the classification result is used to indicate the reply difficulty of the target question text;
[0371] According to the reply difficulty indicated by the classification result, use an answer generation model matching the answer difficulty to generate an initial answer for the target question text based on the text blocks corresponding to the one or more matching scores;
[0372] Perform answer review processing on the initial answer to obtain the target answer for the target question text; the target answer is the initial answer or the answer after the initial answer is fine-tuned.
[0373] In one implementation, the interactive dialogue scenario includes multiple rounds of interactive dialogue. Any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogue is represented as the target round of interactive dialogue. One or more instructions in the computer-readable storage medium are loaded and executed by the processor 1401, and the following steps are further performed:
[0374] Obtain the target question text input by the object in the target round of interactive dialogue, the historical dialogue data in the historical round of interactive dialogue, and the historical object data about the object. The historical object data is generated based on the historical dialogue data;
[0375] Rewrite the target text question based on the target question text, the historical dialogue data, and the historical object data to obtain a new target question text;
[0376] One or more instructions in the computer-readable storage medium are loaded and executed by the processor 1401. When performing vector embedding processing on the target question text using the second retrieval model, the following steps are specifically performed:
[0377] Perform vector embedding processing on the new target question text using the second retrieval model.
[0378] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the computer device provided in the embodiments of the present application are similar to the principle of problem-solving and the beneficial effects of the model training method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity of description, they will not be repeated here.
[0379] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned model training method.
[0380] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of professional skill can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0381] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0382] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any technical person familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims described.
Claims
1. A model training method, characterized in that, it includes: Obtain a knowledge document, and perform text segmentation processing on the knowledge document to obtain a set of text blocks; the set of text blocks includes multiple text blocks, each text block has a theme respectively, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by summarizing the text semantics of the corresponding reference text block; Based on the set of text blocks and the summary information included in each text block, generate a set of data pairs; the set of data pairs includes multiple data pairs, and each data pair is composed of a matching question text and a text block; the question text in each data pair is written by formulating a question for the text block that matches it, and the text block in each data pair serves as the answer source for the question text that matches it; Use the set of data pairs to train a first retrieval model to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors of each text block in the set of text blocks; wherein, the second retrieval model and the retrieval library are used to generate a target answer for a target question text in an interactive dialogue scenario.
2. The method according to claim 1, characterized in that, the number of the knowledge documents is at least one; the performing text segmentation processing on the knowledge document to obtain a set of text blocks includes: Obtain the semantic information of the knowledge document, and perform segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document; By performing semantic summarization on the text semantics of each reference text block in the one or more reference text blocks, obtain the summary information corresponding to each reference text block; Add the summary information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block form a set of text blocks.
3. The method according to claim 2, characterized in that, the semantic information of the knowledge document includes the theme to which the knowledge document belongs, and the knowledge document is composed of one or more characters; the performing segmentation processing on the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document includes: Judge whether the theme to which the knowledge document belongs is single; If the theme to which the knowledge document belongs is single, then count the number of characters included in the knowledge document; When the number of characters included in the knowledge document is less than the character number threshold, use the knowledge document as a reference text block; or, When the number of characters included in the knowledge document is greater than or equal to the character number threshold, perform paragraph cutting processing on the knowledge document according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.
4. The method according to claim 3, wherein, the method further comprises: if the knowledge document includes at least two topics, performing hierarchical expression segmentation on the knowledge document to obtain at least two initial text blocks; each of the at least two initial text blocks has one topic; counting the number of characters included in each of the at least two initial text blocks, and taking the initial text blocks with the number of characters less than the character number threshold as reference text blocks; and, performing paragraph cutting processing on the initial text blocks with the number of characters greater than or equal to the character number threshold according to the text logical relationship to obtain a plurality of reference text blocks corresponding to the knowledge document.
5. The method according to claim 3 or 4, wherein, the text logical relationships are respectively the general - part relationship and the parallel relationship; the general - part relationship indicates that the content structure of the text content is a general - part structure, and the parallel relationship indicates that the content structure of the text content is a parallel structure; the knowledge document or the initial text block to be subjected to paragraph cutting processing is represented as text content; the paragraph cutting processing includes: performing paragraph cutting processing on the text content according to the text logical relationship of the text content to obtain a plurality of reference text blocks corresponding to the knowledge document; wherein, when the text logical relationship of the text content is the general - part relationship, the text content is cut into a general statement text block and one or more specific statement text blocks; the general statement text block is the content with a generalizing effect in the text content, and the general information corresponding to the general statement text block is used to indicate the overall semantics expressed by the text content; the specific statement text blocks are the content with an explanatory effect on the general statement text block in the text content, and the general information corresponding to the specific statement text blocks is used to indicate the specific semantics expressed by the specific statement text blocks and the overall semantics expressed by the text content; when the text logical relationship is the parallel relationship, the text content is cut into at least two specific statement text blocks.
6. The method according to claim 1 or 2, wherein, after performing text segmentation processing on the knowledge document to obtain a text block set, it further comprises: extracting metadata for each text block in the text block set to obtain the metadata of each text block; the metadata of the text block is used to describe the data attributes of the text block, and the data attributes include: the retrieval domain to which the text block belongs, the source of the text block, and the topic to which the text block belongs; wherein, the metadata of each text block is used to update the retrieval library and perform index retrieval for the text block; the update includes: when the data attributes of the text block in the retrieval library are updated, modifying the attribute values of the corresponding data attributes in the metadata of the text block. The index retrieval at least includes: retrieving, in an interactive dialogue scenario, text blocks in the metadata that meet the retrieval requirements indicated by the target question text to generate an answer for the target question text; and, when any text block has been retrieved in the interactive dialogue scenario, retrieving, from the metadata, text blocks in the metadata that are the same as the metadata of the any text block to generate an answer for the target question text.
7. The method according to claim 1, wherein, generating a set of data pairs based on the set of text blocks and the summary information included in each text block includes: obtaining the scenario information corresponding to the interactive dialogue scenario, where the scenario information includes a dialogue style, an interaction mode, and an interaction field; and, determining, based on the scenario information, the set of text blocks, and the summary information included in each text block, a question style and a number of questions adapted to each text block; writing questions for each text block according to the question style and the number of questions adapted to each text block to obtain one or more question texts corresponding to each text block; combining the one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block; generating a set of data pairs based on the one or more data pairs corresponding to each text block.
8. The method according to claim 7, wherein, determining, based on the scenario information, the set of text blocks, and the summary information included in each text block, a question style and a number of questions adapted to each text block includes: mining key elements for each text block according to the set of text blocks and the summary information included in each text block to obtain the key elements of each text block; determining, based on the scenario information and the key elements of each text block, a question style and a number of questions adapted to each text block; wherein, the question style adapted to each text block is one or more, and each question style belongs to one of multiple dialogue styles; the dialogue styles include at least one of the following: factual style, explanatory style, reasoning style, evaluative style, and hypothetical style.
9. The method according to claim 7, wherein, any text block in the set of text blocks is represented as a target text block, and any data pair corresponding to the target text block is represented as a target data pair, and the target data pair includes a target question text matching the target text block; generating a set of data pairs based on the one or more data pairs corresponding to each text block includes: using a key pair verification model to extract, based on the target question text, a first answer matching the target question text from the target text block, and position information of the first answer in the target text block; extracting, according to the position information of the first answer in the target text block, a second answer indicated by the position information from the target text block. Compare the first answer and the second answer, and add the target data pair to the data pair set based on the comparison result; Wherein, if the comparison result indicates that the first answer is the same as the second answer, the target text block is added to the data pair set; if the comparison result indicates that the first answer is different from the second answer, the target text block is not added to the data pair set.
10. The method according to claim 7, characterized in that the text block set includes candidate text blocks, where a candidate text block refers to a text block whose format conforms to the Q&A structure, and the candidate text block includes a question part and an answer part; the method further includes: adding the candidate text block to the data pair set; and, performing question generalization processing on the question part included in the candidate text block to generate a generalized question text, and forming a new data pair corresponding to the candidate text block with the generalized question text and the answer part included in the candidate text block, and adding the new data pair to the data pair set.
11. The method according to claim 1, characterized in that training the first retrieval model using the data pair set to obtain a second retrieval model, including: performing data allocation processing on the data pair set according to a data allocation strategy to obtain a first target set and a second target set; the data allocation strategy includes an answer stratification strategy and a word embedding classification strategy; the first target set is a validation set, and the second target set is a training set; or, the first target set is a training set, and the second target set is a validation set; training the first retrieval model using the training set to obtain a trained first retrieval model; testing the trained first retrieval model using the validation set to obtain a second retrieval model.
12. The method according to claim 11, characterized in that the data allocation strategy is the answer stratification strategy; any text block in the text block set is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero; the performing data allocation processing on the data pair set according to the data allocation strategy to obtain a first target set and a second target set includes: determining the position information of the answer corresponding to the question text in each of the Q target data pairs in the target text block; performing stratification processing on the target text block according to the position information of the answer corresponding to the question text in each target data pair to obtain multiple text sub-layers corresponding to the target text block; one text sub-layer corresponds to at least one of the Q target data pairs; selecting reference data pairs from the target data pairs corresponding to each text sub-layer among the multiple text sub-layers and adding them to the first target set, and adding the target data pairs other than the selected reference data pairs among the multiple text sub-layers to the second target set.
13. The method according to claim 11, characterized in that The data allocation strategy is the word embedding classification strategy; any text block in the text block set is represented as a target text block, and the target text block corresponds to Q target data pairs, where Q is an integer greater than zero, and the question text in the target data pair consists of one or more characters; the data pair set is processed for data allocation according to the data allocation strategy to obtain a first target set and a second target set, including: Performing word vector representation on the Q question texts in the Q target data pairs respectively to obtain word vectors corresponding to each of the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts; Performing clustering processing on the word vectors corresponding to the Q question texts to obtain one or more clustering groups; one clustering group includes one or more target data pairs corresponding to question texts whose vector distances meet the distance requirements; Selecting reference data pairs from the one or more clustering groups and adding them to the first target set, and adding the target data pairs in the one or more clustering groups except the selected reference data pairs to the second target set.
14. The method according to claim 1, wherein, the retrieval library includes feature vectors of P text blocks, where P is a positive integer and P is less than Q; the method further includes: Receiving a target question text input by an object in an interactive dialogue scenario, and performing vector embedding processing on the target question text by using the second retrieval model to obtain a feature vector of the target question text; the feature vector of the target question text is used to characterize the semantic information of the target question text; Performing similarity calculation processing on the feature vector of the target question text and the P feature vectors of the text blocks in the retrieval library to obtain matching scores between each of the P text blocks and the target question text; the matching score is used to indicate the credibility of the text block as the answer source of the target question text; Sorting the P matching scores in descending order and performing gradient operation on each matching score in the sorted score sequence to obtain a plurality of gradient information; Based on the plurality of gradient information, dynamically screening one or more matching scores in the score sequence whose gradient information meets the gradient descent condition, and using the text blocks corresponding to the one or more matching scores as the answer source of the target question text; Generating a target answer for the target question text based on the text blocks corresponding to the one or more matching scores.
15. The method according to claim 14, wherein, generating a target answer for the target question text based on the text blocks corresponding to the one or more matching scores includes: Performing task classification processing on the target question text to obtain a classification result, and the classification result is used to indicate the answer difficulty of the target question text; According to the answer difficulty indicated by the classification result, using an answer generation model matching the answer difficulty to generate an initial answer to the target question text based on the text blocks corresponding to the one or more matching scores; Perform answer review processing on the initial answer to obtain the target answer to the target question text; the target answer is the initial answer or the answer after the initial answer is fine-tuned.
16. The method according to claim 14 or 15, characterized in that the interactive dialogue scenario includes multiple rounds of interactive dialogue, and any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogue is represented as the target round of interactive dialogue; after receiving the target question text input by the object in the interactive dialogue scenario, it further includes: obtain the target question text input by the object in the target round of interactive dialogue, the historical dialogue data in the historical round of interactive dialogue, and the historical object data about the object; the historical object data is generated based on the historical dialogue data; rewrite the target text question based on the target question text, the historical dialogue data, and the historical object data to obtain a new target question text; the vector embedding process of the target question text using the second retrieval model includes: performing vector embedding processing on the new target question text using the second retrieval model.
17. A model training device, characterized in that it includes: an acquisition unit for acquiring a knowledge document and performing text segmentation processing on the knowledge document to obtain a set of text blocks; the set of text blocks includes multiple text blocks, each text block has a theme, and each text block consists of a reference text block belonging to the knowledge document and its corresponding summary information, and the summary information is obtained by summarizing the text semantics of the corresponding reference text block; a processing unit for generating a set of data pairs based on the set of text blocks and the summary information included in each text block; the set of data pairs includes multiple data pairs, and each data pair consists of a matching question text and a text block; the question text in each data pair is written for the text block it matches, and the text block in each data pair serves as the answer source for the question text it matches; the processing unit is further configured to train a first retrieval model using the set of data pairs to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes the feature vectors of each text block in the set of text blocks; wherein, the second retrieval model and the retrieval library are used to generate a target answer for the target question text in the interactive dialogue scenario.
18. A computer device, characterized in that it includes: a processor adapted to execute a computer program; a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, it implements the model training method according to any one of claims 1-16.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the model training method according to any one of claims 1-16.
20. A computer program product, characterized in that the computer program product includes computer instructions, and when the computer instructions are executed by a processor, the model training method according to any one of claims 1-16 is implemented.
Citation Information
Cited By
Water resource scheduling instruction fine tuning data set construction method and device, equipment and medium
CN121745319A