Retrieval model training method and apparatus and computer device

Through text segmentation and data generation technology, the quality of the search model and search library is improved, the problem of poor answer retrieval performance in the existing technology is solved, and more accurate and detailed answer provision is achieved.

WO2025102869A1PCT designated stage expired Publication Date: 2025-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/112266
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-08-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The existing search models have poor answer retrieval performance during the intelligent question and answer process, and the data quality of the search library is poor, resulting in the inability to provide accurate and detailed answers.

Method used

By obtaining knowledge documents, text segmentation is performed to generate a collection of text blocks. Each text block consists of reference text blocks and general information. A data set is generated based on these text blocks, and the data set is used to train the search model to obtain a higher quality search model and search library.

Benefits of technology

Improve the answer retrieval performance of the search model and the data quality of the search library, and can provide more accurate and detailed answers to user questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112266_22052025_PF_FP_ABST
    Figure CN2024112266_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a retrieval model training method and apparatus and a computer device. The method comprises: acquiring a knowledge document, and performing text segmentation processing on the knowledge document to obtain a text block set (501); on the basis of the text block set and summary information comprised in each text block, writing a question for each text block to obtain one or more question texts corresponding to each text block, separately combining the one or more question texts corresponding to each text block with the corresponding text block to obtain one or more data pairs corresponding to each text block and, on the basis of the one or more data pairs corresponding to each text block, generating a data pair set (502); and using the data pair set to train a first retrieval model so as to obtain a second retrieval model (503).
Need to check novelty before this filing date? Find Prior Art

Description

Retrieval model training method, device and computer equipment

[0001] Related applications

[0002] This application claims priority to Chinese patent application No. 2023115290434, filed on November 15, 2023, entitled “A model training method, device, equipment, medium and program product,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computer technology, in particular to the field of artificial intelligence, and specifically to a retrieval model training method, a retrieval model training device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0004] Intelligent question answering refers to the process of answering questions raised by users in accurate and concise natural language. It is a direction in the field of natural language processing (NLP) that has attracted much attention and has broad development prospects.

[0005] Currently, retrieval models are used to retrieve answers to user questions from a retrieval database during intelligent question answering. However, existing retrieval models have been found to have poor answer retrieval performance and poor data quality in the retrieval database. This prevents retrieval models from providing accurate and detailed answers to user questions during intelligent question answering.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide a retrieval model training method, apparatus, computer equipment, computer-readable storage medium, and computer program product.

[0008] A retrieval model training method, executed by a computer device, comprising:

[0009] Acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set; the text block set includes multiple text blocks, each of which is composed of a reference text block belonging to the knowledge document and summary information corresponding to the reference text block, where the summary information is obtained by summarizing the text semantics of the corresponding reference text block;

[0010] Based on the text block set and the general information included in each text block, a question is written for each text block to obtain one or more question texts corresponding to each text block, the one or more question texts corresponding to each text block are respectively combined with the corresponding text block to obtain one or more data pairs corresponding to each text block, and based on the one or more data pairs corresponding to each text block, a data pair set is generated; each data pair consists of a matching question text and a text block; the text block in each data pair serves as the source of the answer to the matching question text; and

[0011] A first retrieval model is trained using a data pair set to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, which includes a feature vector of each text block in the text block set; wherein the second retrieval model and the retrieval library are used to generate a first answer for the first question text in an interactive dialogue scenario.

[0012] A retrieval model training device, comprising:

[0013] An acquisition unit is used to acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set; the text block set includes multiple text blocks, each of which is composed of a reference text block belonging to the knowledge document and general description information corresponding to the reference text block, where the general description information is obtained by summarizing the text semantics of the corresponding reference text block;

[0014] a processing unit configured to write a question for each text block based on the text block set and the general information included in each text block, obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block, obtain one or more data pairs corresponding to each text block, and generate a data pair set based on the one or more data pairs corresponding to each text block; each data pair consists of a matching question text and a text block; the text block in each data pair serves as a source of an answer to the matching question text; and

[0015] The processing unit is also used to train the first retrieval model using a set of data pairs to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, which includes a feature vector of each text block in the text block set; wherein the second retrieval model and the retrieval library are used to generate a first answer for the first question text in an interactive dialogue scenario.

[0016] A computer device includes a memory and one or more processors, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the one or more processors execute the steps of the above-mentioned retrieval model training method.

[0017] One or more non-volatile computer-readable storage media storing computer-readable instructions, on which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the above-mentioned retrieval model training method.

[0018] A computer program product or computer program, wherein the computer program product or computer program includes computer-readable instructions, wherein the computer-readable instructions are stored in a computer-readable storage medium, and one or more processors of a computer device read the computer-readable instructions from the computer-readable storage medium, and the one or more processors execute the computer-readable instructions, so that the computer device performs the steps of the above-mentioned retrieval model training method.

[0019] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0021] FIG1 is a schematic diagram of an intelligent question-answering scenario provided by an exemplary embodiment of the present application;

[0022] FIG2 is a schematic diagram of another intelligent question-answering scenario provided by an exemplary embodiment of the present application;

[0023] FIG3 is a schematic diagram of another intelligent question-answering scenario provided by an exemplary embodiment of the present application;

[0024] FIG4 is a schematic diagram of a model training process provided by an exemplary embodiment of the present application;

[0025] FIG5 is a flow chart of a model training method provided by an exemplary embodiment of the present application;

[0026] FIG6 is a flowchart of a text segmentation process provided by an exemplary embodiment of the present application;

[0027] FIG7 is a schematic diagram of a text segmentation effect provided by an exemplary embodiment of the present application;

[0028] FIG8 is a schematic diagram of a key-value pair verification provided by an exemplary embodiment of the present application;

[0029] FIG9 is a flowchart of text segmentation, key-value pair generation and verification provided by an exemplary embodiment of the present application;

[0030] FIG10 is a schematic diagram of a text distribution provided by an exemplary embodiment of the present application;

[0031] FIG11 is a schematic diagram of another text distribution provided by an exemplary embodiment of the present application;

[0032] FIG12 is a flow chart of a model application provided by an exemplary embodiment of the present application;

[0033] FIG13 is a flow chart of another model training method provided by an exemplary embodiment of the present application;

[0034] FIG14 is a schematic diagram of an intelligent question-answering interaction provided by an exemplary embodiment of the present application;

[0035] FIG15 is a schematic diagram of matching a feature vector of a first question text with a search library in a model application process provided by an exemplary embodiment of the present application;

[0036] FIG16 is a schematic structural diagram of a model training device provided by an exemplary embodiment of the present application;

[0037] FIG17 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0039] In the embodiment of the present application, a retrieval model training scheme is proposed, specifically providing a model training and application scheme for a retrieval model based on artificial intelligence and intelligent question answering. The following briefly introduces the relevant concepts involved in the retrieval model training scheme provided in the embodiment of the present application.

[0040] Intelligent question answering (Q&A), also known as open-ended Q&A or interactive dialogue, belongs to the field of human-computer interaction and is an advanced form of information retrieval system. Its primary purpose is to receive user questions through a computer device and provide accurate and concise natural language responses. With the rapid development and application of artificial intelligence (AI), intelligent question answering technology has become a highly sought-after and promising area of ​​NLP (Natural Language Processing). Intelligent question answering systems primarily consist of three components: question understanding, knowledge retrieval, and answer generation. Question understanding primarily involves techniques such as question classification and keyword extraction, aiming to enable the computer device to understand the semantics of the user-entered question. Knowledge retrieval primarily involves structured and unstructured information retrieval, aiming to retrieve relevant knowledge points from the database based on the understood question. Answer generation primarily involves answer extraction and answer verification, aiming to generate the appropriate answer to the question from the retrieved knowledge points. The core issues of intelligent question answering technology lie in properly understanding the question and ensuring the matching between the question and the answer.

[0041] In actual applications, the intelligent question-answering system has problems such as low question-answer matching, which leads to poor retrieval performance and low accuracy of answers. For example, due to the lack of relevant vertical knowledge in a certain field, the intelligent question-answering system cannot accurately record information on the corpus (i.e., language corpus). Therefore, in vertical applications, it is impossible to directly answer specific questions (such as asking whether an insurance product can cover what disease, the intelligent question-answering system has no relevant field knowledge, and its answer is an illusion (i.e., the answer made up by the intelligent question-answering system itself), not the correct answer). In order to improve the retrieval performance of the intelligent question-answering system and improve the accuracy of answer retrieval, the retrieval model training scheme provided in the embodiment of the present application is mainly aimed at vertical scenarios, combining the powerful text understanding ability of GPT (Generative Pre-trained Transformer, generative pre-trained Transformer model) and the local retrieval method of the traditional DPR (Dense Passage Retrieval, deep text retrieval model) algorithm to form a new intelligent question-answering system for retrieval library generation, training, recall and multi-round question and answer.

[0042] Among them, the core content of the retrieval model training solution provided by the embodiment of the present application may include the creation, update and application of the retrieval library (that is, the application of the retrieval library to retrieve the answers in the intelligent question and answer process). Among them, the creation of the retrieval library may refer to: based on the language corpus of the vertical field, a second retrieval model is trained for the first retrieval model, and a retrieval library is constructed based on the second retrieval model and the language corpus. The update of the retrieval library may refer to: the design of the metadata structure for the corpus, so as to facilitate the addition, subtraction and replacement of data for the retrieval library. The application of the retrieval library may include but is not limited to: general problem processing (such as single-round question retrieval) and user-customized problem processing (such as multi-round question retrieval) based on the second retrieval model and the retrieval library. The following is a brief introduction to the main key technologies and innovative methods of the retrieval model training solution:

[0043] (1) Construction of the second retrieval model and retrieval library.

[0044] The second retrieval model is a retrieval model obtained by training the first retrieval model using training data, while the first retrieval model is a pre-trained model. The retrieval library can be simply understood as a database of text blocks that store the source of answers in the intelligent question-answering system.

[0045] The construction of the second retrieval model and retrieval library mainly includes: text segmentation, data pair construction and retrieval model training. Specifically, first, obtain the knowledge documents in the vertical field. The knowledge documents can be understood as a large amount of corpus belonging to the vertical field. In this way, the knowledge documents can be processed by text segmentation (such as form conversion, knowledge induction, document segmentation and document summarization, etc.) by utilizing the full-text understanding ability and inductive summarization ability of the GPT model to obtain a text block set (or called a search text block set). The text block set includes multiple text blocks, each text block has a single theme, and each text block consists of a reference text block belonging to the knowledge document and the general description information corresponding to the reference text block. The general description information is obtained by summarizing the text semantics of the corresponding reference text block. Then, leveraging the GPT model's content summarization and innovative writing capabilities, one or more data pairs (or key-value pairs) corresponding to each text block in the text block set are constructed based on the key elements of each text block. The language styles of different data pairs corresponding to the same text block can be different, and the question text in each data pair is obtained by writing questions based on the text block that matches the question text. The text block in each data pair serves as the source of the answer to the question text that matches the text block. Next, utilizing the LLM (Large Language Model) model's content screening capabilities, as well as its dialogue and communication capabilities, the data pairs corresponding to each text block are quality checked to obtain a data pair set consisting of data pairs that meet quality requirements. This ensures that the text block can cover the answer to the question, that is, that the text block in the data pair is the source of the answer to the corresponding question text, and that the location information of the answer to the question text can be located in the text block. Finally, the data pair set is used to train the first retrieval model, resulting in a trained second retrieval model and retrieval library.

[0046] In addition, when constructing a text block collection, it also supports the use of a metadata structure to store text blocks, so as to realize the management of the retrieval library through the metadata structure, such as the update of the retrieval library. Among them, metadata (Metadata), also known as intermediary data or relay data, is a kind of data (data about data) for describing data. It is mainly information that describes the properties of data, and is used to support functions such as indicating storage location, historical data, resource search, and file records. In this way, by designing a metadata structure for text blocks, operations such as associated search, search priority, and convenient addition, subtraction, and replacement of the retrieval library can be realized. It should be noted that in subsequent specific embodiments, the contents of updating the retrieval library and associated search of text blocks based on metadata will be introduced in detail, and only a brief explanation is given here.

[0047] (2) Dealing with general issues - application of retrieval database.

[0048] After the second retrieval model and the retrieval library are constructed based on the above implementation method (1), in the actual retrieval process, if the subject has the need to retrieve the answer to the question, the subject can input the first question text into the intelligent question-answering system, and the intelligent question-answering system calls the second retrieval model to perform vector embedding processing on the first question text to obtain the vector representation (or feature vector) of the first question text. In this way, the intelligent question-answering system can perform similarity matching between the vector representation of the first question text and the feature vector of each of the multiple text blocks stored in the retrieval library, so as to match the first question text to the multiple text blocks, and thus generate a matching first answer for the first question text based on the multiple text blocks.

[0049] (3) Processing of user customization issues - application of retrieval library.

[0050] In a multi-round question-and-answer scenario, the embodiments of the present application support constructing historical object data about an object based on the object's historical conversation data. This historical object data can, to a certain extent, characterize the object's query content, query direction, and query style. In this way, the GPT model is used to extract historical object data and fill slots (or simply fill slots) in each round of conversation to keep the object data refreshed, maintain the accuracy and timeliness of the object data, and match the object situation in real time, which is conducive to achieving personalized recommendations and customized question consultation for the object.

[0051] Thus, it can be seen that, on the one hand, the retrieval model training scheme provided by the embodiment of the present application not only ensures that the theme of the text block is single and has general information during text segmentation, thereby ensuring good quality of the text block. In addition, matching questions are generated directly based on the text block to construct data pairs, effectively improving the authenticity and data quality of the data pairs. The data quality is reflected in the fact that the answer to the question text can be extracted from the text block. In this way, when training the retrieval model based on high-quality data pairs, the retrieval model is trained in the direction of ensuring that the feature difference of the same data pair is reduced, so that the retrieval model has good feature representation capabilities for both the question text and the text block, thereby obtaining a second retrieval model with better vector expression capabilities and a retrieval library with better data quality. On the other hand, the embodiment of the present application fully utilizes the rich functions of the GPT model, and by constructing a general retrieval library and a real-time object data capture synchronization mechanism based on GPT, it achieves efficient processing of general questions (i.e., questions in a single round of question and answer) and user-customized questions (i.e., questions in multiple rounds of question and answer). For general questions, the answer is retrieved (specifically, the text block is retrieved and the answer is generated based on the text block) by calculating the similarity between the object's question and the feature vector of the text block in the retrieval library. For user-customized questions, the object data of the object is extracted through the GPT model, and conditional retrieval is performed based on the object data and the first question text of the object, so as to achieve personalized question responses for different objects, thereby improving the accuracy of the intelligent question-answering system and the experience of object question retrieval.

[0052] The intelligent question-and-answer system provided in the embodiments of this application is an automated question-and-answer solution based on artificial intelligence technology that can understand, parse, and answer user questions. This makes the retrieval model training solution provided in the embodiments of this application applicable to a variety of interactive dialogue scenarios that require the use of intelligent question-and-answer systems to answer inquiries. Interactive dialogue scenarios may include, but are not limited to: ① Customer Support: The intelligent question-and-answer AI system can serve as a customer support tool to answer common user questions, reduce the workload of customer service staff, and improve customer satisfaction. ② Internal Enterprise Knowledge Base: Enterprises can use the intelligent question-and-answer AI system to build internal knowledge bases to help employees quickly find the information they need and improve work efficiency. ③ Virtual Assistant: The intelligent question-and-answer AI system can serve as a virtual assistant for individuals or enterprises, providing functions such as daily task management, scheduling, and reminder services. ④ Online Education: The intelligent question-and-answer AI system can be applied in the online education field to provide students with personalized learning resources and real-time Q&A services. ⑤ E-commerce: The intelligent question-and-answer AI system can help users answer questions during the shopping process, provide shopping recommendations, and enhance the shopping experience. ⑥ Financial Services: The intelligent question-and-answer AI system can provide real-time consulting services to customers of financial institutions such as banks and insurance companies, answering questions about accounts, transactions, and products. ⑦ Medical Consultation: The AI ​​Q&A system can provide patients with basic medical consultation services, answering questions about illnesses, treatments, medications, and more. ⑧ Tourism Consultation: The AI ​​Q&A system can provide tourists with real-time travel information, answering questions about attractions, hotels, transportation, and more. ⑨ News and Information Retrieval: The AI ​​Q&A system can help users quickly find the news and information they need, improving information retrieval efficiency.

[0053] It should be understood that the above description is only an exemplary product performance and interactive dialogue scenario provided in the embodiments of the present application, and does not limit the product performance and interactive dialogue scenario of the retrieval model training solution provided in the embodiments of the present application. The intelligent question-answering system provided in the embodiments of the present application can provide efficient, accurate and convenient question-answering services in various interactive dialogue scenarios, showing high value and practicality in various interactive dialogue scenarios, and helping to improve user experience and satisfaction.

[0054] To facilitate understanding of the retrieval model training scheme provided in the embodiment of the present application, the interactive dialogue scenario involved in the embodiment of the present application is briefly introduced below in combination with the scenario diagram shown in Figure 1: As shown in Figure 1, the system includes an object 101, a terminal 102 and a server 103. The embodiment of the present application does not limit the number and naming of objects 101, terminals 102 and servers 103.

[0055] Among them, the terminal 102 can refer to a terminal device with an interactive dialogue function, and the object 101 can conduct a single or multiple rounds of interactive dialogue with the terminal 102 to obtain relevant knowledge. The terminal 102 can include but is not limited to: a smartphone (such as a smartphone deploying the Android system, or a smartphone deploying the Internetworking Operating System (IOS)), a tablet computer, a portable personal computer, a mobile Internet device (Mobile Internet Devices, MID), a vehicle-mounted device, a head-mounted device, an intelligent chat robot and an aircraft. The embodiment of this application does not limit the type of terminal device, which is explained here. The server 103 is a server corresponding to the terminal 102, which is used to interact with the terminal 102 for data to provide computing and application service support for the terminal 102. The server 103 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 102 and the server 103 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application.

[0056] In a specific implementation, the retrieval model training object uses the server 103 to obtain knowledge documents in vertical scenarios / fields (such as education, medical care, insurance, entertainment, banking, and real estate) from different platforms or systems through the network. Then, the server 103 can perform document segmentation processing, build a data pair set, train the model, and build a retrieval library on the knowledge documents in accordance with the retrieval model training solution provided in the embodiment of the application to obtain a trained second retrieval model and retrieval library.

[0057] Furthermore, the trained second retrieval model and retrieval library can be directly deployed in the server 103. In this way, when the subject initiates an interactive dialogue through the terminal 102, the subject's first question text can be sent through the terminal 102 to the server 103 where the trained second retrieval model is deployed. In this way, the server 103 uses the trained second retrieval model to perform vector embedding processing on the first question text to obtain a vector representation of the first question text. Then, based on the vector representation of the first question text, the server 103 can retrieve one or more text blocks that match the vector representation of the first question text from the retrieval library, thereby generating a corresponding first answer for the first question text based on the one or more text blocks. Finally, the server 103 sends the first answer to the terminal 102 so that the subject can obtain the first answer to the first question text through the terminal 102.

[0058] Of course, the trained second retrieval model can also be deployed in the terminal 102. In this case, the interactive dialogue process will vary depending on the deployment location of the retrieval library. Optionally, as shown in Figure 2, when the retrieval library is maintained by the server 103, the terminal 102 receives the first question text of the object and calls the deployed second retrieval model to perform vector embedding processing on the first question text. After obtaining the vector representation of the first question text, the vector representation can be sent to the server 103. The server 103 retrieves one or more text blocks that match the vector representation of the first question text from the retrieval library based on the vector representation of the first question text. Then, the server 103 can return the one or more text blocks to the terminal 102, and the terminal 102 generates and outputs the corresponding first answer for the first question text based on the one or more text blocks. It is worth noting that the process of generating the corresponding first answer for the first question text based on the one or more text blocks described above can also be performed by the server 103. In this case, the server 103 directly outputs the first answer to the terminal 102 for display, which can alleviate the burden on the terminal 102 to a certain extent. Optionally, as shown in Figure 3, when the retrieval library is maintained by the terminal 102, after receiving the first question text of the object and calling the deployed second retrieval model to perform vector embedding processing on the first question text, the terminal 102 can directly execute the retrieval of one or more text blocks that match the vector representation of the first question text from the retrieval library based on the vector representation of the first question text, and generate and output a corresponding first answer for the first question text based on the one or more text blocks.

[0059] Furthermore, the trained second retrieval model can be deployed in the terminal 102 as a plug-in or application. For example, if the trained second retrieval model is deployed in the terminal 102 as a system-level plug-in, any application deployed in the terminal 102 can call the plug-in to implement an interactive dialogue and provide a first answer to the subject's first question text. For another example, if the trained second retrieval model is deployed in a specific application, then after launching the application through the terminal 102, the trained second retrieval model can be called in the application to implement an interactive dialogue. An application can refer to a computer program designed to perform one or more specific tasks. By categorizing applications according to different dimensions (such as the application's operating mode and functionality), the types of the same application under different dimensions can be obtained. For example, according to the application's operating mode, applications may include but are not limited to: clients installed in the terminal, mini-programs (which are subprograms of the client) that can be used without downloading and installing, web (World Wide Web) applications opened through a browser, and so on. For another example, according to the application's functional type, applications may include but are not limited to: IM (Instant Messaging) applications, content interaction applications, and so on. Instant messaging applications refer to internet-based applications for instant messaging and social interaction. These applications may include, but are not limited to, social applications with messaging functionality, map applications with social interaction functionality, and gaming applications. Content interaction applications refer to applications that enable content interaction, such as online banking, sharing platforms, personal spaces, and news applications. This embodiment of the present application does not limit the specific type of application running on terminal 102 and deployed with the trained second retrieval model, but is described herein.

[0060] It should be noted that the above Figures 1, 2 and 3 are only schematic diagrams of the exemplary scenario architecture provided in the embodiments of the present application; in actual applications, the architecture can undergo adaptive changes.

[0061] It should also be noted that the collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information requires the knowledge or consent of the individual subject (or the presence of a legal basis for information acquisition), and subsequent data use and processing shall be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as when obtaining the first question text of the object, the user's permission or consent must be obtained, and the collection, use and processing of relevant data (such as the collection and release of the barrage posted by the object, etc.) must comply with the relevant laws, regulations and standards of the relevant region.

[0062] Based on the retrieval model training scheme described above, the embodiment of the present application proposes a more detailed retrieval model training method. The retrieval model training method proposed in the embodiment of the present application will be introduced in detail below with reference to the accompanying drawings. As can be seen from the aforementioned related description, the retrieval model training method provided in the embodiment of the present application mainly includes two parts: retrieval model training and model application for the retrieval model, wherein the retrieval model training part also includes the construction of a retrieval library. For ease of understanding, different embodiments will be used later to introduce the specific implementation processes of retrieval model training and model application.

[0063] Please refer to Figure 4, which is a flowchart of a retrieval model training method provided by an exemplary embodiment of the present application. This flowchart mainly provides the overall training process of the retrieval model training from the perspective of retrieval model training. As shown in Figure 4, the general process of the retrieval model training method may include:

[0064] First, the scenario information of the interactive dialogue scenario is determined based on the type of intelligent question-answering system (such as intelligent robots in shopping malls, food delivery robots in restaurants, insurance inquiry robots, etc.). Then, based on the scenario information of the interactive dialogue scenario, knowledge documents related to the scenario information are collected and organized as a corpus for retrieval model training and retrieval library to overcome the shortcomings of the GPT model that does not have relevant knowledge of vertical scenarios and cannot achieve accurate information recording and recommendation. Among them, the scenario information of the interactive dialogue scenario may include but is not limited to: ① The functional positioning (or interactive field) of the intelligent question-answering system, mainly including product positioning and user groups, so that the purpose of the questions and the questioning angle of the intelligent question-answering system can be determined according to the functional positioning. For example: for intelligent insurance consulting AI question-answering products, the user group with inquiry needs needs needs to be positioned as users who need to buy insurance and have questions about insurance products. ②Interaction style, or conversation style, mainly refers to the type or style of questions asked during human-computer interaction. Conversational styles can include at least one of the following: factual style (or factual questions, which require the intelligent question-answering system to answer questions raised by users according to actual facts), explanatory style (or explanatory questions, which require the intelligent question-answering system to explain questions raised by users), inferential style (or inferential questions, which require the intelligent question-answering system to have certain reasoning capabilities to answer questions raised by users), evaluative style (or evaluative questions, which require the intelligent question-answering system to give an evaluative answer to a certain point in the questions raised by users), and hypothetical style (or hypothetical questions, which require the intelligent question-answering system to answer questions raised by users in a hypothetical manner). ③Interaction mode refers to the conversation mode used in the human-computer interaction process. For example, the interaction mode between the user and the intelligent question-answering system may include, but is not limited to, text-to-voice interaction mode, text-to-text interaction mode, voice-to-voice interaction mode, and voice-to-text interaction mode. For example, the text-to-speech interaction mode indicates that users can enter their own questions in the intelligent question-answering system in the form of text input, and when the intelligent question-answering system answers the questions raised by the users, the answers are played through voice broadcast.

[0065] Then, the GPT model is used to judge the document accuracy and question-answer completeness for the knowledge document, specifically, document segmentation (or called cutting), rewriting and summarization are performed on the knowledge document to divide the complete knowledge document into manageable fragments that can be processed separately (which can be called text blocks in the embodiment of this application). This is conducive to separate processing of text blocks of limited length, which greatly facilitates the management of text blocks. Among them, the logic of document accuracy and question-answer completeness judgment for knowledge documents supported by the embodiment of this application can roughly include: ① Topic consistency, that is, extracting each text block according to the theme, that is, it is necessary to ensure that each text block after segmentation has a theme to ensure the singleness of the theme of the text block, so that each text block has the advantage of a clear theme. ② Word count limit, considering that the input text length that the GPT model can accept is limited and it cannot accept a knowledge base that is too long or too large, it cannot directly answer specific questions in vertical applications. Therefore, this application needs to ensure that the number of characters included in each text block obtained by segmentation is less than the number threshold, and the number threshold is specifically determined by the number of words allowed by the GPT model. ③ Logical relationship, the embodiment of the present application supports summarizing the text semantics of each reference text block (i.e., obtained by directly segmenting from the knowledge document) according to the text logical relationship of the knowledge document (such as the general-specific relationship or the parallel relationship, etc.) to obtain the general information of each reference text block, and adding the general information to the corresponding reference text block to obtain the text block. In this way, the semantics of each text block in the knowledge document can be clarified, and the clarity of the logical relationship of the knowledge document as a whole can be improved. ④ Question and answer structure, the embodiment of the present application supports directly storing the text block originally structured as a question and answer structure in the form of a question and answer structure. Subsequently, the text block of the question and answer structure can be directly used to generalize the question to generate the question variant corresponding to the text block of the question and answer structure, thereby improving the richness of the data pair and gradually improving the retrieval accuracy. It can be seen that the embodiment of the present application uses the language ability of the GPT model for the first time to design text blocks for vertically related knowledge documents, and provides a new text segmentation method, so that the text blocks have the same important knowledge points and complete information, clear themes and clear logic. It not only ensures the coherence, integrity and relevance between text blocks, but also greatly improves the data quality of the retrieval library.

[0066] Secondly, after the knowledge document is segmented based on the above steps to obtain multiple text blocks, the embodiment of the present application supports the storage of metadata structure for each document block, so as to facilitate the subsequent update of the retrieval library based on metadata. In addition, it also supports question generation for each text block, specifically using the text block as the source of the answer, generating one or more question texts that match the text block, so that a data pair is obtained by combining each question text and the text block, thereby obtaining one or more data pairs corresponding to each text block. A data pair includes a question text and a text block, the question text is generated based on the text block, and the text block is the source of the answer to the question text. Furthermore, after obtaining one or more data pairs corresponding to each text block, in order to ensure that each data pair is not an illusion (i.e., a false construction of the GPT model), the embodiment of the present application also needs to perform key-value pair verification on each data pair to ensure that each data pair obtained by verification is real, thereby improving the authenticity and reliability of the corpus of the retrieval model training and the retrieval library.

[0067] Finally, the multiple data pairs (belonging to the data pair set) after the above processing are used to perform retrieval model training on the first retrieval model to obtain a trained first retrieval model (referred to as the second retrieval model in the embodiment of the present application) and a retrieval library. Among them, the method of determining the retrieval library may specifically include: optionally, obtaining the feature vector of the text block obtained by performing vector embedding processing on the text block in each data pair during the last iterative training in the retrieval model training process, and after deduplication processing on the feature vector of the text block, adding the feature vector of the text block to the retrieval library. Optionally, the second retrieval model is used to perform vector embedding processing on each text block after the above text segmentation processing, and the feature vector of the obtained text block is added to the retrieval library.

[0068] Based on the above-mentioned brief introduction of the overall process of retrieval model training in FIG4, the specific implementation process of retrieval model training is described in detail in conjunction with FIG5. FIG5 shows a flow chart of a model method provided by an exemplary embodiment of the present application. The retrieval model training method flow shown in FIG5 is mainly about the retrieval model training and the construction of the retrieval library of the retrieval model. The retrieval model training method can be executed by a computer device, which can be the server 103 shown in FIG1. ​​The retrieval model training method may include but is not limited to steps S501-S503:

[0069] S501: Acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set.

[0070] As described above, knowledge documents can be understood as a large amount of corpus belonging to vertical fields. The vertical fields here are also called vertical scenarios, vertical fields, which refer to specialized and segmented industries or market fields. These fields usually have specific needs, user groups and business models. Compared with broad general fields, vertical fields are more focused on specific market segments. Specifically, a vertical field can refer to a specific field / scenario, then the knowledge documents belonging to the vertical field can refer to the language corpus (referred to as corpus) belonging to that specific field. For example, if the vertical scenario is an insurance scenario, then the knowledge documents belonging to the vertical scenario may include insurance policies, insurance contracts and insurance clause texts related to insurance. For another example, if the vertical scenario is a financial scenario, then the knowledge documents belonging to the vertical scenario may include deposit rules related to funds or finance, financial product descriptions, etc. It should be understood that the embodiments of the present application do not limit the vertical scenarios, the number of knowledge documents belonging to the vertical scenarios (such as the number of knowledge documents is at least one, such as one or more insurance policies about medical insurance, etc.) and types. For the sake of convenience, the vertical scenario will be described as an insurance scenario, and the knowledge document will be a medical insurance policy belonging to the insurance scenario. This is specifically explained here. Furthermore, the embodiments of the present application do not limit the method of obtaining knowledge documents. For example, the method of obtaining knowledge documents may include but is not limited to: a computer device obtains from a platform or system related to the vertical scenario through a network, or a computer device directly obtains from content publicly available on the Internet, etc.

[0071] After obtaining the knowledge documents in vertical scenarios, considering that knowledge documents often have long content, such as medical insurance policies often have dozens or even hundreds of pages, and knowledge documents with too long content often include many knowledge points (such as insurance clauses about different types of diseases, etc.), if the entire knowledge document is directly used for retrieval model training and retrieval library construction, the model performance of the retrieval model and the data quality of the retrieval library will be poor due to factors such as the complex subject matter and rich content of the knowledge document.

[0072] To this end, the embodiment of the present application supports the use of a text segmenter to perform text segmentation processing on knowledge documents, so as to split knowledge documents with longer content into small blocks or fragments with less content (i.e., the text blocks mentioned above). In this way, the use of text blocks with less content is more conducive to the training of retrieval models and the construction of retrieval libraries. Among them: ① The text blocks involved in the embodiment of the present application can be represented as Chunks; in the fields of information retrieval and natural language processing, Chunks refer to a smaller fragment or part in a text, and in a retrieval library, Chunks refer to the division of larger documents or data sets into smaller, easier to process and analyze parts. In this way, dividing longer content into Chunks and then processing them can improve the efficiency of retrieval and analysis, while Chunks with less content are easier to extract valuable information. ② The text segmenter involved in the embodiment of the present application is an algorithm or method for splitting large text into smaller blocks or fragments; its goal is to create manageable fragments (i.e., text blocks) that can be processed individually, which is usually necessary when processing large documents or data sets.

[0073] In a specific implementation, the principle diagram of using a text segmenter to perform text segmentation processing on a knowledge document can be referred to the relevant content of the text block design part shown in FIG4 , and the specific implementation process of the text segmentation processing can be shown in FIG5 , including but not limited to steps s11-s13, wherein:

[0074] s11. Acquire the semantic information of the knowledge document, and segment the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document. The knowledge document is a text content composed of one or more characters, and the text content may also include images. The semantic information of the knowledge document can be used to indicate the semantics expressed by the knowledge document, such as the subject to which the knowledge document belongs, the specific content elaborated by the knowledge document, the text type to which the knowledge document belongs, and so on. The embodiment of the present application does not limit the method for acquiring the semantic information of the knowledge document. For example, the text segmenter provided in the embodiment of the present application has a semantic extraction function, so the text segmenter can be directly used to acquire the semantic information of the knowledge document. For another example, some existing semantic extraction networks (such as bag-of-words models, etc.) or tools can be used to extract the semantic information of the knowledge document.

[0075] Taking into account that knowledge documents are often long texts, and too long content will affect the subsequent retrieval and summarization process, in order to refine the knowledge points and meet the data quality requirements of the subsequent retrieval process, ensure that the subject of the segmented text block is clear, and improve the document accuracy of the text block (that is, the content expressed by each text block is single), the embodiment of the present application needs to ensure that each text block has only a single subject. The so-called text block has a single subject can be simply understood as that the semantics expressed by the text block is unique and not mixed. For example, if the content of the text block is "Let's go hiking tomorrow. The weather is very good today", the text block includes two themes, namely the theme of "Let's go hiking tomorrow" and the theme of "The weather is very good today". Therefore, after obtaining the semantic information of the knowledge document, the text segmenter will determine whether the subject to which the knowledge document belongs is single based on the semantic information of the knowledge document.

[0076] On the one hand, if the subject to which the knowledge document belongs is single, indicating that the knowledge document only expresses the same subject, then the number of characters included in the knowledge document is determined to ensure that the number of characters included in the segmented text block meets the requirements of the text segmenter, to avoid the text block content being too long and causing the text segmenter to be unable to accept it, and to avoid the retrieval library's requirements for the number of words in the text block (specifically, the retrieval library's requirements for the length of the feature vector embedding of the stored text block). When the number of characters included in the knowledge document is less than the character number threshold, the knowledge document is used as a reference text block; or, when the number of characters included in the knowledge document is greater than or equal to the character number threshold, the knowledge document is segmented according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document. Among them, the specific value of the character number threshold is related to the category of the text segmenter. For example, when the text segmenter is a GPT model, most of the time it is necessary to ensure that the number of words in the text block is within 512 tokens (token can be understood as a character unit, and a character unit can include one or more characters).

[0077] On the other hand, if the subject of the knowledge document is at least two, indicating that the subject of the knowledge document is not clear, the knowledge document is segmented hierarchically according to the subject type to obtain at least two initial text blocks, each of which has a subject. Then, the number of characters included in each of the at least two initial text blocks segmented by the hierarchical expression is counted, and the initial text block with a character count less than the character count threshold is used as a reference text block, and the initial text block with a character count greater than or equal to the character count threshold is subjected to paragraph segmentation according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document. That is to say, when it is determined that the subject of the knowledge document is not single, the knowledge document can be segmented hierarchically according to the subject type to obtain multiple initial text blocks with a single subject, and the number of characters for each initial text block is counted, and the initial text block with a character count less than the character count threshold is used as a reference text block, whereas the initial text block with a character count greater than or equal to the character count threshold is subjected to paragraph segmentation to obtain multiple reference text blocks corresponding to the knowledge document.

[0078] It can be understood that by segmenting the knowledge document multiple times based on thematic consistency and word count requirements to obtain multiple reference text blocks corresponding to the knowledge document, each segmented reference text block has the following characteristics: it contains the same important knowledge point (e.g., a single theme) and the word count is less than the character count threshold. In this way, when constructing subsequent data pairs and retrieval libraries based on multiple reference text blocks, the authenticity and rationality of the data pairs can be greatly improved, thereby improving the data quality of the retrieval library.

[0079] From the above two aspects, it can be seen that both the knowledge document and the initial text block may be segmented according to the text logical relationship. For ease of explanation, the embodiment of the present application represents the knowledge document or initial text block to be segmented as text content. The text logical relationship of the text content is a logical relationship used to characterize the overall structure of the text content, which can include: a general-specific relationship and a parallel relationship. The general-specific relationship indicates that the content structure of the text content is a general-specific structure; that is, the beginning of the text content is often a paragraph summarizing the entire text, and the paragraphs after the beginning of the text content are often some explanatory paragraphs of the paragraph summarizing the entire text. The parallel relationship indicates that the content structure of the text content is a parallel structure; that is, the various parts included in the text content have a parallel logical relationship. Therefore, depending on the different text logical relationships of the text content, the paragraph segmentation processing method for the text content varies, among which: ① When the text logical relationship of the text content is a general-specific relationship, the text segmenter can segment the text content into a general text block and one or more detailed text blocks according to the general-specific structure, that is, the text content is segmented into a general text block and one or more detailed text blocks. Among them, the general text block is the content in the text content that has a summarizing effect, that is, the part of the text content that has a summarizing effect, and the sub-text block is the content in the text content that has an explanatory effect on the general text block, such as a sub-text block that mainly explains a knowledge point under the general text block. ② When the logical relationship of the text is a parallel relationship, the text segmenter can split the text content into at least two sub-text blocks according to the parallel structure, that is, the text content is cut into at least two sub-text blocks; at least two sub-text blocks have an independent relationship (or independence), and the independent relationship here is reflected in that the content / viewpoints / facts expressed by at least two sub-text blocks are different, so that the independence between the sub-text blocks can be clearly identified in the subsequent analysis to avoid confusion in themes and logic.

[0080] It can be understood that by segmenting knowledge documents based on their logical relationships to obtain multiple reference text blocks corresponding to the knowledge documents, each segmented reference text block possesses logical clarity. This significantly improves the authenticity and rationality of data pairs and retrieval database construction based on these multiple reference text blocks, thereby enhancing the data quality of the retrieval database.

[0081] In summary, the embodiment of the present application mainly performs multiple segmentation of the knowledge document based on the dimensions of topic consistency, word count requirements, and text logical relationships to obtain multiple reference text blocks corresponding to the knowledge document; this allows each reference text block obtained by segmentation to have the following characteristics: all of them are the same important knowledge point (such as a single topic), the word count is less than the character count threshold, and the logic is clear. s12. By semantically summarizing the text semantics of each reference text block in one or more reference text blocks, the general description information corresponding to each reference text block is obtained.

[0082] s13. Add the summary information corresponding to each reference text block to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block. The text blocks corresponding to each reference text block constitute a text block set.

[0083] In steps s12-s13, it should be noted that the text segmentation processing provided in the embodiment of the present application is not only for segmentation of text content (such as knowledge documents or initial text blocks), but also adds a small summary and rewrite content (referred to as general information in the embodiment of the present application) to the corresponding reference text block based on the original text of the reference text block. The summary and rewrite content mainly describes the logical position of the reference text block in the overall document, so that the information of each text block is sufficiently complete. Taking the reference text blocks as the general description text block and the detailed description text block mentioned in step s11 as an example, the summary and rewrite content of the reference text block is introduced. Among them: the general description information obtained by the text segmenter for semantically summarizing the general description text block is used to indicate the overall semantics expressed by the text content corresponding to the general description text block, so that the generality of the general description text block can be clearly identified in subsequent analysis. Similarly, the general information obtained by the text segmenter through semantic summarization of the sub-text block is used to indicate the sub-semantics and logical structure expressed by the sub-text block (that is, the sub-text block has an explanatory effect on the general text block) and the overall semantics expressed by the text content to which the sub-text block belongs, so as to ensure the information integrity of each text block chunk that is finally cut.

[0084] Based on the above introduction to the general description information of the reference text block, the text segmenter of the embodiment of the present application performs semantic summarization on each of the multiple reference text blocks corresponding to the knowledge document, and can obtain the text block corresponding to each reference text block, thereby obtaining a text block set. Among them, the text block set includes multiple text blocks, each text block has a theme, and each text block is composed of a reference text block belonging to the knowledge document and its corresponding general description information (i.e. the summary rewritten content mentioned above), and the general description information is obtained by summarizing the text semantics of the corresponding reference text block. In addition, the embodiment of the present application supports adding the general description information of the reference text block to the target position in the reference text block to obtain the text block corresponding to the reference text block. The target position here can include: the beginning, middle or end of the paragraph of the reference text block, which is not limited to this.

[0085] It should be noted that the text segmenter that performs the text segmentation processing shown in the above steps s11-s13 can be a generative pre-trained model. Among them, from the above description of the generative pre-training model - GPT model, it can be seen that the GPT model has extremely strong full-text comprehension ability, summarization and summarization ability, can adapt to multiple languages ​​and document formats, and has a large amount of knowledge background at the bottom layer. The embodiment of the present application can use the GPT model as a semantic-level text segmenter to implement text segmentation processing for knowledge documents. In practical applications, the embodiment of the present application supports the use of the GPT model to implement the above-mentioned text segmentation processing process through Prompt guidance and fine-tuning. Among them: ① Prompt can be called a prompt word. In the AI ​​large model, the role of Prompt is mainly to prompt the AI ​​model (such as the GPT model involved in this application) the context of the input information and the parameter information of the input model, so that the AI ​​large model can realize the corresponding function under the guidance / prompt of Prompt; in the embodiment of the present application, the GPT model is mainly prompted by Prompt, so that the GPT model can perform the above-mentioned text segmentation processing process according to the Prompt prompt. ② Fine-tuning can refer to the process of pre-training the GPT model by using training data so that the GPT model has certain specific capabilities; for example, in the embodiment of the present application, the GPT model can be fine-tuned so that the GPT model has the ability to extract semantic information from knowledge documents.

[0086] In conjunction with Figure 7, and taking the text segmenter as the GPT model as an example, an exemplary process of using the GPT model to perform text segmentation processing on knowledge documents is introduced. As shown in Figure 7, assuming that the text logical relationship of the knowledge document is a general-specific relationship, then according to the processing logic of the text segmentation processing shown in Figure 6 above, the segmentation effect of the knowledge document guided by the GPT model through Prompt is as follows: through semantic understanding and segmentation through the general-specific relationship, the first paragraph of the knowledge document is divided into the first text block Chunk, and the first text block Chunk is a general text block, and its corresponding general information is used to characterize the logical structure (such as the summary part) and semantic information of the first text block Chunk in the knowledge document. Each paragraph after the first paragraph in the knowledge document is divided into a text block Chunk. Of course, if the topics of two adjacent paragraphs are the same and the number of characters is less than the character number threshold, then the same text block Chunk can also include at least two paragraphs. The embodiment of the present application does not limit the number of paragraphs included in the text block Chunk. In addition, each text block Chunk located after the first text block Chunk in the knowledge document has corresponding general description information, which is used to represent the semantics of the corresponding text block and the overall semantics of the entire knowledge document.

[0087] It should be understood that FIG7 illustrates the addition of summary information to a reference text block to generate a corresponding text block, using the example of adding the summary information to a front position in the reference text block (e.g., after the first few characters of the reference text block). In actual applications, the summary information may also be added to the beginning or end of a paragraph in the reference text block, without limitation.

[0088] Taking the text segmentation method provided in the embodiment of the present application and the traditional text segmentation as an example, the advantages of the embodiment of the present application are explained: the traditional text segmentation method may include: LangChain algorithm, nltk algorithm and spacy algorithm, etc. Among them, the traditional LangChain uses a character list as a parameter to put all paragraphs together as much as possible, resulting in the same topic, such as the summary paragraph in the total-specific structure being divided into multiple Chunks, resulting in too fine text block segmentation, causing information loss. The segmentation method in the traditional nltk algorithm is mainly based on punctuation marks (such as periods, question marks, exclamation marks, etc.) to segment the text into sentences, resulting in the original document having complex structure and language characteristics. The text cannot be segmented. The segmentation method in the traditional spacy algorithm is mainly based on rules and statistical methods, which realizes the word segmentation and sentence segmentation of the text by identifying spaces, punctuation marks, dependencies and grammatical rules of a specific language. It does not refer to the semantics of the document itself to achieve segmentation, resulting in unclear logic and information loss.

[0089] However, to ensure the information integrity and coherence of each reference text block in the knowledge document, the embodiment of the present application uses the GPT model to reasonably summarize and generalize each reference text block to obtain general information, and then adds the general information to the corresponding reference text block to provide a brief explanation of the corresponding reference text block. In this way, each text block has complete information and a reasonable knowledge density, and ensures that the feature vector corresponding to the text block is displayed as coherent information when it is placed in the search library, which is more conducive to data management and retrieval in the search library. That is to say, compared with some existing text segmentation methods that mainly rely on document format (such as line breaks, number of markdowns, punctuation marks, etc.) or use traditional NLP methods for segmentation, the segmentation processing method provided in the embodiment of the present application can effectively overcome information loss (such as in the text segmentation process, due to the use of rigid standards or traditional algorithms to cut documents without essentially understanding the meaning of the document, especially when the concepts in the text span multiple parts, some important information may be lost), context loss (such as the segmented text fragments may lose the original context, resulting in difficulty in understanding) and strong language field dependence (such as traditional text segmentation methods may be effective for documents in a specific language or field, but may not be effective in other cases. For documents with complex structures and formats, segmentation may become more difficult). The defects make the text blocks all have the same important knowledge point and complete information, clear themes, and clear logic, greatly improving the data quality of the retrieval library.

[0090] S502: Based on the text block set and the general information included in each text block, write questions for each text block to obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block, and obtain one or more data pairs corresponding to each text block. Based on the one or more data pairs corresponding to each text block, generate a data pair set.

[0091] After obtaining the text block set obtained by segmenting the knowledge document based on the aforementioned step S501, the embodiment of the present application supports the use of a key pair generator to generate questions for each text block in the text block set, so that the text block and the question text generated based on the text block are composed into a data pair, so that one or more data pairs corresponding to each text block constitute a data pair set. In this way, the data pair set can include multiple data pairs, each data pair consists of a matching question text and a text block, the question text in each data pair is obtained by writing questions for the text block that matches it, and the text block in each data pair serves as the source of the answer to the matching question text.

[0092] It should be noted that the embodiment of the present application takes into account that the data format of the input data of the retrieval model is a data pair, so the step of generating a data pair set based on the text block set is proposed. In detail, considering that traditional retrieval models such as TF-IDF (termfrequency_inversedocumentfrequency) or BM25 (Best Matching 25, a relevance scoring algorithm) algorithm, etc., all represent the question and context into a sparse high-dimensional space vector by efficiently matching keywords. This method of searching based solely on word matching does not consider the relevance between the semantics of text blocks and has significant limitations. The DPR model / algorithm is represented by a dense spatial vector that can contain semantic information. By optimizing the maximum inner product of the feature vectors of the question text and the related text block (passage), the goal is to compare the similarity of the feature vectors between the question text and the corresponding text block in all data pairs (question-passage, or represented as kv key-value pairs (abbreviated as kv pairs), k is the abbreviation of key, which can be represented as the question text in the embodiment of this application, and v is the abbreviation of value, which can be represented as the text block that contains the answer to the matching question text in the embodiment of this application) in the data pair set (i.e., a batch of data pairs). This seemingly simple method has a high retrieval accuracy. For example, the top-20 article retrieval accuracy is 9%-19% higher than the article retrieval accuracy of Lucene-BM25. Based on this, the present application adopts the DPR model as a retrieval model to realize answer retrieval in interactive dialogue scenarios, specifically the retrieval of text blocks containing answers to questions. The core of applying the DPR algorithm as a retrieval model is the construction of kv key-value pairs of questions and retrieval answers (specifically, the text blocks containing retrieval answers).

[0093] Furthermore, the embodiment of the present application supports the use of a key pair generator to generate a matching question text for each text block in the text block set, and generates a data pair corresponding to the text block based on the text block and the question text that matches the text block, thereby obtaining a data pair set. Among them, considering that the GPT model has extremely strong text comprehension and dialogue interaction capabilities, the embodiment of the present application supports the use of Prompt guidance (for relevant content about Prompt, please refer to the aforementioned related description, which will not be repeated here), and uses the GPT model as a key pair generator to implement question writing for text blocks.

[0094] The GPT model is used to write questions for each text block based on the text block set and the general information included in each text block to obtain one or more question texts corresponding to each text block. The one or more question texts corresponding to each text block are respectively combined with the corresponding text block to obtain one or more data pairs corresponding to each text block. Based on the one or more data pairs corresponding to each text block, the process may include:

[0095] (1) The GPT model obtains context information corresponding to the interactive dialogue scenario, which includes the dialogue style, interaction mode, and interaction domain. For details about the context information in the interactive dialogue scenario, please refer to the relevant content of the embodiment of FIG4 above and will not be repeated here. Furthermore, based on the context information, the set of text blocks, and the general information included in each text block, the GPT model determines the question style and number of questions (or the number of questions) that are suitable for each text block.

[0096] Specifically, the GPT model mines key elements for each text block based on a collection of text blocks and the general information included in each text block, obtaining the key elements of each text block. Key elements of a text block are factors that are crucial to the text block, i.e., factors that characterize the text block. These factors may include, but are not limited to, keywords, contextual relationships, and logical relationships. Then, based on the context information and the key elements of each text block, the GPT model determines the appropriate question style and number of questions for each text block. Each text block may be suitable for one or more question styles, each belonging to one of multiple conversational styles, including at least one of the following: factual, explanatory, inferential, evaluative, and hypothetical. Furthermore, the weighted number of questions for each question style corresponding to a text block may vary depending on the interactive conversational scenario. A higher weighted number of questions indicates a higher number of questions, while a lower weighted number indicates a lower number of questions. For example, in an interactive conversational scenario involving insurance inquiries, the number of explanatory questions is higher than the number of inferential questions. In this way, it is possible to combine the scene information and the content of the text block to achieve an accurate match of the question style and the number of questions that are suitable for the text block, and then use the question style and the number of questions to achieve accurate question writing. Furthermore, by using the general information to mine the key elements of the text block, and then using the obtained key elements to determine the question style and the number of questions that are suitable for the text block, it is possible to eliminate the interference of redundant information when determining the question style and the number of questions that are suitable for the text block, and by using a smaller number of key elements for adaptation, it is possible to effectively improve the efficiency of determining the question style and the number of questions that are suitable for the text block. (2) According to the question style and the number of questions that are suitable for each text block, write questions for each text block to obtain one or more question texts corresponding to each text block. In the embodiment of the present application, the GPT model is mainly used to implement semantic extraction and other processing of the text block to generate one or more question texts for the text block, that is, the step of obtaining one or more question texts corresponding to each text block in the embodiment of the present application can be achieved by the GPT model. For example, the three question texts generated for the text block using the GPT model can be as shown in Table 1 below:

[0097] It should be understood that Table 1 introduces the product positioning as medical insurance as an example, and does not limit the products applicable to the embodiments of the present application and the problems generated. This is specially explained here.

[0098] In addition, referring to the relevant description of text segmentation in the embodiment shown in Figure 4 above, it can be seen that there may be some text blocks that have a question-answer structure, that is, such text blocks themselves can be used as a data pair. For the sake of convenience, the text block set includes a candidate text block, and the candidate text block refers to a text block whose format conforms to the question-answer structure, and the candidate text block includes a question part and an answer part. For example, for candidate text blocks of this type of question-answer structure, the embodiment of the present application supports the construction of new data pairs through question generalization. This method of constructing data pairs through question generalization helps to generate key-value pairs of various dialogue styles and type variations to a certain extent, improves the richness of key-value pairs, and effectively reduces the workload compared to using a model to generate question text. Specifically, when it is detected that a candidate text block is included in the text block set, the candidate text block can be added to the data pair set, and the question part included in the candidate text block can be generalized to generate a generalized question text, and the generalized question text and the answer part included in the candidate text block are used to form a new data pair corresponding to the candidate text block, and the new data pair is added to the data pair set.

[0099] It's important to note that this approach to question generalization applies not only to candidate text blocks with question-answer structures within a text block set, but also to the iteration phase of the DPR model. Specifically, it uses failed recall data pairs to guide a few shots of the question, generalizing the question for these few shots. These new data pairs, after generalization, are then added to the training set. This helps leverage failure cases to gradually improve the DPR model's vector representation capabilities, thereby increasing the model's retrieval accuracy.

[0100] (3) Combining the one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block, and generating a data pair set based on the one or more data pairs corresponding to each text block.

[0101] It is worth noting that in order to avoid the GPT model from generating hallucinations in the process of generating question texts based on text blocks, that is, the question texts generated based on text blocks are false, the embodiment of the present application mandates that the answers corresponding to the generated question texts must come from the corresponding text blocks, and can locate the position information of the answers to the question texts from the corresponding text blocks, further improving the pertinence and availability of the questions. Therefore, after the key pair generator generates a sufficient amount of data pairs through the above description, the embodiment of the present application also supports the use of a key pair verifier to perform key pair verification on each data pair. Only data pairs that meet the verification requirements can be added to the data pair set to ensure that the answers to the question texts included in the data pairs in the data pair set can be extracted from the corresponding text, that is, to determine whether the data pairs in the data pair set are training data that can be used for subsequent training and verification. In the embodiment of the present application, it supports the use of the GPT model as a key pair verifier to implement key pair verification of the data pairs corresponding to each text block, and to perform secondary verification on the data pairs by utilizing the powerful text understanding ability of the GPT model to ensure the availability of the data pairs and the rationality of the question generation.

[0102] For example, assuming that any text block in the text block set is represented as a first text block, any data pair corresponding to the first text block is represented as a first data pair, and the first data pair includes a first question text that matches the first text block, then the data pair is double-checked using the GPT model, and the specific process of generating a data pair set can be seen in Figure 8. As shown in Figure 8, first, based on the first question text, a key pair verification model (i.e., the key pair verifier mentioned above, in the embodiment of the present application, the key pair verifier can be a GPT model) is used to extract the first answer that matches the first question text from the first text block, and mark the position information of the first answer in the first text block. Among them, the position information of the first answer in the first text block can be generated based on the position of the starting character and the ending character of the first answer in the first text block. For example, when the characters included in the first text block are marked in order from left to right, the position information of the starting character "B" of the first answer "B disease" shown in Figure 8 in the first text block is 23, and the position information of the ending character "disease" in the first text block is 25.

[0103] Then, based on the position information of the first answer in the first text block, the second answer indicated by the position information is extracted from the first text block. For example, the second answer extracted from the first text block based on the position information 23 and the position information 25 is "B disease". Secondly, the first answer is compared with the second answer extracted based on the position information to obtain a comparison result. The comparison between the first answer and the second answer here may include, but is not limited to: determining whether the character string composed of the characters included in the first answer and the character string composed of the characters included in the second answer are exactly the same, and keyword extraction technology (such as Ner technology) may be used to extract keywords from the first answer and the second answer respectively, and compare whether the keywords of the two answers match (i.e., are the same).

[0104] Finally, the first data pair is added to the data pair set based on the comparison result. If the comparison result indicates that the first answer and the second answer are the same, indicating that the first answer to the first question text in the first data pair can be extracted from the first text block in the first data pair, the first data pair is added to the data pair set. Conversely, if the comparison result indicates that the first answer and the second answer are different, indicating that the first answer to the first question text in the first data pair cannot be extracted from the first text block in the first data pair, the first data pair is not added to the data pair set, i.e., the first data pair is an illusion created by the GPT model, and the first data pair is discarded.

[0105] It can be seen that the embodiment of the present application determines the logic and rules for setting question generation for Prompt, and uses the GPT model for the first time to write questions for text blocks without question text. Compared with the traditional method of extracting question and answer records from historical data (such as common customer service question and answer manuals, using answers as retrieval libraries and user questions as corresponding questions to form key-value pairs) or manually writing questions based on text blocks, it greatly improves the productivity of questions and saves time and resources for question generation. In addition, in order to avoid the problem of hallucination in the process of generating question text for text blocks by the GPT model as much as possible, the embodiment of the present application also constructs a key pair verifier to perform a second verification on the generated data pairs. Only data pairs that have been successfully verified (that is, the text blocks in the data pairs are indeed the source of the answers to the corresponding question texts) can be added to the data pair set, which can effectively ensure the availability and data quality of the key pairs, thereby improving the model performance of the second retrieval model trained using the data pair set and the data quality of the retrieval library.

[0106] To facilitate understanding of the specific process shown in the above-mentioned steps S501-S502, the complete process of text segmentation processing, key-value pair generation and key-value pair verification is given again below in conjunction with Figure 9. As shown in Figure 9, first, after the text segmenter receives the knowledge document, it can perform text segmentation processing on the knowledge document to obtain one or more text blocks, which constitute a text block set. Then, a key pair generator is used to generate questions for each text block in the text set to obtain one or more data pairs corresponding to each text block, and if the text block itself is a question-answer structure, the question part of the text block is directly generalized to generate data pairs. Finally, a key pair verifier is used to perform a secondary verification on the data pairs corresponding to each text block, and the data pairs that are not hallucinated by the key pair generator are retained to form a data pair set. The text segmenter, key pair generator, and key pair verifier in the above process can all be GPT models. However, when the intelligent question-answering system executes the above processes, it will fine-tune the GPT data and adjust the prompt for different processes and functions. By using the GPT model to complete the functions of each link, the computing and storage costs are greatly reduced, while the model accuracy is significantly improved, broadening the capabilities of the AI ​​system (i.e., the intelligent question-answering system).

[0107] S503: Use the data pair set to train the first retrieval model to obtain a second retrieval model.

[0108] It should be understood that the training data for retrieval model training often includes a training set for retrieval model training and a validation set for retrieval model testing / validation. Therefore, after constructing the data pair set based on the aforementioned steps, it is necessary to use a text distributor to perform data distribution processing (i.e., text distribution processing) on ​​the data pair set so as to extract some data pairs from the data pair set as the validation set of the first retrieval model, and the remaining data pairs in the data pair set except for the extracted data pairs are used as the training set of the first retrieval model, thereby implementing the training of the first retrieval model based on the training set and the validation set to obtain a trained first retrieval model (i.e., the second retrieval model). By first performing data distribution processing to obtain a training set and a validation set, and then training the first retrieval model using a training-first-then-testing approach, it is possible to identify and correct model errors during the training process, improve the accuracy of the model, and use an independent test set in the testing phase to evaluate the generalization ability of the model, ensure that the model does not depend on the training data, and detect and reduce overfitting.

[0109] In the specific implementation: (1) the text allocator can perform data allocation processing on the data pair set according to the data allocation strategy to obtain a first data pair set and a second data pair set. When the first data pair set is a verification set, the second data pair set is a training set complementary to the first data pair set. Conversely, when the first data pair set is a training set, the second data pair set is a verification set complementary to the first data pair set. The embodiment of the present application does not limit the types of the first data pair set and the second data pair set.

[0110] Among them, the data allocation strategy described above may include: answer stratification strategy and word embedding classification strategy. Among them: ① The answer stratification strategy mainly stratifies the text block based on the position of the answer in the text block, and extracts a certain proportion of data pairs from each layer to form a data pair set. Specifically, as can be seen from the above description, the answers corresponding to the question texts in multiple data pairs generated by the same text block may be at different positions in the text block. Then, the text block can be stratified according to the position information of different answers in the text block (consisting of the starting position and the ending position), and then a certain proportion of data pairs corresponding to the question texts are randomly extracted from each layer as the first data pair set, and the unextracted data pairs in each layer are used as the second data pair set. In this way, it can be ensured that the answer knowledge points involved in the question texts in the data pair set (such as the first data pair set or the second data pair set) are similar to the overall question set, thereby improving the comprehensiveness of the data set pairs. ② The word embedding classification strategy mainly clusters the word vectors of the question texts in the data pair to extract a certain proportion of data pairs to form the data pair set. Specifically, word embeddings can be used based on the representation of text (or characters, strings). Word embeddings are used to represent words or phrases as fixed-size vectors that can capture features such as similarities between words and contextual relationships. First, word embedding methods (such as Word2Vec, GloVe, FastText, BERT, etc.) are used to convert each character in the string that makes up the question text into a vector. These character vectors are then clustered (such as K-means or hierarchical clustering). The clustering results are used as a basis for partitioning, and strings with similar representations are grouped together. Extraction from these groups is then possible, preserving the representations of different questions to the greatest extent possible and increasing the diversity and balance of the data set.

[0111] Below, taking any text block in the text block set as the first text block, the first text block corresponds to Q first data pairs, Q is an integer greater than zero, and the question text in the first data pair consists of one or more characters as an example, the specific implementation of the above two data allocation strategies is introduced in detail.

[0112] In one implementation, the data allocation strategy is an answer stratification strategy. In this implementation, the text allocator can determine the position information within the first text block of the answer corresponding to the question text in each of the Q first data pairs corresponding to the first text block. Then, based on the position information within the first text block of the answer corresponding to the question text in each first data pair, the first text block is stratified to obtain multiple text sub-layers corresponding to the first text block; each text sub-layer corresponds to at least one first data pair in the Q first data pairs. Finally, from the first data pairs corresponding to each of the multiple text sub-layers, a reference data pair is selected and added to the first data pair set, and the first data pairs in the multiple text sub-layers, excluding the selected reference data pairs, are added to the second data pair set. Thus, when constructing a target set (such as the first data pair set or the second data pair set) using the answer stratification strategy, it is possible to ensure that the first data pairs in the target set come from different layers within the first text block, thereby ensuring that the answer knowledge points involved in the question text in the target set are located at each layer of the first text block, thereby increasing the similarity between the answer knowledge points involved in the question text in the target set and the overall question set.

[0113] As shown in FIG10 , assuming that the first text block includes a total of 10 characters, and the first text block corresponds to 4 first data pairs (i.e., Q=4), namely, first data pair 1, first data pair 2, first data pair 3, and first data pair 4. Then, the position information of the answer corresponding to the question text in each of the 4 first data pairs in the first text block can be expressed as follows: the position information of the answer to the question text in first data pair 1 in the first text block is the starting position 1 (1 is the order of the first character of the answer in the first text block) and the ending position 4 (4 is the order of the last character of the answer in the first text block); the position information of the answer to the question text in first data pair 1 in the first text block is the starting position 1 and the ending position 4, the position information of the answer to the question text in first data pair 3 in the first text block is the starting position 5 and the ending position 9, and the position information of the answer to the question text in first data pair 4 in the first text block is the starting position 5 and the ending position 9. Furthermore, based on the positional information of the answers corresponding to the question texts in the four first data pairs within the first text block, the first text block is layered, resulting in two text sub-layers. Text sub-layer 1 has positional information of start position 1 and end position 4, corresponding to first data pair 1 and first data pair 2. Similarly, text sub-layer 2 has positional information of start position 5 and end position 9, corresponding to first data pair 3 and first data pair 4. Furthermore, a certain proportion (e.g., 50%) of first data pairs can be extracted from first data pair 1 and first data pair 2 corresponding to text sub-layer 1 and added to the first data pair set, and the unextracted first data pairs from text sub-layer 1 can be added to the second data pair set. Similarly, a certain proportion (e.g., 100%) of first data pairs can be extracted from first data pair 3 and first data pair 4 corresponding to text sub-layer 2 and added to the first data pair set, and the unextracted first data pairs from text sub-layer 2 can be added to the second data pair set.

[0114] In other implementations, the data allocation strategy is a word embedding classification strategy. In this implementation, the text allocator can perform word vector representation on the Q question texts in the Q first data pairs corresponding to the first text block, and obtain the word vector corresponding to each question text in the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts. Specifically, the closer the vector distance between at least two word vectors, the higher the similarity between the at least two question texts corresponding to the at least two word vectors, indicating that the questions that the at least two question texts want to ask may be closer or the same. Then, the word vectors corresponding to the Q question texts are clustered to obtain one or more cluster groups, and a cluster group includes one or more first data pairs corresponding to question texts whose vector distance meets the distance requirement. Here, the vector distance meeting the distance requirement can mean that the vector distance is less than the distance threshold. In other words, it supports dividing the first data pairs corresponding to the question texts corresponding to the word vectors whose vector distance is less than the distance threshold into one cluster group, so that the question texts included in the first data pairs in each cluster group are similar. Finally, reference data pairs are selected from one or more of the cluster groups and added to the first data pair set, and first data pairs other than the selected reference data in one or more cluster groups are added to the second data pair set. Thus, when the same cluster group includes multiple similar first data pairs, extracting first data pairs from different cluster groups to form a target set can largely preserve the representations of different problems and increase the diversity and balance of the target set. Specifically, the vector distance can be Euclidean distance, Manhattan distance, Hamming distance, etc., and the distance threshold can be configured according to the actual application scenario, and can specifically be a preset numerical threshold.

[0115] As shown in Figure 11, it is assumed that the first text block includes a total of 10 characters, and the first text block corresponds to 4 first data pairs (i.e., Q=4), namely, first data pair 1, first data pair 2, first data pair 3, and first data pair 4. Then, after the question texts in the 4 first data pairs are represented by word vectors and the word vectors corresponding to the 4 question texts are obtained, the word vectors corresponding to the 4 question texts can be clustered to obtain one or more cluster groups. Assuming that the vector distance between the word vector of the question text in the first data pair 1 and the word vector of the question text in the first data pair 3 is less than the distance threshold, it is determined that the first data pair 1 and the first data pair 3 are divided into the same cluster group (such as cluster group 1). Similarly, assuming that the vector distance between the word vector of the question text in the first data pair 2 and the word vector of the question text in the first data pair 4 is less than the distance threshold, it is determined that the first data pair 2 and the first data pair 4 are divided into the same cluster group (such as cluster group 2). In this way, a certain proportion of first data pairs can be extracted from cluster group 1 and cluster group 2 respectively and added to the first data pair set, and the remaining first data pairs that are not extracted can be added to the second data pair set.

[0116] It should be noted that Figures 10 and 11 above use a single text block as an example to illustrate the process of assigning text to one or more data pairs corresponding to that single text block in a data pair set. However, in actual applications, the method of assigning text to the data pairs corresponding to each text block in a data pair set is the same as the process shown in Figures 10 and 11 above, and will not be further described here.

[0117] (2) The first retrieval model is trained using the training set to obtain a trained first retrieval model. The training process of the first retrieval model using the training set may include multiple rounds of iterative training, and the first retrieval model after the last round of iterative training is used as a second retrieval model with better model prediction performance. The second retrieval model is used to implement question retrieval in subsequent model applications.

[0118] When the first retrieval model can be a dual-tower model, such as a DPR model, each round of iterative training for the DPR model may include: using the data pairs included in the training set as input data for the DPR model. Based on the dual-tower model structure of the DPR model (i.e., comprising sub-model 1 for vector representation of the question and sub-model 2 for vector representation of the text block), sub-model 1 of the DPR model can be used to perform vector embedding processing on the question text in the data pair to obtain a feature vector for the question text. Simultaneously, sub-model 2 of the DPR model can also be used to perform vector embedding processing on the text block in the data pair to obtain a feature vector for the text block. The retrieval model is then trained to ensure that the feature difference between the matching question text and text block in each data pair is less than a preset threshold. Specifically, the model parameters of the DPR model are optimized in a direction that reduces the difference between the feature vector of the question text and the feature vector of the text block in the same data pair. This optimization process is repeated until the DPR model has good vector representation capabilities for both the question text and the text block, specifically, the vector representations of the question text and the text block in the same data pair are similar.

[0119] (3) The trained first retrieval model is tested using a validation set to obtain a second retrieval model. After the second retrieval model is obtained through the above-mentioned retrieval model training process, the second retrieval model corresponds to a retrieval library, which includes the feature vectors obtained by vector embedding processing of each text block in the text block set by the second retrieval model. In this way, the second retrieval model and the retrieval library can be used together to generate a first answer for the first question text input by the object in an interactive dialogue scenario. The construction method of the retrieval library corresponding to the second retrieval model may include: optionally, the feature vectors of the text blocks included in the retrieval library are obtained during the retrieval model training process. Specifically, considering that in the last iterative training for the first retrieval model, the first retrieval model has a better vector representation for the text blocks included in the data pairs in the training set. Therefore, the embodiment of the present application supports adding the feature vectors of the text blocks included in each data pair in the training set during the last round of iterative training to the retrieval library to construct the retrieval library. It is worth noting that, since the same text block often corresponds to multiple data pairs, before adding the feature vectors of the text blocks included in each data pair in the training set to the retrieval library, it is also necessary to perform deduplication processing on the feature vectors of the text blocks included in each data pair. Specifically, if the same text block is included in different data pairs, then the feature vector of the text block included in one data pair can be selected and added to the retrieval library. Optionally, the feature vectors of the text blocks in the retrieval library can also be obtained by performing vector embedding processing on the text blocks in the text block set using the second retrieval model; that is, considering that each text block in the text block set is independent and unique, after obtaining the second retrieval model, the second retrieval model can be directly used to perform vector embedding processing on the text blocks in the text block set to obtain the feature vectors without performing deduplication operations.

[0120] In summary, the embodiments of the present application, on the one hand, ensure that the subject matter of the text block is single and contains general information during text segmentation, thereby ensuring good quality of the text block. On the other hand, matching questions are directly generated based on the text block to construct data pairs, effectively improving the authenticity and data quality of the data pairs. On the other hand, when training a retrieval model based on high-quality data pairs, the retrieval model is trained in a direction that ensures that the feature difference of the same data pair is reduced, so that the retrieval model has good feature representation capabilities for both question text and text blocks, thereby obtaining a second retrieval model with better vector expression capabilities and a retrieval library with better data quality.

[0121] The embodiments shown in Figures 4 and 5 above mainly introduce the construction of data pairs, retrieval model training and the construction of the retrieval library. The specific content of the model application will be introduced below in conjunction with the embodiments of Figures 12 and 13. As shown in Figure 12, during the model application process, after the intelligent question-answering system receives the first question text of the object, it can use the second retrieval model to perform vector embedding processing on the first question text to obtain the feature vector of the first question text, and perform similarity calculation on the feature vector of the first question text and the feature vector of the text block included in the retrieval library to determine one or more text blocks that match the first question text. It is worth noting that in the retrieval filtering stage, in addition to using the similarity calculation method described above, other matching methods can also be used, such as extracting key information from the first question text and performing keyword matching on the retrieval library, so as to determine the uniqueness of the keyword and the necessity of recalling the text block through the number of matched keywords.

[0122] Then, the intelligent question-answering system will classify the first question text into tasks to determine the difficulty level of the first question text. It is worth noting that in addition to classifying tasks according to the difficulty of the question, the first question text can also be classified according to other dimensions, such as according to the urgency of the task, the text theme of the first question text, etc. Finally, one or more text blocks that match the first question text are sent to the corresponding answer model according to the difficulty level of the first question text, so as to generate a corresponding first answer for the first question text based on the one or more text blocks. For example, the answer generation model corresponding to ordinary customer service, the answer generation model corresponding to professional customer service, and the answer generation model corresponding to question retrieval, these answer generation models can correspond to different GPT models, and ordinary customer service and professional customer service are mainly used for simple judgments that require judgment or known core instructions based on user information, while specific questions (such as the coverage details of insurance products) or difficult questions can be answered using question retrieval. It can be seen that, considering the performance and memory reasons of a single GPT model, it is impossible to carry an extremely complex intelligent question-answering system. Therefore, the embodiment of the present application supports the use of multiple GPTs for different role-playing and function implementation (for example, different GPT models for handling questions of different difficulty levels (such as ordinary customer service, professional customer service, etc.), and GPT models for text segmentation, and GPT models for key pair generation and key pair verification, etc.), so as to jointly complete the overall intelligent AI question-answering system.

[0123] In addition, as described above, the intelligent question-answering system provided by the embodiment of the present application also supports multi-round question-answering. Specifically, it combines the historical object data of the same subject during multiple rounds of interactive dialogue to achieve customized responses for the subject. In the process of model application, after the intelligent question-answering system obtains the correct first question text, it can rewrite the first question text based on the historical object data of the subject, so that the rewritten question text is more closely matched with the subject's object data, thereby better providing a personalized and customized first answer for the subject.

[0124] Based on the general introduction to model application in Figure 12 above, the complete implementation process of retrieval model training and model application is described in detail below in conjunction with Figure 13; Figure 13 shows a flow chart of a model method provided by an exemplary embodiment of the present application; the retrieval model training method flow shown in Figure 13 is mainly about the model application process of the second retrieval model. The retrieval model training method can be executed by a computer device, specifically by an intelligent question-answering system installed in the computer device. The retrieval model training method may include but is not limited to steps S1301-S1308:

[0125] S1301: Acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set.

[0126] S1302: Based on the text block set and the general information included in each text block, write questions for each text block to obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block, and obtain one or more data pairs corresponding to each text block. Based on the one or more data pairs corresponding to each text block, generate a data pair set.

[0127] S1303: Use the data pair set to train the first retrieval model to obtain a second retrieval model.

[0128] It should be noted that, for the specific implementation process shown in steps S1301-S1303, reference can be made to the relevant description of the specific implementation process shown in steps S501-S503 in the embodiment shown in FIG5 , and no further details are given here.

[0129] It should also be noted that in traditional retrieval schemes, all data are processed in the same way and have the same application status, but in fact, the information reliability and application breadth of data from different data sources are different, and in order to ensure the granularity of information, long texts belonging to the same topic will be segmented, so that the multiple text blocks after segmentation come from the same paragraph of text, that is, the topics of multiple text blocks may be the same, but their meanings and relevance may be different. In order to make full use of these differences between different text blocks in the retrieval library, the embodiment of the present application performs text segmentation processing on the knowledge document to obtain a text block set, and then sets a new index design for each text block in the text block set, so that the search for a specified text block (such as by source or subject, etc.) can be realized based on metadata, and it is helpful to realize the management of the retrieval library according to metadata (such as updating or adding and deleting the feature vectors of text blocks in the retrieval library, etc.).

[0130] In a specific implementation, after obtaining a set of text blocks, metadata can be extracted for each text block in the set of text blocks to obtain metadata for each text block; the metadata of any text block can be used to describe the data attributes of any text block, and the data attributes can at least include: the search domain to which the text block belongs, the source of the text block, and the subject to which the text block belongs. That is, the index design mechanism set in the embodiment of the present application introduces three concepts for text blocks, namely: search domain (or simply "domain"), source, and subject. Among them: ① "Domain" is the entire scope of the search library during the search. For example, when the search library includes feature vectors of text blocks of multiple insurance products, each insurance product can be a "domain". It is feasible to use data from multiple domains to train the retrieval model of the first retrieval model during the retrieval model training process. ② "Source" can refer to the source information or source location of the text block. For example, the source of the text block can be a summary of questions obtained during the customer service Q&A process. For another example, the source of the text block can be information obtained from insurance terms, etc. ③ "Subject" is the main knowledge content of the text block Chunk, and the relevance of the text block and the question can be confirmed through the subject.

[0131] In an embodiment of the present application, the metadata of each text block is used to update the search library and perform index retrieval on the text block. Specifically, the following steps are implemented: ① Updating the search library may include: when the data attributes of a text block in the search library are updated, modifying the attribute values ​​of the corresponding data attributes in the metadata of the text block to update the text block. For example, when the search library is continuously updated, the new version number or version time (i.e., version update time) of the text block can be added to the data attribute "source" of the text block, making the search library manageable, addable, subtractable, and replaceable. ② Index retrieval of the text block may include at least the following: 1. In an interactive dialogue scenario, based on the retrieval requirements indicated by the first question text, retrieve the text blocks whose metadata meet the retrieval requirements from the metadata to generate an answer for the first question text. For example, during the model application process, if the object specifies a query for a certain insurance product, the intelligent question-answering system can only search from the text blocks within that insurance product, rather than from the text blocks within other insurance products. In this case, it is necessary to filter the text blocks that meet the specified insurance product based on the data attribute "domain" as the text blocks that match the object's first question text. 2. When any text block has been retrieved in the interactive dialogue scenario, a text block with metadata identical to that of any text block is retrieved from the metadata to generate an answer for the first question text. For example, based on the aforementioned text segmentation process, for a long text belonging to the same topic, the multiple text blocks after the long text is segmented all use the same topic. If, during the model application process, a text block belonging to the same topic is recalled, other text blocks belonging to the same topic can be found based on the data attribute "topic" as auxiliary information for retrieval (e.g., as auxiliary information for generating answers to the retrieved text blocks).

[0132] It is worth noting that in the actual use of metadata to implement index retrieval and retrieval library updates, the data attributes "domain" and "source" are strongly related to the retrieval business and are explicit information (i.e., directly marked in the text block or added to the text block), so they can be used directly. If the data attribute "topic" exists in the text block (such as the title of the text block), it can also be used directly. However, if the data attribute "topic" does not exist in the text block, that is, it is implicit information, then the aforementioned GPT model can be used to perform real-time topic induction.

[0133] In summary, it can be seen that the embodiment of the present application constructs the above-described indexing mechanism for the text blocks in the retrieval library, so that the metadata structure can be used to realize the index retrieval of the text blocks and the update of the retrieval library, thereby making the retrieval library clearer and improving the retrieval speed and efficiency during the retrieval process, and facilitating the management of the retrieval library.

[0134] S1304: Receive a first question text input by the subject in the interactive dialogue scenario, and perform vector embedding processing on the first question text using a second retrieval model to obtain a feature vector of the first question text.

[0135] As described above, the second retrieval model involved in the embodiment of the present application can be a dual-tower model, such as a DPR model that includes a sub-model for vectorizing questions and a sub-model for vectorizing text blocks. Therefore, in an interactive dialogue scenario where the model is applied, after the intelligent question-answering system receives the first question text input by the object, it can use the sub-model for vectorizing questions in the second retrieval model to perform vector embedding processing on the first question text to obtain a feature vector of the first question text, which is used to represent the semantic information of the first question text.

[0136] It should be noted that the embodiments of the present application do not limit the interaction method between the intelligent question-answering system and the object, that is, they do not limit the method in which the intelligent question-answering system receives the first question text of the object. For example: as shown in the first figure of Figure 14, the object can directly input the first question text in text form on the display screen of the intelligent question-answering system. As shown in the second figure of Figure 14, the object can ask questions in voice form. At this time, the intelligent question-answering system can collect the voice signal of the object and convert the voice signal into the first question text in text form. As shown in the third figure of Figure 14, the intelligent question-answering system can also actively output one or more candidate questions in semantic or text form, so that the object selects a candidate question from one or more candidate questions displayed on the display screen of the intelligent question-answering system as the first question text to be consulted, and so on.

[0137] S1305: Perform similarity calculation on the feature vector of the first question text and the feature vectors of the P text blocks in the search library to obtain matching scores between the P text blocks and the first question text.

[0138] S1306: Sorting the P matching scores in descending order, and performing a gradient operation on each matching score in the sorted score sequence to obtain a plurality of gradient information.

[0139] S1307: Based on the plurality of gradient information, dynamically select one or more matching scores whose gradient information satisfies the gradient descent condition from the score sequence, and use the text blocks corresponding to the one or more matching scores as the answer source of the first question text.

[0140] In steps S1305-S1307, considering that the goal of training the second retrieval model in the embodiment of the present application is to make the feature difference between the question text and the text block belonging to the same data pair smaller, that is, the feature vector of the question text belonging to the same data pair and the feature vector of the text block are relatively similar, thereby ensuring that the question text belonging to the same data pair can definitely find the corresponding answer from the text block. Therefore, the intelligent question-answering system in the embodiment of the present application receives and calls the second retrieval model to perform vector embedding processing on the first question text, obtains the feature vector of the first question text, and can map the feature vector of the first question text to the retrieval library; specifically, the feature vector of the first question text and the feature vector of each text block in the P text blocks in the retrieval library are subjected to similarity calculation processing to obtain the matching scores between the P text blocks and the first question text respectively; the matching score corresponding to each text block is used to indicate the credibility of the corresponding text block as the source of the answer to the first question text, or to indicate the similarity between the feature vector of the corresponding text block and the feature vector of the first question text; in this way, the higher the matching score corresponding to the text block, the higher the possibility that the text block is the source of the answer to the first question text.

[0141] Considering that the common retrieval method in the industry is to extract a specified number of text blocks from the retrieval library and give them to the intelligent question-answering system for answering; but in fact, directly recalling a specified number of text blocks may result in too much invalid information in the recalled text blocks due to the specified number, resulting in problems such as overly redundant or even self-contradictory answers. In order to improve the flexibility of recalling the number of text blocks, the embodiment of the present application designs a method of dynamically recalling text blocks based on matching scores (or called a new split-score dynamic selection form), ensuring that all recalled text blocks are as concise and clear as possible around the first question text, thereby improving the effectiveness and accuracy of the first answer generated for the first question text, and further improving the retrieval accuracy.

[0142] The general principle of the method for dynamically adjusting the number of text chunks recalled based on the matching score score may include: for the first question text posed by the subject, after calculating the matching scores between the first question text and P text chunks in the search library using the aforementioned description, the P matching scores may be arranged from high to low (or from large to small) to obtain a score sequence, which includes the P matching scores in descending order. Then, a descending gradient is calculated for each of the score sequence. Specifically, the latter matching score is subtracted from the former matching score of two adjacent matching scores in the score sequence to obtain P-1 gradient information. These gradient information are then normalized and fitted to find one or more significant descending gradients outside of [mean - 3 * standard deviation, mean + 3 * standard deviation]. These significant descending gradients indicate a significant decrease in the similarity between the answer and the question. Then, the matching score corresponding to the first significant descending gradient in the one or more significant descending gradients is found in the score sequence (i.e., one or more matching scores that meet the gradient descending condition), and the text chunk corresponding to the matching score and the text chunk preceding the matching score in the score sequence are determined to be more likely to be the source of the answer to the first question text.

[0143] It should be noted that the above is based on the direct sorting of the P matching scores corresponding to the P text blocks in the retrieval rate; in actual applications, in order to save computing overhead, the embodiment of the present application supports specifying the output of the top-k text blocks (such as k = 10) and the matching scores corresponding to the top-k text blocks to achieve dynamic recall. For example, as shown in Figure 15, it is assumed that after the feature vector of the first question text and the feature vectors of the P text blocks in the retrieval library are processed for similarity calculation, the text blocks corresponding to the top-5 matching scores and the corresponding matching scores are screened, and the text blocks are sorted according to the matching scores as follows: text block 1 → text block 2 → text block 3 → text block 4 → text block 5, and the score sequence obtained by sorting the matching scores corresponding to each text block is: score 1 → score 2 → score 3 → score 4 → score 5. Then the gradient information is calculated as: gradient 1, gradient 2, gradient 3, and gradient 4, and normal fitting is performed on these gradient information to obtain [mean - 3 * standard deviation, mean + 3 * standard deviation]. If it is determined that the significantly decreasing gradients falling outside [mean - 3 * standard deviation, mean + 3 * standard deviation] are gradient 3 and gradient 4. Further, the first significantly decreasing gradient of the two significantly decreasing gradients is determined to be gradient 3, and the matching score corresponding to the first significantly decreasing gradient "gradient 3" is found in the score sequence as score 4. It is determined that text block 3, text block 2, and text block 1 before text block 4 corresponding to score 4 in the score sequence are more likely to be the source of the answer to the first question text, and the recalled text blocks are determined to be: text block 1, text block 2, and text block 3.

[0144] It can be seen that the text block recall method provided in the embodiment of the present application is not simply limited to recalling the first few text blocks from the retrieval library, but dynamically recalls text blocks according to the descending gradient. Moreover, when the answer to the first question text is relatively clear, the matching score between the first question text and the text block including the answer will be significantly increased, then one or more most relevant text blocks can be returned, further improving the recall accuracy and reducing the dilution of valid information by redundant answers.

[0145] S1308: Generate a first answer for the first question text based on the text blocks corresponding to the one or more matching scores.

[0146] As described in the embodiment shown in FIG12 above, the intelligent question-answering system provided by the embodiment of the present application is also configured with answer generation models corresponding to different difficulty levels. Then, after obtaining one or more matching scores that meet the gradient descent condition based on the aforementioned steps S1304-S1307 (specifically, the matching scores before the first significant gradient descent), the one or more matching scores can be assigned to the answer generation model corresponding to the difficulty level according to the difficulty level of the first question text to generate the first answer. In a specific implementation, the intelligent question-answering system performs task classification processing on the first question text to obtain a classification result. The classification result is used to indicate the difficulty of answering the first question text (i.e., the difficulty level mentioned above). Then, the intelligent question-answering system generates an initial answer to the first question text based on the text blocks corresponding to one or more matching scores according to the difficulty of the answer indicated by the classification result.

[0147] Furthermore, in order to ensure the accuracy of the initial answer generated by the answer generation model, the embodiment of the present application also supports answer quality inspection / review of the initial answer, and uses the reviewed initial answer as the first answer that matches the first question text. Specifically, the intelligent question and answer system supports answer review processing of the initial answer to obtain the first answer to the first question text; in the actual review process, if the initial answer is a suitable answer (such as a style that is consistent with the first question text), then the initial answer is the first answer. If the initial answer is not a suitable answer, then the answer after the initial answer is fine-tuned is determined to be the first answer. In detail, the intelligent question and answer system can include a quality inspection model with a quality inspection function, and the intelligent question and answer system can call the quality inspection model to perform quality inspection on the initial answer. The quality inspection here is mainly performed by inputting the first question text into the quality inspection model to facilitate the verification of the question style, and when the form or style of the initial answer is inappropriate, the initial answer is rewritten to obtain the first answer.

[0148] Furthermore, traditional Q&A is mainly targeted at single-round Q&A. When faced with multi-round Q&A, it is usually split into single rounds or forced intervention by manual strategies; this results in the loss of a lot of key information between multi-round Q&A, and the correlation between multi-round Q&A is not strong, resulting in a stiff response. To improve the flexibility and relevance of multi-round Q&A, the embodiments of this application support the use of the long memory and extraction capabilities of the GPT model, using different GPT models for role-playing, and combining the object data of the object to play different advantages in different links of the multi-round Q&A, thereby achieving customized answers for different objects.

[0149] Taking an interactive dialogue scenario including multiple rounds of interactive dialogues, where any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogues is represented as a target round interactive dialogue, as an example, the operations performed by the intelligent question-answering system in the target round interactive dialogue in the multiple rounds of interactive dialogues may include: first, obtaining a first question text input by the subject in the target round interactive dialogue, historical dialogue data (such as historical question texts and answers to historical question texts) from previous round interactive dialogues (such as interactive dialogues located between target round interactive dialogues in the multiple rounds of interactive dialogues), and historical object data about the subject (such as personalized data about the subject extracted during the previous round interactive dialogues, such as the subject's query focus and basic information about the subject (such as age and gender, etc.)), where the historical object data is generated based on the historical dialogue data. Then, based on the first question text, the historical dialogue data, and the historical object data, the target text question is rewritten to obtain a new first question text. In this way, the second retrieval model can be used to perform vector embedding processing on the new first question text, and the similarity between the feature vector of the new first question text and the feature vector of the text block in the retrieval library can be calculated to obtain multiple matching scores, and operations such as screening the text block that matches the first question text based on the multiple matching scores can be performed (i.e., the specific process shown in the above steps S1304-S1307). It is worth noting that the target round interactive dialogue in the above description is any round of interactive dialogue except the first round of interactive dialogue in the multi-round interactive dialogue. Then, for the first round of interactive dialogue in the multi-round interactive dialogue, it can be regarded as a single round of interactive dialogue, and there is no object data that can be introduced.

[0150] As described above, the embodiment of the present application uses multiple GPT models to implement different processes of multi-round question and answer. The following describes the general process of triggering multi-round interactive dialogues of the intelligent question and answer system from the perspective of different GPT models:

[0151] First, after receiving the first question text, the intelligent question-answering system can use GPT model A as a process controller to input the current question text, historical dialogue data (if the current interactive dialogue is the first round, the historical dialogue data is empty), and historical object data (if the current interactive dialogue is the first round, the historical object data is empty) raised by the subject during the current interactive dialogue into GPT model A. In this way, GPT model A performs the following operations: extracting new object data based on the current question text, historical dialogue data, and historical object data raised by the subject during the current interactive dialogue; determining whether the historical object data needs to be updated or added based on the current interactive dialogue; and rewriting the current question text based on the new object data and the current question text to generate a simple and understandable new question text.

[0152] Then, the intelligent question answering system calls the second retrieval model to perform vector embedding processing on the new question text, and executes the specific steps shown in the aforementioned steps S1304-S1307 to obtain one or more text blocks that match the new question text.

[0153] Finally, the intelligent question-answering system uses GPT model B as a logical question distribution to determine whether the rewritten question (i.e., the new question text) can be answered directly based on one or more text blocks matching the new question text. If it can, it indicates that the new question text is relatively easy to answer, such as questions about price or policy amount. In this case, GPT model C (as shown in Figure 12, a general customer service representative) is used as a general answerer to directly generate a first answer based on one or more text blocks matching the new question text. Conversely, if it cannot, it indicates that the new question text is somewhat difficult to answer, such as determining whether a certain industry is eligible for insurance. This determines whether more specialized logical reasoning is required. If it is not, it indicates that the new question text is somewhat difficult to answer, but the difficulty level is not high. GPT model D (as shown in Figure 12, a professional customer service representative) can be handed over to generate a first answer based on one or more text blocks matching the new question text. If it is required, it indicates that the new question text is difficult to answer, and GPT model E (as shown in Figure 12, a question retrieval model) is used in combination with the object data to perform logical reasoning and output the answer.

[0154] In the embodiment of the present application, from the user experience level, the intelligent question-answering system can support multiple rounds of interactive dialogues and user-customized recommendations, which greatly improves the fluency, naturalness and accuracy of the answers. From the underlying technical level, the embodiment of the present application includes at least the following advantages: 1. From the perspective of text segmentation, a new text segmentation technology is proposed. This text segmentation technology can ensure that the subject of the segmented text block is single and carries general information, so that the text block has advantages such as clear subject and clear logic, thereby improving the data quality of the retrieval library built based on the text block. 2. From the perspective of data pair generation, a new data pair generation scheme is proposed, specifically using the functions and advantages of the emerging technology GPT to build more real and reliable data pairs based on text blocks, that is, to ensure that the text blocks in the same data pair must be the source of the answer to the question text. 3. From the perspective of retrieval model training, when using high-quality data pairs to train the first retrieval model, it is also supported to train the model in a direction that reduces the feature differences between the question text and the text block in the same data pair. This ensures that the trained first retrieval model (i.e., the second retrieval model) has good vector representation capabilities for both the question text and the text block, and that the vector representation of the question text is closer to the vector representation of the text block that provides the answer to the question text. Thus, during model application, a good and accurate text block is determined to match the first question text. 4. From the perspective of text block recall, it supports a dynamic recall count based on matching scores, enabling dynamic recall of different numbers of text blocks for different questions and ensuring the matching of the recalled text blocks to the corresponding questions, thereby improving the effective information concentration of the recall. 5. From the perspective of the retrieval library, it not only supports the storage of metadata structures for text blocks in the retrieval library, thereby facilitating index retrieval and retrieval library updates (such as iterative updates) based on the metadata of the text blocks, but also supports multiple formats of retrieval, including tables, images, and videos. In summary, the embodiments of the present application support the use of multiple GPT models to play different roles for link control, function implementation, and response quality inspection, etc., and are committed to improving the retrieval capability, link control capability, and continuous iteration capability of the intelligent question-answering system from all aspects (such as combining multiple rounds of question-answering with object data), and generating sustainable iterative solutions.

[0155] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, accordingly, the device of the embodiment of the present application is provided below. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the module or unit function.

[0156] FIG16 shows a schematic diagram of the structure of a retrieval model training device provided by an exemplary embodiment of the present application; the retrieval model training device can be used to perform some or all of the steps in the method embodiments shown in FIG5 or FIG13. Referring to FIG16, the device includes the following units:

[0157] The acquisition unit 1601 is configured to acquire a knowledge document and perform text segmentation processing on the knowledge document to obtain a text block set; the text block set includes multiple text blocks, each of which is composed of a reference text block belonging to the knowledge document and its corresponding summary information, where the summary information is obtained by summarizing the text semantics of the corresponding reference text block;

[0158] Processing unit 1602 is configured to write a question for each text block based on the text block set and the general information included in each text block to obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block to obtain one or more data pairs corresponding to each text block, and generate a data pair set based on the one or more data pairs corresponding to each text block; each data pair consists of a matching question text and a text block; the text block in each data pair serves as a source of an answer to the question text that matches the text block; and

[0159] Processing unit 1602 is also used to train the first retrieval model using a data pair set to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each data pair less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes a feature vector of each text block in the text block set; wherein the second retrieval model and the retrieval library are used to generate a first answer for the first question text in an interactive dialogue scenario.

[0160] In one implementation, the number of knowledge documents is at least one; the processing unit 1602 is used to perform text segmentation processing on the knowledge document, and when obtaining a text block set, it is specifically used to: obtain semantic information of the knowledge document, and segment the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document, obtain general description information corresponding to each reference text block by semantically summarizing the text semantics of each reference text block in the one or more reference text blocks, add the general description information corresponding to each reference text block to the target position in the corresponding reference text block, and obtain a text block corresponding to each reference text block; the text blocks corresponding to each reference text block constitute a text block set.

[0161] In one implementation, the semantic information of the knowledge document includes the subject to which the knowledge document belongs, and the knowledge document is composed of one or more characters; the processing unit 1602 is used to segment the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document, specifically for: if the subject to which the knowledge document belongs is single, then count the number of characters included in the knowledge document; when the number of characters included in the knowledge document is less than the character number threshold, the knowledge document is used as a reference text block; when the number of characters included in the knowledge document is greater than or equal to the character number threshold, the knowledge document is segmented according to the text logical relationship to obtain multiple reference text blocks corresponding to the knowledge document.

[0162] In one implementation, the processing unit 1602 is further used to: if the subject to which the knowledge document belongs is at least two subjects, perform hierarchical expression segmentation on the knowledge document to obtain at least two initial text blocks; at least two initial text blocks each have a subject, count the number of characters included in each of the at least two initial text blocks, and use the initial text block with a character number less than a character number threshold as a reference text block, perform paragraph segmentation processing on the initial text block with a character number greater than or equal to the character number threshold according to the text logical relationship, and obtain multiple reference text blocks corresponding to the knowledge document.

[0163] In one implementation, the knowledge document or initial text block subjected to paragraph segmentation processing is represented as text content; the text logical relationship is a general-specific relationship or a parallel relationship; the general-specific relationship indicates that the content structure of the text content is a general-specific structure, and the parallel relationship indicates that the content structure of the text content is a parallel structure; the paragraph segmentation processing includes: according to the text logical relationship of the text content, paragraph segmentation processing is performed on the text content to obtain a reference text block corresponding to the knowledge document, wherein, when the text logical relationship of the text content is a general-specific relationship, the text content is segmented into a general description text block and one or more specific description text blocks; the general description text block is content in the text content that has a generalizing effect, and the general description information corresponding to the general description text block is used to indicate the overall semantics expressed by the text content; the specific description text block is content in the text content that has an explanatory effect on the general description text block, and the general description information corresponding to the specific description text block is used to indicate the specific description semantics expressed by the specific description text block and the overall semantics expressed by the text content; when the text logical relationship is a parallel relationship, the text content is segmented into at least two specific description text blocks.

[0164] In one implementation, the processing unit 1602 is further used to: extract metadata from each text block in the text block set to obtain metadata for each text block; the metadata of the text block is used to describe the data attributes of the text block, and the data attributes include: the retrieval domain to which the text block belongs, the source of the text block, and the subject to which the text block belongs, wherein the metadata of each text block is used to update the retrieval library and perform index retrieval on the text block; the update includes: when the data attributes of the text block in the retrieval library are updated, modifying the attribute values ​​of the corresponding data attributes in the metadata of the text block, and the index retrieval at least includes: in an interactive dialogue scenario, according to the retrieval requirements indicated by the first question text, retrieving from the metadata the text blocks whose metadata meet the retrieval requirements to generate an answer for the first question text; and, when any text block has been retrieved in the interactive dialogue scenario, retrieving from the metadata the text blocks whose metadata are the same as the metadata of any text block to generate an answer for the first question text.

[0165] In one implementation, the processing unit 1602 is used to generate a data pair set based on the text block set and the general information included in each text block, and is specifically used to: obtain scene information corresponding to the interactive dialogue scene, the scene information including the dialogue style, interaction method and interaction field; based on the scene information, the text block set and the general information included in each text block, determine the question style and number of questions that are suitable for each text block; according to the question style and number of questions that are suitable for each text block, write questions for each text block to obtain one or more question texts corresponding to each text block.

[0166] In one implementation, the processing unit 1602 is used to determine the question style and number of questions suitable for each text block based on the scene information, the text block set and the general information included in each text block. It is specifically used to: mine the key elements of each text block according to the text block set and the general information included in each text block to obtain the key elements of each text block; and determine the question style and number of questions suitable for each text block based on the scene information and the key elements of each text block.

[0167] In one implementation, any text block in the text block set is represented as a first text block, and any data pair corresponding to the first text block is represented as a first data pair, wherein the first data pair includes a first question text that matches the first text block; processing unit 1602 is used to generate a data pair set based on one or more data pairs corresponding to each text block, specifically for: extracting a first answer that matches the first question text and position information of the first answer in the first text block from the first text block based on the first question text using a key pair verification model; extracting a second answer indicated by the position information from the first text block according to the position information of the first answer in the first text block; if the first answer is the same as the second answer, the first data pair is added to the data pair set; if the first answer is different from the second answer, the first data pair is not added to the data pair set.

[0168] In one implementation, the text block set includes candidate text blocks, where the candidate text block refers to a text block whose format conforms to the question-answer structure, and the candidate text block includes a question part and an answer part; the processing unit 1602 is also used to: add the candidate text block to the data pair set, perform question generalization processing on the question part included in the candidate text block, generate a generalized question text, and use the generalized question text and the answer part included in the candidate text block to form a new data pair corresponding to the candidate text block, and add the new data pair to the data pair set.

[0169] In one implementation, the processing unit 1602 is used to train the first retrieval model using a data pair set to obtain a second retrieval model, specifically for: performing data allocation processing on the data pair set according to a data allocation strategy to obtain a first data pair set and a second data pair set; the data allocation strategy includes an answer stratification strategy and a word embedding classification strategy; the first data pair set is a validation set, and the second data pair set is a training set; or, the first data pair set is a training set, and the second data pair set is a validation set, the training set is used to train the first retrieval model to obtain a trained first retrieval model, and the validation set is used to test the trained first retrieval model to obtain a second retrieval model.

[0170] In one implementation, the data allocation strategy is an answer stratification strategy, and any text block in the text block set is represented as a first text block. The first text block corresponds to Q first data pairs, where Q is an integer greater than zero. The processing unit 1602 is used to perform data allocation processing on the data pair set according to the data allocation strategy to obtain the first data pair set and the second data pair set. The processing unit 1602 is specifically used to: determine the position information of the answer corresponding to the question text in each first data pair in the Q first data pairs in the first text block; perform stratification processing on the first text block according to the position information of the answer corresponding to the question text in each first data pair in the first text block to obtain multiple text sub-layers corresponding to the first text block, where each text sub-layer corresponds to at least one first data pair in the Q first data pairs; select a reference data pair from the first data pairs corresponding to each text sub-layer in the multiple text sub-layers and add it to the first data pair set; and add the first data pairs of the multiple text sub-layers except the selected reference data pairs to the second data pair set.

[0171] In one implementation, the data allocation strategy is a word embedding classification strategy, and any text block in the text block set is represented as a first text block. The first text block corresponds to Q first data pairs, where Q is an integer greater than zero, and the question text in the first data pair consists of one or more characters. The processing unit 1602 is used to perform data allocation processing on the data pair set according to the data allocation strategy, and when obtaining the first data pair set and the second data pair set, it is specifically used to: respectively represent the Q question texts in the Q first data pairs by word vectors to obtain the word vector corresponding to each question text in the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between different question texts, cluster the word vectors corresponding to the Q question texts to obtain one or more cluster groups, and a cluster group includes one or more first data pairs corresponding to question texts whose vector distance meets the distance requirement, select reference data pairs from the one or more cluster groups and add them to the first data pair set, and add the first data pairs in the one or more cluster groups except the selected reference data to the second data pair set.

[0172] In one implementation, the retrieval library includes feature vectors of P text blocks, where P is a positive integer and P is less than Q. The processing unit 1602 is further configured to: receive a first question text input by an object in an interactive dialogue scenario, perform vector embedding processing on the first question text using a second retrieval model, and obtain a feature vector of the first question text, wherein the feature vector of the first question text is used to represent semantic information of the first question text; perform similarity calculation on the feature vector of the first question text and the feature vectors of the P text blocks in the retrieval library to obtain matching scores between the P text blocks and the first question text; the matching scores are used to indicate the credibility of the text blocks as the source of the answer to the first question text; the P matching scores are sorted in descending order, and a gradient operation is performed on each matching score in the sorted score sequence to obtain multiple gradient information; based on the multiple gradient information, dynamically select one or more matching scores from the score sequence whose gradient information satisfies a gradient descent condition; the text blocks corresponding to the one or more matching scores are used as the source of the answer to the first question text; and a first answer is generated for the first question text based on the text blocks corresponding to the one or more matching scores.

[0173] In one implementation, the processing unit 1602 is used to generate a first answer for a first question text based on one or more text blocks corresponding to matching scores, and is specifically used to: perform task classification processing on the first question text to obtain a classification result, the classification result is used to indicate the difficulty of answering the first question text, according to the answer difficulty indicated by the classification result, use an answer generation model that matches the answer difficulty to generate an initial answer to the first question text based on the text blocks corresponding to one or more matching scores, perform answer review processing on the initial answer, and obtain a first answer to the first question text; the first answer is the initial answer, or an answer after the initial answer is fine-tuned.

[0174] In one implementation, the interactive dialogue scenario includes multiple rounds of interactive dialogues, and any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogues is represented as a target round interactive dialogue. The processing unit 1602 is further used to: obtain a first question text input by the object in the target round interactive dialogue, historical dialogue data and historical object data about the object in the historical round interactive dialogue; the historical object data is generated based on the historical dialogue data, and the target text question is rewritten based on the first question text, the historical dialogue data and the historical object data to obtain a new first question text. The processing unit 1602 is used to use the second retrieval model to perform vector embedding processing on the first question text, specifically to: use the second retrieval model to perform vector embedding processing on the new first question text.

[0175] According to one embodiment of the present application, the various units in the retrieval model training device shown in Figure 16 can be separately or all merged into one or several other units to constitute, or one (some) of the units can be further split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effect of the embodiment of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the function of multiple units is realized by one unit. In other embodiments of the present application, the retrieval model training device can also include other units. In practical applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, a computer program (including program code) that can execute the steps involved in the corresponding method shown in Figures 5 and 13 can be run on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), an access storage medium (RAM), a read-only storage medium (ROM), etc., to construct a retrieval model training device as shown in Figure 16, and to implement the retrieval model training method of the embodiment of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above-mentioned computing device via the computer-readable recording medium, and executed therein.

[0176] In an embodiment of the present application, after obtaining the knowledge document of the vertical category, first, the knowledge document can be subjected to text segmentation processing to obtain a text block set including multiple text blocks; wherein each text block has a theme to ensure that the theme of each document block is clear, and each text block includes not only the reference text block belonging to the knowledge document, but also general information for summarizing the text semantics of the reference text block, so that each text block has the advantage of clear logic. Then, support is provided for generating a data pair set including multiple data pairs based on the text block set, specifically the text blocks in the text block set; wherein each data pair is composed of a question text and a text block, and the question text in each data pair is obtained by writing questions for the text block that matches it, which greatly improves the productivity of questions, and the text block of each data pair serves as the source of the answer to the question text that matches it, ensuring the authenticity and availability of the data pair. Finally, the first retrieval model can be trained using a data pair set to obtain a second retrieval model and a retrieval library (including a feature vector obtained by vector embedding processing of each text block using the second retrieval model); during the training process, the feature difference between the matching question text and the text block in each data pair needs to be less than a preset threshold to ensure that the features of the question text and the text block in the same data pair are similar, which is conducive to mapping the features of the first question text to a retrieval library with better data quality during subsequent answer retrieval, and being able to find text blocks with similar features from the retrieval library to generate the first answer for the first question text, thereby improving the accuracy and professionalism of answer retrieval in the intelligent question and answer process.

[0177] Figure 17 shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Referring to Figure 17, the computer device includes a processor 1701, a communication interface 1702 and a computer-readable storage medium 1703. The processor 1701, the communication interface 1702 and the computer-readable storage medium 1703 can be connected via a bus or other means. The communication interface 1702 is used to receive and send data. The computer-readable storage medium 1703 can be stored in the memory of the computer device, and the computer-readable storage medium 1703 is used to store a computer program. The computer program includes program instructions, and the processor 1701 is used to execute the program instructions stored in the computer-readable storage medium 1703. The processor 1701 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0178] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, one or more instructions suitable for being loaded and executed by the processor 1701 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage, and optionally, at least one computer-readable storage medium located away from the aforementioned processor.

[0179] In one embodiment, the computer-readable storage medium stores one or more instructions, and the processor 1401 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned retrieval model training method embodiment.

[0180] Based on the same inventive concept, the principles and beneficial effects of solving the problem by the computer device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problem by the retrieval model training method in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.

[0181] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described retrieval model training method.

[0182] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0183] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD) or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0184] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0185] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A retrieval model training method, characterized in that: Executed by computer equipment, including: Acquire a knowledge document, and perform text segmentation processing on the knowledge document to obtain a text block set; the text block set includes multiple text blocks, each of which is composed of a reference text block belonging to the knowledge document and general description information corresponding to the reference text block, and the general description information is obtained by summarizing the text semantics of the corresponding reference text block; Based on the text block set and the general information included in each text block, write a question for each text block to obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block, obtain one or more data pairs corresponding to each text block, and generate a data pair set based on the one or more data pairs corresponding to each text block; each data pair consists of a matching question text and a text block; the text block in each data pair serves as a source of answers to the question text matching the text block; and The first retrieval model is trained using the data pair set to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each of the data pairs less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes a feature vector of each text block in the text block set; wherein the second retrieval model and the retrieval library are used to generate a first answer for the first question text in an interactive dialogue scenario.

2. The method according to claim 1, characterized in that The number of the knowledge document is at least one; the text segmentation process is performed on the knowledge document to obtain a text block set, including: Acquiring semantic information of the knowledge document, and segmenting the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document; Obtaining summary information corresponding to each reference text block by semantically summarizing the text semantics of each reference text block in the one or more reference text blocks; and The general description information corresponding to each reference text block is added to the target position in the corresponding reference text block to obtain the text block corresponding to each reference text block; the text blocks corresponding to each reference text block constitute a text block set.

3. The method according to claim 2, characterized in that The semantic information of the knowledge document includes the subject to which the knowledge document belongs, and the knowledge document is composed of one or more characters; the segmentation process of the knowledge document based on the semantic information of the knowledge document to obtain one or more reference text blocks corresponding to the knowledge document includes: If the subject to which the knowledge document belongs is single, counting the number of characters included in the knowledge document; and When the number of characters included in the knowledge document is less than a character number threshold, the knowledge document is used as a reference text block.

4. The method according to claim 3, characterized in that The method further comprises: When the number of characters included in the knowledge document is greater than or equal to a character number threshold, the knowledge document is segmented into paragraphs according to the text logic relationship to obtain a plurality of reference text blocks corresponding to the knowledge document.

5. The method according to claim 3, characterized in that The method further comprises: If the subject to which the knowledge document belongs is at least two subjects, the knowledge document is segmented by hierarchical expression to obtain at least two initial text blocks; the at least two initial text blocks each have a subject; Counting the number of characters included in each of the at least two initial text blocks, and taking the initial text block whose number of characters is less than a character number threshold as a reference text block; and The initial text blocks whose number of characters is greater than or equal to the character number threshold are segmented into paragraphs according to the text logical relationship to obtain a plurality of reference text blocks corresponding to the knowledge document.

6. The method according to claim 4 or 5, characterized in that The knowledge document or initial text block to be processed by paragraph segmentation is represented as text content; the text logical relationship is a general-specific relationship or a parallel relationship; the general-specific relationship indicates that the content structure of the text content is a general-specific structure, and the parallel relationship indicates that the content structure of the text content is a parallel structure; the paragraph segmentation process includes: According to the textual logical relationship of the textual content, the textual content is segmented into paragraphs to obtain a reference text block corresponding to the knowledge document; Wherein, when the textual logical relationship of the textual content is the general-specific relationship, the textual content is cut into a general description text block and one or more specific description text blocks; the general description text block is the content in the textual content having a summarizing effect, and the general description information corresponding to the general description text block is used to indicate the overall semantics expressed by the textual content; the specific description text block is the content in the textual content having an explanatory effect on the general description text block, and the general description information corresponding to the specific description text block is used to indicate the specific semantics expressed by the specific description text block and the overall semantics expressed by the textual content; and When the text logical relationship is a parallel relationship, the text content is cut into at least two sub-text blocks.

7. The method according to any one of claims 1 to 6, characterized in that: After the text segmentation process is performed on the knowledge document to obtain a text block set, the method further includes: Extracting metadata from each text block in the text block set to obtain metadata for each text block; the metadata for the text block is used to describe data attributes of the text block, and the data attributes include: the search domain to which the text block belongs, the source of the text block, and the subject to which the text block belongs; The metadata of each text block is used to update the search library and perform index search on the text block; the update includes: when the data attribute of the text block in the search library is updated, modifying the attribute value of the corresponding data attribute in the metadata of the text block; and The index retrieval includes: in an interactive dialogue scenario, according to the retrieval requirements indicated by the first question text, retrieving text blocks whose metadata meet the retrieval requirements from metadata to generate an answer for the first question text.

8. The method according to claim 7, characterized in that The index search also includes: When any text block has been retrieved in the interactive dialogue scene, a text block having metadata identical to that of the any text block is retrieved from the metadata to generate an answer for the first question text.

9. The method according to any one of claims 1 to 8, characterized in that: The step of writing a question for each text block based on the text block set and the general information included in each text block to obtain one or more question texts corresponding to each text block includes: Acquire scenario information corresponding to the interactive dialogue scenario, wherein the scenario information includes a dialogue style, an interaction mode, and an interaction field; Determine the question style and number of questions suitable for each text block based on the scenario information, the text block set and the general information included in each text block; and Questions are written for each text block according to the question style and question quantity that are suitable for each text block, so as to obtain one or more question texts corresponding to each text block.

10. The method according to claim 9, characterized in that The step of determining the question style and the number of questions suitable for each text block based on the scenario information, the text block set and the general information included in each text block includes: According to the text block set and the general information included in each text block, key element mining is performed on each text block to obtain the key element of each text block; and Based on the scenario information and key elements of each text block, the question style and number of questions suitable for each text block are determined.

11. The method according to any one of claims 1 to 10, characterized in that: Any text block in the text block set is represented as a first text block, and any data pair corresponding to the first text block is represented as a first data pair, wherein the first data pair includes a first question text matching the first text block; The step of generating a data pair set based on one or more data pairs corresponding to each of the text blocks comprises: Extracting a first answer matching the first question text and position information of the first answer in the first text block from the first text block based on the first question text by using a key pair verification model; According to the position information of the first answer in the first text block, the position the second answer indicated by the setting information; and If the first answer is the same as the second answer, the first data pair is added to the data pair set.

12. The method according to any one of claims 1 to 11, characterized in that: The text block set includes candidate text blocks, wherein the candidate text blocks refer to text blocks whose format conforms to the question-answer structure, and the candidate text blocks include a question part and an answer part; the method further includes: Adding the candidate text block to the data pair set; and The question part included in the candidate text block is generalized to generate a generalized question text, and the generalized question text and the answer part included in the candidate text block are used to form a new data pair corresponding to the candidate text block, and the new data pair is added to the data pair set.

13. The method according to any one of claims 1 to 12, characterized in that: The step of training the first retrieval model using the data pair set to obtain the second retrieval model comprises: Performing data allocation processing on the data pair set according to the data allocation strategy to obtain a first data pair set and a second data pair set; the first data pair set is a validation set, and the second data pair set is a training set; or, the first data pair set is a training set, and the second data pair set is a validation set; Using the training set to train the first retrieval model to obtain a trained first retrieval model; and The validation set is used to test the trained first retrieval model to obtain a second retrieval model.

14. The method according to claim 13, characterized in that The data allocation strategy is an answer stratification strategy; any text block in the text block set is represented as a first text block, and the first text block corresponds to Q first data pairs, where Q is an integer greater than zero; the data pair set is subjected to data allocation processing according to the data allocation strategy to obtain a first data pair set and a second data pair set, including: Determine the position information of the answer corresponding to the question text in each of the Q first data pairs, in the first text block; According to the position information of the answer corresponding to the question text in each of the first data pairs in the first text block, the first text block is layered to obtain a plurality of text sub-layers corresponding to the first text block; one text sub-layer corresponds to at least one first data pair in the Q first data pairs; and A reference data pair is selected from the first data pairs corresponding to each text sub-layer in the plurality of text sub-layers and added to the first data pair set; and the first data pairs in the plurality of text sub-layers except the selected reference data pair are added to the second data pair set.

15. The method according to claim 13, characterized in that The data allocation strategy is a word embedding classification strategy; any text block in the text block set is represented as a first text block, the first text block corresponds to Q first data pairs, Q is an integer greater than zero, and the question text in the first data pair consists of one or more characters; the data pair set is subjected to data allocation processing according to the data allocation strategy to obtain a first data pair set and a second data pair set, including: Representing the Q question texts in the Q first data pairs with word vectors respectively, and obtaining a word vector corresponding to each question text in the Q question texts; the vector distance between the word vectors corresponding to different question texts is used to indicate the similarity between the different question texts; Performing clustering processing on the word vectors corresponding to the Q question texts to obtain one or more cluster groups; a cluster group includes the first data pairs corresponding to one or more question texts whose vector distances meet the distance requirement; and A reference data pair is selected from the one or more cluster groups and added to the first data pair set; and first data pairs of the one or more cluster groups except the selected reference data pair are added to the second data pair set.

16. The method according to any one of claims 1 to 15, characterized in that: The search library includes feature vectors of P text blocks, where P is a positive integer and P is less than Q; the method further includes: In an interactive dialogue scenario, a first question text input by a subject is received, and the first question text is subjected to vector embedding processing by using the second retrieval model to obtain a feature vector of the first question text; the feature vector of the first question text is used to represent semantic information of the first question text; Compare the feature vector of the first question text with the feature vectors of the P text blocks in the search library Similarity calculation processing is performed to obtain matching scores between the P text blocks and the first question text respectively; the matching scores are used to indicate the credibility of the text blocks as the source of the answer to the first question text; Sorting the P matching scores in descending order, and performing a gradient operation on each matching score in the sorted score sequence to obtain multiple gradient information; Based on the plurality of gradient information, dynamically select one or more matching scores whose gradient information satisfies the gradient descent condition from the score sequence, and use the text blocks corresponding to the one or more matching scores as the source of the answer to the first question text; and A first answer is generated for the first question text based on one or more text blocks corresponding to the matching scores.

17. The method according to claim 16, characterized in that Generating a first answer for the first question text based on one or more text blocks corresponding to the matching scores includes: Performing task classification processing on the first question text to obtain a classification result, wherein the classification result is used to indicate the difficulty of answering the first question text; According to the answer difficulty indicated by the classification result, using an answer generation model that matches the answer difficulty to generate an initial answer to the first question text based on one or more text blocks corresponding to the matching scores; and The initial answer is reviewed to obtain a first answer to the first question text; the first answer is the initial answer, or an answer after the initial answer is fine-tuned.

18. The method according to claim 16 or 17, characterized in that The interactive dialogue scenario includes multiple rounds of interactive dialogues, and any round of interactive dialogue except the first round of interactive dialogue in the multiple rounds of interactive dialogues is represented as a target round of interactive dialogue; After receiving the first question text input by the object in the interactive dialogue scene, the method further includes: Acquire a first question text input by the object in the target round interactive dialogue, historical dialogue data in the historical round interactive dialogue, and historical object data about the object; the historical object data is generated based on the historical dialogue data; Based on the first question text, the historical dialogue data and the historical object data, the target text question is rewritten to obtain a new first question text; and The adopting the second retrieval model to perform vector embedding processing on the first question text includes: The second retrieval model is used to perform vector embedding processing on the new first question text.

19. A retrieval model training device, characterized in that: include: An acquisition unit, used for acquiring a knowledge document and performing text segmentation processing on the knowledge document to obtain a text block set; The text block set includes a plurality of text blocks, each of which is composed of a reference text block belonging to the knowledge document and general description information corresponding to the reference text block, wherein the general description information is obtained by summarizing the text semantics of the corresponding reference text block; A processing unit is used to write a question for each text block based on the text block set and the general information included in each text block to obtain one or more question texts corresponding to each text block, combine the one or more question texts corresponding to each text block with the corresponding text block respectively to obtain one or more data pairs corresponding to each text block, and generate a data pair set based on the one or more data pairs corresponding to each text block; each data pair consists of a matching question text and a text block; the text block in each data pair serves as a source of an answer to the matching question text; and The processing unit is also used to train the first retrieval model using the data pair set to obtain a second retrieval model; the training is used to make the feature difference between the matching question text and the text block in each of the data pairs less than a preset threshold; the second retrieval model corresponds to a retrieval library, and the retrieval library includes a feature vector of each text block in the text block set; wherein the second retrieval model and the retrieval library are used to generate a first answer for the first question text in an interactive dialogue scenario.

20. A computer device, characterized in that: include: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the retrieval model training method according to any one of claims 1 to 18 is implemented.

Citation Information

Patent Citations

  • Intelligent question and answer method and system combining retrieval and generation

    CN112364150A

  • Semantic retrieval and question and answer processing method and device for long text and electronic equipment

    CN115630136A

  • Intelligent question-answering system and method suitable for vertical field and application of intelligent question-answering system and method

    CN116805001A

  • Knowledge base construction method and question and answer dialogue method and system based on generative large language model

    CN117056471A

  • Questions and answers generation

    US20110125734A1

Cited By

  • Text partitioning method and device, storage medium and electronic equipment

    CN120448524A

  • Evidence tracing method and device for reply content of medical model, medium and product

    CN120561252A

  • Knowledge retrieval method based on multilayer index

    CN120633843A

  • Power business question answering method and device, terminal equipment and storage medium

    CN120705283A

  • A power service problem solving method and device, a terminal device, and a storage medium

    CN120705283B