Question and answer pair generation method and system
By building a knowledge graph at the document level and unit level, and automatically generating Q&A pairs, it solves the problem of cumbersome and expensive generation of Q&A pairs in the existing technology, and realizes the richness and diversity of Q&A pairs, and is suitable for scenarios such as intelligent customer service and FAQ systems.
Patent Information
- Application Number
- CN202510414934.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The process of generating a question-and-answer pair in the prior art is cumbersome, time-consuming, costly, and low-quality, making it difficult to build a comprehensive and accurate FAQ library.
By constructing a knowledge graph at the document level and unit level, sample text data sets and text block data sets respectively to generate Q&A pairs, and use large language models and template generation methods to automatically generate Q&A pairs.
It improves the richness, comprehensiveness and diversity of Q&A pairs, reduces the workload of manual writing, improves generation efficiency and accuracy, and is suitable for scenarios such as intelligent customer service, knowledge base and FAQ system.
Smart Images

Figure CN120258146A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular, to a method and system for generating question-and-answer pairs. Background Art
[0002] In the Retrieval-Augmented Generation (RAG) technology, constructing a comprehensive and accurate Frequently Asked Questions (FAQ) library that includes Question-Answer Pairs (QA pairs) plays an important and non-negligible role in improving service quality and optimizing user experience.
[0003] In related technologies, QA pairs are mainly generated through manual annotation. For example, traditional methods usually sample manual annotation to generate QA pairs and build an FAQ library based on the QA pairs.
[0004] However, generating QA pairs through manual annotation not only has a cumbersome process, is time-consuming and laborious, but also has drawbacks such as high costs and low quality.
[0005] It should be noted that the content of the above related technologies is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the filing date of this specification, nor does it mean that it can become the prior art of this specification. Summary of the Invention
[0006] This specification provides a method and system for generating question-and-answer pairs to avoid at least one of the above technical problems.
[0007] In a first aspect, this specification provides a method for generating question-and-answer pairs, including:
[0008] Constructing a knowledge base based on the obtained knowledge data, where the knowledge base includes a text data set at the document level and a text block data set at the unit level;
[0009] Constructing a knowledge graph corresponding to each of the text data set and the text block data set, where the knowledge graph is used to represent the relevance between different documents in the data set, and the data set includes the text data set and the text block data set; and
[0010] Sampling the documents in the text data set and the text block data set respectively to obtain sampled documents, and generating question-and-answer pairs based on the sampled documents and the knowledge graph corresponding to the sampled documents.
[0011] In a second aspect, this specification provides a system for generating question-and-answer pairs, including:
[0012] At least one storage medium storing at least one instruction set for generating question-answer pairs;
[0013] At least one processor communicatively connected to the at least one storage medium, wherein when the at least one processor runs, it reads the at least one instruction set and executes the method described in the first aspect according to the instructions of the at least one instruction set.
[0014] In a third aspect, this specification provides a computer-readable non-transitory storage medium, wherein at least one instruction set is stored in the computer-readable non-transitory storage medium, and the at least one instruction set is executed by at least one processor to implement the method described in the first aspect.
[0015] As can be seen from the above technical solutions, the method and system for generating question-answer pairs provided in this specification divide the document into a coarse-grained document and a fine-grained document to construct a coarse-grained knowledge graph and a fine-grained knowledge graph, and on this basis, generate coarse-grained question-answer pairs and fine-grained question-answer pairs respectively. This can improve the richness, comprehensiveness, and diversity of the generated question-answer pairs.
[0016] Other functions of the method and system for generating question-answer pairs provided in this specification will be partially listed in the following description. The creative aspects of the method and system for generating question-answer pairs provided in this specification can be fully explained through practice or the use of the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic diagram of the application scenario of the method for generating question-answer pairs provided in the embodiments of this specification;
[0019] Figure 2 It is a schematic diagram of the structure of the system for generating question-answer pairs provided in the embodiments of this specification;
[0020] Figure 3 It is a schematic flowchart of the method for generating question-answer pairs provided in the embodiments of this specification;
[0021] Figure 4 It is a schematic diagram of the principle of the method for generating question-answer pairs provided in an embodiment of this specification;
[0022] Figure 5Schematic diagram of the principle of the method for generating question-and-answer pairs provided in an embodiment of this specification;
[0023] Figure 6 Schematic diagram of the principle of normalizing similar questions for multiple question-and-answer pairs provided in an embodiment of this specification. Detailed implementation manners
[0024] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0025] It should be understood that the terms "include" and "have" and any variations thereof in the embodiments of this specification are intended to cover but not exclude inclusion. For example, a product or device including a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0026] The term "and / or" in the embodiments of this specification describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0027] The term "multiple" in the embodiments of this specification means two or more, and other quantifiers are similar.
[0028] The terms "first", "second", "third", etc. in this specification are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be interchanged under appropriate circumstances, for example, they can be implemented in an order other than those given in the illustration or description of the embodiments of this specification.
[0029] The term "unit / module" used in this specification refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to that element.
[0030] For the convenience of readers' understanding of this specification, the application scenarios of this specification are now introduced.
[0031] The technical solutions provided in this specification are applicable to scenarios that require generating question-and-answer pairs. Exemplarily, the technical solutions provided in this specification can be applied to intelligent customer service, knowledge bases and FAQ systems, the education field, search engine optimization (SEO), medical and health consultations, legal consultations, games and entertainment, news and media, etc., and are not listed one by one here.
[0032] For example, taking the application of the technical solutions provided in this specification to the scenario of intelligent customer service as an example:
[0033] Through the technical solutions provided in this specification, question-and-answer pairs in the intelligent customer service scenario can be generated. For example, the question in the question-and-answer pair is: How to reset the password? The answer in the question-and-answer pair is: Please visit the login page, click "Forgot Password", and follow the prompts.
[0034] Relatively speaking, by applying the technical solutions provided in this specification to the scenario of intelligent customer service, the customer service efficiency can be improved, and the workload of manual customer service can be reduced. Moreover, it can provide services all day long.
[0035] Another example is taking the application of the technical solutions provided in this specification to knowledge bases and FAQ systems as an example:
[0036] In an enterprise or a product support website, through the technical solutions provided in this specification, common questions and answers can be sorted out to form a knowledge base or an FAQ system including question-and-answer pairs. For example, the FAQ system includes question-and-answer pairs, and the question in the question-and-answer pair is: How long is the warranty period of the product? The answer in the question-and-answer pair is: This product provides one-year free warranty service.
[0037] Relatively speaking, by applying the technical solutions provided in this specification to the scenario of knowledge bases and FAQ systems, it can help users quickly find the information they need. And it can reduce repeated questions and improve the user experience.
[0038] Regarding the description of the application of the technical solutions provided in this specification to other scenarios, reference can be made to the above examples about the scenarios of intelligent customer service, knowledge bases and FAQ systems, and they are not listed one by one here.
[0039] Figure 1 It is a schematic diagram of the application scenario of the method for generating question-and-answer pairs (hereinafter simply referred to as the generation method) of the embodiments of this specification. Among them, the generation method of this specification can be applied to, for example, Figure 1 the scenario 100 as shown. As Figure 1 shown, the scenario 100 may include a target user 101, a client 102, a server 103, and a network 104.
[0040] The target user 101 can be the user who triggers the generation of question-and-answer pairs. For example, the target user 101 can perform a target operation on the client 102 to trigger the generation of question-and-answer pairs.
[0041] The client 102 can be an electronic device that provides an interactive function to the target user 101. For example, the client 102 can provide an interactive interface to the target user 101, and the target user 101 can perform interactive operations on the interactive page. In some embodiments, in response to detecting an operation for generating question-and-answer pairs triggered by the target user 101, the client 102 executes the generation method described in this specification. At this time, the client 102 may store data or instructions for executing the generation method described in this specification and may execute or be used to execute the data or instructions. In some embodiments, the client 102 may include a hardware device with data information processing capabilities and necessary programs for driving the hardware device to work to execute the generation method described in this specification.
[0042] In some embodiments, the client 102 may include a mobile device, a tablet computer, a laptop computer, a built-in device of a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant, a gaming device, a navigation device, etc., or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an in-vehicle computer, an in-vehicle TV, etc.
[0043] In some embodiments, the client 102 may be installed with one or more applications (APPs). The APP can provide the target user 101 with the ability to interact with the outside world through the network 104 and an interface. The APP includes, but is not limited to: web browser APP programs, search APP programs, chat APP programs, shopping APP programs, video APP programs, financial management APP programs, instant messaging tools, email clients, social platform software, and so on.
[0044] As Figure 1 shown, the client 102 can be communicatively connected to the server 103. Among them, the server 103 can be communicatively connected to one client 102 or multiple clients 102. In some embodiments, the client 102 can interact with the server 103 through the network 104 to receive or send messages, etc. For example, the client 102 can interact with the server 103 through the network 104 to send a request for generating question-and-answer pairs to the server 103.
[0045] Server 103 can be a server that provides various services. For example, server 103 can be a cloud server or a local server. Server 103 can communicate with a client 102 and receive data sent by the client 102, or can communicate with multiple clients 102 and receive data sent by each client 102 respectively.
[0046] In some embodiments, the generation method described in this specification can be executed on server 103. At this time, server 103 can store data or instructions for executing the generation method described in this specification, and can execute or be used to execute the data or instructions. Server 103 can include a hardware device with data information processing capabilities and necessary programs for driving the hardware device to work.
[0047] Network 104 is a medium for providing a communication connection between client 102 and server 103. Network 104 can facilitate the exchange of information or data. As Figure 1 shown, client 102 and server 103 can be respectively connected to network 104, and transmit information or data to each other through network 104.
[0048] In some embodiments, network 104 can be any type of wired or wireless network, or a combination thereof. For example, network 104 can include a cable network, a wired network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network TM, a short-range wireless network (ZigBee TM), a near field communication (NFC) network, or a similar network.
[0049] In some embodiments, network 104 can include one or more network access points. For example, network 104 can include a wired or wireless network access point, such as a base station or an Internet exchange point, through which one or more components of client 102 and server 103 can be connected to network 104 to exchange data or information.
[0050] It is worth noting that Figure 1The numbers of the client 102, the server 103, and the network 104 in [it] are merely illustrative. According to implementation requirements, there can be any number of clients 102, servers 103, and networks 104. Moreover, the generation method provided in this specification can be executed entirely on the client 102, entirely on the server 103, or partially on the client 102 and partially on the server 103.
[0051] That is to say, Figure 1 and the above description for Figure 1 is only used to exemplarily elaborate on the application scenarios to which the generation method in this specification may apply, and should not be construed as a limitation on the application scenarios.
[0052] Figure 2 The hardware structure diagram of a generation system (hereinafter simply referred to as the generation system) 200 provided according to an embodiment of this specification is shown. The generation system 200 can execute the generation method described in this specification. The generation method is introduced in other parts of this specification. When the generation method is executed on the client 102, the generation system 200 can be the client 102. When the generation method is executed on the server 103, the generation system 200 can be the server 103. When the generation method is partially executed on the client 102 and partially on the server 103, the generation system 200 can be a system including the client 102 and the server 103.
[0053] As Figure 2 shown, the generation system 200 can include at least one storage medium 203 and at least one processor 202. In some embodiments, the generation system 200 can further include a communication port 204 and an internal communication bus 201. The generation system 200 can further include I / O components 205.
[0054] The internal communication bus 201 can connect different system components. For example, the internal communication bus 201 can connect the storage medium 203, the processor 202, the communication port 204, and the I / O components 205.
[0055] The I / O components 205 support input / output between the generation system 200 and other components.
[0056] The communication port 204 is used for data communication between the generation system 200 and the outside world. For example, the communication port 204 can be used for data communication between the generation system 200 and the network 104. The communication port 204 can be a wired communication port or a wireless communication port.
[0057] The storage medium 203 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 2031, a read-only storage medium (ROM) 2032, or a random access storage medium (RAM) 2033. The storage medium 203 further includes at least one instruction set stored in the data storage device. The instruction set includes computer program code, and the computer program code may include programs, routines, objects, components, data structures, procedures, modules, etc. for executing the generation method provided in this specification.
[0058] At least one processor 202 may be communicatively connected to at least one storage medium 203. The at least one processor 202 is configured to execute the above at least one instruction set. When the generation system 200 runs, the at least one processor 202 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the generation method provided in this specification. The processor 202 may execute all steps included in the generation method. The processor 202 may be in the form of one or more processors. In some embodiments, the processor 202 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.
[0059] For purposes of illustration only, only one processor 202 is shown in the generation system 200 in the drawings. However, it should be noted that the generation system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 202 of the generation system 200 executes step A and step B, it should be understood that step A and step B may also be jointly or separately executed by two different processors 202 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0060] Please refer to Figure 3 , Figure 3 which is a flowchart of the method for generating question-and-answer pairs provided in the embodiments of this specification. As Figure 3 shown, the method includes the following S301 to S304:
[0061] S301: Construct a knowledge base based on the obtained knowledge data. The knowledge base includes a collection of text data at the document level and a collection of text chunk data at the chunk level.
[0062] Knowledge data can also be referred to as knowledge sources, which can be understood as information and data used to generate question-and-answer pairs.
[0063] Regarding obtaining knowledge data, the following examples can be used for implementation:
[0064] In one example, the generation system can be connected to a collection device and receive the knowledge data collected and sent by the collection device.
[0065] In another example, the generation system can provide a tool for loading data, and the user can use this tool for loading data to transfer the knowledge data to the generation system.
[0066] Among them, the tool for loading data can be an interface for connecting to an external device, such as an interface for connecting to other storage devices, and obtain the knowledge data transmitted by the external device through this interface; the tool for loading data can also be a display device. For example, the generation system can output an interface for the function of loading data on the display device, and the user can import the knowledge data into the generation system through this interface.
[0067] In some embodiments, the knowledge data includes data related to the target field collected from multiple channels, and the target field is the field corresponding to the question-and-answer pair.
[0068] Combined with the above analysis, it can be seen that the generation method provided in this specification can be applied to different fields and scenarios. For different fields and scenarios, the data used to generate the data in that field may not be the same. For example, the knowledge data for generating question-and-answer pairs in the intelligent customer service scenario and the medical and health consultation scenario is different.
[0069] Therefore, in this embodiment, the generation system can collect information and data related to the corresponding field through various channels to obtain the knowledge data for generating question-and-answer pairs in the corresponding field.
[0070] That is to say, the generation system can sample multiple channels to obtain multiple data sources based on the actual scenarios and field requirements such as the generation of question-and-answer pairs, and finally generate question-and-answer pairs based on multiple data sources. This can improve the accuracy, comprehensiveness, diversity, and practicality of the generated question-and-answer pairs.
[0071] In some embodiments, the knowledge data includes: structured data, semi-structured data, and unstructured data.
[0072] Exemplarily, based on whether the knowledge data is structured, the generation system can classify the knowledge data into three categories: structured data, semi-structured data, and unstructured data.
[0073] Among them, structured data includes tables, icons, reports, etc. Relatively speaking, this type of data has a fixed format and organization method, which is convenient for formatted understanding. Relatively speaking, due to its clear fields and relationships, structured data can support the efficient generation of question-answer pairs.
[0074] Semi-structured data includes Hyper Text Markup Language (HTML) pages, Extensible Markup Language (XML) files, Markdown (MD) files, etc. Relatively speaking, this type of data usually contains recognizable fields or tags, as well as clear titles, paragraphs, and hierarchical structures. When generating question pairs, its tags such as titles can be recognized to assist in generation.
[0075] Unstructured data is pure text without a fixed format. For example, user feedback and comments such as user reviews, social media content, forum posts, etc.
[0076] In this embodiment, since the knowledge data includes structured data, semi-structured data, and unstructured data, the generated question-answer pairs can cover various aspects of content to ensure that each small knowledge point can be converted into clear questions and answers.
[0077] Correspondingly, after obtaining the knowledge data, the generation system can construct two different granularity data sets based on the knowledge data: a text data set at the document level and a text block data set at the unit level. Among them, the text data set at the document level can also be called a text data set with a complete text granularity, and the text block data set at the unit level can also be called a text block data set with a fragment text granularity.
[0078] Exemplarily, the document level can be understood as the overall perspective of the entire document. A document from the overall perspective can be a complete article, a book chapter, a web page content, or any form of long text. Relatively speaking, at this level, when the generation system generates question-answer pairs, it pays more attention to the overall structure, theme, logical relationship, and global information of the entire document.
[0079] The unit level can be understood as a document that divides the entire document into smaller units, and these units can be sentences, paragraphs, fragments, or other custom text blocks. Relatively speaking, at this level, when the generation system generates question-answer pairs, it pays more attention to the local information of the entire document, and the efficiency is relatively higher.
[0080] Relatively speaking, such as Figure 4As shown, after obtaining knowledge data, the generation system can process the knowledge data, such as including document parsing of the knowledge data. Through document parsing, the generation system can generate a data set of documents including two different granularities.
[0081] That is to say, in this step, the generation system can analyze the knowledge data to extract text data at the overall granularity from it to obtain a text data set; extract text block data at the local granularity from it to obtain a text block data set. Correspondingly, the generation system finally obtains a data set including two different granularities (that is, a coarse-grained text data set and a fine-grained text block data set). So that the generation system can generate both global question-and-answer pairs and local question-and-answer pairs.
[0082] In some embodiments, S301 may include the following steps 11 to 13:
[0083] Step 11: Parse the knowledge data to obtain data in a preset text format.
[0084] Combined with the above analysis, it can be seen that the knowledge data may be from multiple different knowledge sources and may have different formats. The generation system can perform document parsing on the knowledge data to parse the knowledge data into a unified text format, such as a preset text format.
[0085] Among them, the preset text format can be determined by the generation system based on requirements, historical records, experiments, etc., and this embodiment does not make any limitations.
[0086] Step 12: Convert the data in the preset text format into data in Markdown format.
[0087] After parsing the knowledge data in different text formats into a unified text format, the generation system can convert the data in the unified text format into Markdown format. So as to facilitate the generation system to extract structured information such as titles, paragraphs, and sentences later, and make full use of the document content to generate question-and-answer pairs.
[0088] Combined with the above analysis, it can be seen that the knowledge data may include dense information such as tables and charts. In some embodiments, the generation system can extract this part of the dense information and convert the extracted information into data in Markdown format. To avoid the drawback that in related technologies, such as in traditional retrieval-enhanced generation question-and-answer assistants, dense information such as tables and charts causes bottlenecks in the question-and-answer effect. The generation system can generate question-and-answer pairs through the extracted information, and can effectively handle various retrieval and interpretive questions about contents such as tables and icons.
[0089] Among them, retrieval-augmented generation is a technique that combines retrieval and generation, aiming to enhance the generation ability of large language models (LLMs) by dynamically retrieving relevant information. Such methods usually rely on external knowledge bases or document collections. When dealing with natural language tasks, they can first retrieve the context or facts related to the input from the knowledge base, and then use the large language model to generate accurate answers or texts based on the retrieval results. In this way, the model can incorporate external knowledge during the generation process, improving its performance in complex tasks, especially in scenarios that require specific domain knowledge or real-time information.
[0090] Large language models can be understood as neural network models with a huge scale and numerous parameters, used to process and generate natural language texts. Such models usually rely on the Transformer architecture and can be pre-trained on a large amount of text data to capture the deep syntactic and semantic rules of language. These models can perform zero-shot or few-shot learning through the method of learning from context.
[0091] Step 13: Build a knowledge base based on the data in Markdown format.
[0092] After converting the unified text format to Markdown format, the generation system can generate a data set including two different granularity texts based on the data in Markdown format.
[0093] Combining the above analysis of steps 11 to 13, it can be seen that in this embodiment, the generation system constructs a knowledge base on the basis of Markdown format by converting the format of knowledge data. Relatively speaking, the unified format is convenient for the generation system to process data, especially the Markdown format is convenient for extracting structured information such as titles, paragraphs, and sentences. Therefore, the efficiency, accuracy, and reliability of generating question-answer pairs in the later stage can be improved.
[0094] S302: Construct knowledge graphs corresponding to the text data set and the text block data set respectively, where the knowledge graph is used to represent the relevance between different documents in the data set, and the data set includes the text data set and the text block data set.
[0095] Continuing to combine the above examples and Figure 4 , knowledge processing also includes the construction of knowledge graphs.
[0096] Exemplarily, for the coarse-grained and fine-grained data sets, the generation system constructs corresponding knowledge graphs respectively. That is, the knowledge graph includes a knowledge graph corresponding to the coarse-grained text data set and a knowledge graph corresponding to the fine-grained text block data set.
[0097] A knowledge graph can be understood as integrating information from different document sources into a structured network for easy querying and reasoning. For example, the knowledge graph corresponding to a text data set can be understood as the generation system integrating the information of the coarse-grained documents in the text data set into a structured network, and this structured network can be used to represent the associations between the information of the coarse-grained documents. The knowledge graph corresponding to a text block data set can be understood as the generation system integrating the information of the fine-grained documents in the text block data set into a structured network, and this structured network can be used to represent the associations between the information of the fine-grained documents.
[0098] S303: Sample the documents in the text data set and the text block data set respectively to obtain sampled documents.
[0099] Continuing with the above example and Figure 4 , knowledge processing also includes sampling of documents.
[0100] Exemplarily, the sampling can include sampling in two non-interfering aspects.
[0101] For example, on the one hand, the generation system can sample the text data set to complete subsequent generation of coarse-grained question-and-answer pairs based on the corresponding sampling results, such as the sampled text obtained by sampling the text data set.
[0102] On the other hand, the generation system can sample the text block data set to complete subsequent generation of fine-grained question-and-answer pairs based on the corresponding sampling results, such as the sampled text obtained by sampling the text block data set.
[0103] This embodiment does not limit the way the generation system samples to obtain sampled documents. For example, the generation system can sample to obtain sampled documents based on a traversal method.
[0104] S304: Generate question-and-answer pairs according to the sampled documents and the knowledge graphs corresponding to the sampled documents.
[0105] Exemplarily, combining the above analysis, this step can be understood as:
[0106] If the sampled document is a coarse-grained sampled document obtained by sampling the text data set, the generation system can generate coarse-grained question-and-answer pairs based on this coarse-grained sampled document and the coarse-grained knowledge graph.
[0107] If the sampled document is a fine-grained sampled document obtained by sampling the text block data set, the generation system can generate fine-grained question-and-answer pairs based on this fine-grained sampled document and the fine-grained knowledge graph.
[0108] Based on the above analysis of S301 to S304, it can be seen that in this embodiment, the generation system can divide the document into a coarse-grained document and a fine-grained document to construct a coarse-grained knowledge graph and a fine-grained knowledge graph, and on this basis, generate coarse-grained Q&A pairs and fine-grained Q&A pairs respectively. It can avoid the disadvantages of high cost, insufficient coverage, and uneven quality faced in the related art when highly relying on manual annotation to generate Q&A pairs. It can realize the automatic generation of Q&A pairs and improve the richness, comprehensiveness, and diversity of the generated Q&A pairs.
[0109] In some embodiments, S304 may include the following steps 21 and 22:
[0110] Step 21: Determine the target generation method from the preset generation methods according to the data type of the sampled document, where the preset generation methods include the large language model generation method and the template generation method.
[0111] Continuing to combine the above examples and Figure 4 , the generation system can generate Q&A pairs based on the large language model generation method or generate Q&A pairs based on the template generation method. The generation system can automatically generate Q&A pairs through the large language model method and the template generation method, reducing the workload of manual writing and updating.
[0112] Exemplarily, the generation system can select a relatively more suitable generation method from the two different generation methods based on the data type, actual application scenario, etc., and sample the more suitable generation method to generate Q&A pairs. And the two different methods include: the large language model generation method and the template generation method.
[0113] As Figure 5 shown, after obtaining knowledge data (such as multiple knowledge sources), the generation system can automatically generate Q&A pairs based on the knowledge data. And combining Figure 4 it can be known that the automatic generation of Q&A pairs by the generation system includes knowledge processing and automatically generating Q&A pairs based on the large language model generation method or the template generation method.
[0114] The large language model generation method can be understood as using the powerful natural language generation ability of the large language model to automatically generate Q&A pairs.
[0115] The template generation method can be understood as pre-defining one or more groups of question templates, and then filling the pre-defined question templates according to the content of the sampled document to automatically generate Q&A pairs. Continuing to combine the above examples and Figure 4 , in the template generation method, it can at least include two links: structured parsing and template filling. For example, the generation system can perform structured parsing on the sampled document and perform template filling operations based on the parsing results.
[0116] The generation system can determine the target generation method by referring to the following examples:
[0117] For example, the large language model generation method can handle unstructured and semi-structured data and generate natural and fluent questions and answers. Therefore, if the data type of the sampled document is unstructured or semi-structured, the generation system can determine the large language model generation method as the target generation method.
[0118] Another example is that the large language model generation method has stronger flexibility, higher diversity, and higher automation. Therefore, for practical application scenarios with stronger flexibility, higher diversity, and higher automation, the generation system can determine the large language model generation method as the target generation method.
[0119] For another example, the template generation method is relatively more suitable for structured data to ensure the accuracy and consistency of generating question-answer pairs. For example, the templates can be "{What is the {Y} of {X}?}", "{What is the function of {X}?}", "{What is the reason for encountering the {X} error?}", and so on. For complete information such as tables and icons, using the template generation method can relatively more quickly generate question-answer pairs in batches.
[0120] Also, the template generation method has relatively higher controllability and lower computational resource consumption. Therefore, for practical application scenarios with higher controllability and lower computational resource consumption, the generation system can determine the template generation method as the target generation method.
[0121] Step 22: Adopt the target generation method to generate question-answer pairs according to the sampled document and the knowledge graph corresponding to the sampled document.
[0122] Correspondingly, after determining the target generation method, the generation system can automatically generate question-answer pairs using the target generation method.
[0123] Combining the above analysis of Steps 21 and 22, it can be seen that in this embodiment, the generation system can select the corresponding target generation method based on the corresponding requirements to achieve the diversity and flexibility of generating question-answer pairs. In addition, it can also achieve the automatic generation of question-answer pairs, which can reduce the workload of manually writing question-answer pairs. Moreover, it can also improve the efficiency, accuracy, and reliability of generating question-answer pairs.
[0124] Combining the above analysis, it can be seen that the generation system can determine the large language model generation method as the target generation method, or it can also determine the template generation method as the target generation method. In some embodiments, if the target generation method is the large language model generation method, Step 22 may include the following Steps 221 and 222:
[0125] Step 221: Obtain relevant documents corresponding to the sampled document from the knowledge graph corresponding to the sampled document.
[0126] Combined with the above example, if the sampled document is a coarse-grained document and the knowledge graph is a coarse-grained knowledge graph, the generation system can, based on the coarse-grained document, determine documents that are logically related to the coarse-grained document from the coarse-grained knowledge graph to obtain coarse-grained relevant documents.
[0127] If the sampled document is a fine-grained document and the knowledge graph is a fine-grained knowledge graph, the generation system can, based on the fine-grained document, determine documents that are logically related to the fine-grained document from the fine-grained knowledge graph to obtain fine-grained relevant documents.
[0128] Step 222: Input the sampled document, relevant documents, and a preset prompt into a large language model to obtain question-answer pairs, where the question-answer pairs include questions, answers, and corresponding bases, and the preset prompt is used to guide the large language model to generate question-answer pairs based on the sampled document and relevant documents.
[0129] Exemplarily, the generation method of the large language model can be further understood as using the powerful natural language generation ability of the large language model to guide the large language model to automatically generate corresponding question-answer pairs through a corresponding prompt (such as the preset prompt). To make full use of the in-depth understanding and generation ability of the large language model to create questions closely related to the sampled document, and then generate accurate and detailed answers.
[0130] Moreover, in this embodiment, the data structure of the question-answer pair can include three-dimensional content: question (Query, Q), (original text) basis (Evidence, E), and answer (Answer, A). Correspondingly, the data structure of the question-answer pair in this embodiment can be represented by QEA (Question-Evidence-Answer).
[0131] Continuing to combine the above example and Figure 4 , the generation method of the large language model includes the design of the question-answer pair structure, and specifically, it is a structure of question + basis + answer (i.e., QEA).
[0132] That is to say, in this embodiment, the generation system can, based on a preset prompt, guide the large language model to generate corresponding evidence when generating questions and answers. This is in line with the "Chain of Thought" (COT) habit of the large language model, making the generated questions and answers more coherent and logical, thereby improving the overall quality of the question-answer pairs. Additionally, when necessary, the accuracy and relevance of the evidence can be checked to more effectively evaluate and optimize the generated question-answer pairs, ensuring the reliability and practicality of their content.
[0133] It can be understood that the questions can include single-hop questions and multi-hop questions. In some embodiments, for single-hop questions, the generation system can be implemented by traversing the documents to ensure that each document has a chance to be selected. For multi-hop questions, the generation system can be implemented by traversing + knowledge graph. In this way, the relevance between documents can be understood through the knowledge graph, so as to select those logically related documents together to determine the question-answer pairs.
[0134] Combined with the above analysis, it can be seen that the preset prompt can guide the large language model to generate corresponding evidence when generating questions and answers. In some embodiments, the preset prompt can also guide at least one of the style, wording, and quantity of the question-answer pairs.
[0135] Continuing with the above example and Figure 4 , in the template generation method, a controllable generation strategy can also be included. The generation system can generate question-answer pairs based on the controllable generation strategy. And the controllable generation strategy includes control strategies for one or more of the style, wording, and quantity of the question-answer pairs.
[0136] For example, when using the large language model to generate question-answer pairs, the generation system can, through a preset prompt, guide the large language model to generate question-answer pairs that conform to a specific style or wording. In this way, the output of the large language model can be restricted by the preset prompt to meet the expected style or wording requirements.
[0137] Another example is that the generation system can also control the number of question-answer pairs generated by the large language model based on the sampled documents through a preset prompt. Specifically, the preset prompt can include an indication of the number of question-answer pairs to be generated by the large language model; or, the preset prompt can instruct the large language model to automatically detect the information volume of the sampled documents, such as by the method of keyword density extraction and the method of document length weighting, and determine the corresponding number according to the information volume.
[0138] In this embodiment, the generation system can improve the flexibility, diversity, and reliability of generating Q&A pairs by guiding the large language model to automatically generate the style, wording, and quantity of Q&A pairs through a preset prompt, so that the generated Q&A pairs can meet the requirements of different dimensions.
[0139] Combined with the above analysis, it can be seen that the quantity of Q&A pairs can be determined by the preset prompt. The quantity of Q&A pairs corresponding to one sampled document may be one or multiple. Usually, the quantities of the sampled documents corresponding to the coarse-grained data set and the fine-grained data are multiple. Therefore, the quantity of Q&A pairs finally generated by the generation system is usually multiple.
[0140] In some embodiments, when the quantity of Q&A pairs is multiple, the generation system can automatically optimize the multiple Q&A pairs, such as performing normalization of similar questions (which can be simply referred to as normalization) to obtain the Q&A pairs corresponding to various types of questions, so that the same type of questions correspond to one answer.
[0141] Continuing to combine the above example and Figure 5 , after generating the Q&A pairs, the generation system can automatically optimize the Q&A pairs. And as Figure 4 shown, the automatic optimization of Q&A pairs can include Q&A pair normalization, that is, the generation system can perform normalization of similar questions on the Q&A pairs.
[0142] Exemplarily, the ways users ask questions are diverse, which may lead to different expressions of the same question. For example, "How to register an account?" and "What are the steps for account registration?" are actually the same question. In this embodiment, the generation system can reduce the redundancy of Q&A pairs and improve the query efficiency during the query application process by performing normalization of similar questions (or called normalization of homologous questions), and merging similar questions into one standard question.
[0143] Combined with the above example, it can be seen that a Q&A pair can include a question and an answer. In some embodiments, the above "performing normalization of similar questions on multiple Q&A pairs to obtain the Q&A pairs corresponding to various types of questions" may include the following steps 31 to 33:
[0144] Step 31: Cluster similar questions into a question cluster.
[0145] Exemplarily, the generation system can cluster multiple similar questions (or multiple questions with the same essence) into a question cluster by means of clustering.
[0146] In some embodiments, step 31 may include the following steps 311 to 313:
[0147] Step 311: For any new question, convert the new question into a new question vector.
[0148] Exemplarily, when the generation system obtains each new problem, it can convert the new problem into a vector to obtain a corresponding new problem vector.
[0149] For example, as Figure 6 shown, the new problem can be understood as the currently generated problem, which can be simply referred to as the current problem as Figure 6 shown. The generation system can perform vector conversion on the new problem to obtain a new problem vector. Correspondingly, the new problem vector can be referred to as the current problem vector as Figure 6 shown.
[0150] Among them, the generation system can use the bgm-m3 model to convert each new problem into a high-dimensional vector representation to obtain a corresponding new problem vector.
[0151] Step 312: If there is a target problem cluster in the existing problem clusters that includes a similar problem vector corresponding to the new problem vector, then merge the new problem into the target problem cluster.
[0152] Step 313: If there is no problem cluster in the existing problem clusters that includes a similar problem vector corresponding to the new problem vector, then create a problem cluster corresponding to the new problem.
[0153] Exemplarily, after determining the vector corresponding to the new problem, the generation system can determine whether there is already a problem cluster corresponding to the new problem vector in the existing problem clusters. If so, execute Step 312; otherwise, if not, execute Step 313.
[0154] Continuing with the above example and Figure 6 , for each generated problem, after the generation system converts it into a corresponding vector, it can store the corresponding vector in a vector library (Elasticsearch, ES). So as to utilize the efficient retrieval ability of the vector library for similarity clustering.
[0155] Correspondingly, for the current problem vector, the generation system can retrieve similar problems through the vector library and dynamically update the clustering result so as to merge similar problems into one problem cluster.
[0156] Continuing to refer to Figure 6 , the generation system can perform similarity clustering on the current problem vector. If there is already an existing problem cluster to which the current problem vector belongs, the generation system adds the current problem to the existing problem cluster. Conversely, if there is no problem cluster to which the current problem vector belongs, then a new problem cluster including the current problem is newly created.
[0157] Specifically, for the newly generated question in the first time, the generation system converts it into a vector and determines a question cluster for this question. For the newly generated question in the second time, the generation system converts it into a vector and calculates the similarity between it and the vector of the newly generated question in the first time. If the similarity is relatively high, that is, the newly generated questions in the first and second times may be the same question, then the generation system merges the newly generated question in the second time and the question cluster generated in the first time into one question cluster. And so on, which will not be enumerated one by one here.
[0158] Combined with the above analysis of steps 311 to 313, it can be seen that in the embodiment, by converting the question into the form of a vector to determine the corresponding question cluster, the accuracy, reliability, and effectiveness of determining the question cluster can be improved.
[0159] Step 32: For any question cluster, perform quality scoring on the Q&A pairs within any question cluster.
[0160] Exemplarily, the generation system takes the cluster as a unit and performs quality scoring on all the question pairs within each question cluster respectively.
[0161] This embodiment does not limit the way of quality scoring. For example, the generation system can perform quality scoring based on scoring dimensions such as question quality, answer quality, Q&A consistency, and statement quality.
[0162] Among them, the question quality can include integrity, consistency, and rationality. Integrity can be understood as that the question should clearly raise a doubt, have the form of an interrogative sentence, and avoid declarative sentences or imperative sentences. Consistency can be understood as that the question should closely revolve around the theme and avoid deviating from the theme or introducing irrelevant content. Rationality can be understood as that the question should be universal, avoid being too esoteric or special cases, and ensure that it is a normal question that can be answered.
[0163] The answer quality can include key point coverage and key point detail. Key point coverage can be understood as that the answer should comprehensively cover all the key points required to answer the question and avoid missing important information. Key point detail can be understood as that the key points should be elaborated in detail, providing sufficient information to support the answer, and avoiding being too brief or vague.
[0164] Question consistency can include semantic consistency. Semantic consistency can be understood as that the question and the answer should describe the same entity or event, ensuring that the answer directly addresses the question and avoiding answering the wrong question.
[0165] The statement quality can include readability. Readability can be understood as that the question and the answer should be well-organized, the arguments should be clear, the structure should be easy to understand, and the language should be fluent and natural.
[0166] Step 33: Determine the Q&A pair with the highest quality score as the Q&A pair corresponding to any question cluster.
[0167] Exemplarily, for each question cluster, after the generation system determines the respective quality scores corresponding to each Q&A pair within the question cluster, it can determine the highest quality score therefrom, and further determine the Q&A pair corresponding to the highest quality score as the representative Q&A pair of the question cluster.
[0168] Continuing to refer to the above example and Figure 6 , the generation system can generate a question cluster set including multiple question clusters. For each question cluster in the question cluster set, the generation system can determine the Q&A pairs of the question cluster by means of normalization and optimization. That is, the generation system can determine the representative Q&A pairs corresponding to each question cluster in the question cluster set based on the method of normalization and optimization.
[0169] Combined with the above analysis of steps 31 to 33, in this embodiment, the generation system clusters similar questions into question clusters to determine relatively better-quality Q&A pairs within the question clusters for representing the Q&A pairs of the question clusters, which can ensure the effective and reliable operation of applications such as queries while avoiding duplication and redundancy of Q&A pairs.
[0170] In some embodiments, in certain actual application scenarios, in order to meet corresponding requirements, the generation system can correspondingly adjust the Q&A pairs so that the adjusted Q&A pairs meet the requirements of the actual application scenarios.
[0171] Exemplarily, the generation system can obtain the standard speech patterns corresponding to the actual application scenarios and embed the standard speech patterns in the Q&A pairs.
[0172] The standard speech pattern can be understood as a set of standardized language expressions generated to achieve specific goals. It can be determined according to scenario requirements, user groups, and communication purposes to ensure the accuracy and efficiency of information transmission.
[0173] For example, the generation system can include a processing interface (or referred to as a fixed speech pattern interface, fixed speech pattern template, etc.). After the Q&A pairs are generated, the generation system can integrate the generated Q&A pairs with a pre-defined fixed speech pattern interface in a modular manner to form the final Q&A pairs.
[0174] In this embodiment, by combining the generated Q&A pairs with the fixed speech pattern interface to ensure that the Q&A pairs meet the corresponding requirements of the corresponding scenarios, the user experience and the accurate transmission of information can be improved.
[0175] In some embodiments, after the Q&A pairs are generated, in combination with Figure 4It can be seen that the generation system can also rewrite and expand the question-answer pairs (which can be simply referred to as rewriting or expanding). Rewriting and expanding can be understood as expanding the coverage of question-answer pairs through diverse expressions without changing the original meaning of the question and ensuring that the question can be answered normally. Among them, the generation system can rewrite and expand the question-answer pairs based on a 7B-scale model (specifically, it can be the Qwen2.5-7B model). And the Qwen2.5-7B model can be a fine-tuned model. For example, fine-tune the Qwen2.5-7B model based on task requirements in actual scenarios, etc., and rewrite and expand the question-answer pairs based on the fine-tuned Qwen2.5-7B model.
[0176] Exemplarily, rewriting and expanding can include: sentence pattern transformation, synonym replacement, and user-feedback-driven rewriting.
[0177] Among them, sentence pattern transformation can be understood as transforming the sentence pattern of the question. For example, changing a declarative sentence to an interrogative sentence, or changing the active voice to the passive voice, etc.
[0178] Synonym replacement can be understood as generating question-answer pairs with similar semantics but different expressions by replacing some words in the question or answer.
[0179] User-feedback-driven rewriting can be understood as analyzing the actual search records and feedback of users, summarizing the common expressions used by users, and rewriting based on this. This method can ensure that the question-answer pairs are closer to the actual needs of users.
[0180] In some embodiments, combined with Figure 4 and Figure 5 It can be seen that after automatically generating question-answer pairs, the generation system can also control the quality of the question-answer pairs. On this basis, the generation system can build a FAQ library based on question-answer pairs with relatively high quality.
[0181] Among them, the FAQ library can be understood as a form of help resource in the form of a document or web page, usually used in websites, software applications, or other products, with the purpose of providing users with quick answers to the most common questions about the product or service. By setting up the FAQ library, organizations or individuals can effectively reduce the workload of customer service and at the same time make it more convenient and faster for users to find the information they need. The FAQ library generally includes various questions that users may encounter, such as usage guides, troubleshooting methods, situation descriptions, etc., and their answers.
[0182] As can be seen from the above analysis, the generation method provided in this specification can automatically generate question-and-answer pairs. Therefore, the generation system can further automatically generate an FAQ library on this basis, significantly shortening the construction cycle of the FAQ library. In addition, by generating through multiple data sources and combining the large language model generation method and the template generation method to automatically generate question-and-answer pairs, the diversity and comprehensive coverage of the FAQ library can be achieved.
[0183] Exemplarily, the quality control of the question-and-answer pairs by the generation system may include the following steps 1 to 3:
[0184] Step 1: Determine the measured quality score and confidence level of the question-and-answer pair respectively, where the confidence level is used to characterize the consistency between the measured quality score and the labeled quality score.
[0185] Continuing to combine the above example and Figure 4 , the quality control of the question-and-answer pair may include a confidence level measurement strategy and a quality standard system. Among them, the confidence level is determined by the generation system based on the confidence level measurement strategy, and the measured quality score is determined by the generation system based on the quality standard system.
[0186] Regarding the method for determining the measured quality score, reference can be made to the description of the generation system's quality scoring of the question-and-answer pairs within any question cluster in the above example, which will not be elaborated here.
[0187] The labeled quality score can be understood as the quality score obtained by manually or by other means (such as a network model, etc.) to label the question-and-answer pair.
[0188] Correspondingly, this step can be understood as the generation system determining the measured quality score of the question-and-answer pair and, on this basis, determining the similarity degree between the measured quality score and the labeled quality score. Relatively speaking, the higher the consistency between the measured quality score and the labeled quality score, the relatively higher the quality of the question-and-answer pair.
[0189] This embodiment does not limit the manner in which the generation system determines the confidence level. Exemplarily, the generation system can determine the confidence level based on a large model of prompting, or can also determine the confidence level based on methods such as the likelihood-based method, the sampling-based method, the training-based method, etc.
[0190] Among them, the likelihood-based method can be understood as measuring the confidence of the model in the given input by calculating the likelihood value output by the model. In large models, likelihood is usually used to evaluate the fitting degree of the model to the input data. If the model can well explain the input data, the likelihood value will be higher; conversely, if the model's explanation of the input data is poor, the likelihood value will be lower.
[0191] The sampling-based method can be understood as a technique for evaluating the confidence of large models by sampling from the model multiple times. This method is particularly suitable for dealing with uncertainty problems. Especially in large models (such as deep neural networks, generative models, etc.), the probability distribution output by the model may be very complex and difficult to directly analyze. By sampling multiple times, multiple prediction results can be generated, thereby quantifying the confidence of the model.
[0192] The training-based method can be understood as, based on the prompt-based method, collecting and processing training data and performing targeted training and optimization on the general large model to obtain better results. The disadvantage is that it requires a large amount of labeled data and will reduce the generalization of the solution.
[0193] The generation system determines the confidence and the measured quality score through the above methods to achieve the quality evaluation of the question-and-answer pair, and can realize the automated evaluation (Automated Evaluation) of the question-and-answer pair. And it can quickly and efficiently measure the quality, accuracy, fluency and other performances of the question-and-answer pair.
[0194] Step 2: If the confidence reaches the preset confidence threshold, build a FAQ library based on the question-and-answer pair.
[0195] Step 3: If the confidence is less than the preset confidence threshold, obtain the manual detection result of the question-and-answer pair, and when the manual detection result indicates that the question-and-answer pair meets the quality requirements, build a FAQ library based on the question-and-answer pair.
[0196] This embodiment does not limit the preset confidence threshold, which can be determined by the generation system based on requirements, historical records, experiments, etc.
[0197] Correspondingly, the generation system can compare the confidence with the preset confidence threshold to determine the size relationship between the confidence and the preset confidence threshold.
[0198] Exemplarily, continuing to combine the above example and Figure 5 It can be seen that the generation system can judge the size relationship between the confidence and the preset confidence threshold. If the confidence is greater than or equal to the preset confidence threshold, execute Step 2; conversely, if the confidence is less than the preset confidence threshold, execute Step 3.
[0199] For example, combining Figure 5As shown, if the confidence level is greater than or equal to the preset confidence threshold, it indicates that the quality of the question-and-answer pair is relatively high, that is, the judgment result is high confidence. Correspondingly, the generation system can build an FAQ library based on the question-and-answer pairs with relatively high quality.
[0200] For another example, combined with Figure 5 As shown, if the confidence level is less than the preset confidence threshold, it indicates that the quality of the question-and-answer pair may be relatively low, that is, the judgment result is medium confidence. To further determine whether the quality of the question-and-answer pair meets the quality requirements. The generation system can be determined by combining manual detection. For example, the generation system can output the question-and-answer pair for manual detection of the quality of the question-and-answer pair.
[0201] If the manual detection result of the manual feedback received is that the quality of the question-and-answer pair meets the quality requirements, as Figure 5 shown, the manual detection result is that it meets the quality requirements, then the generation system determines the question-and-answer pair that meets the quality requirements as a high-quality question-and-answer pair, and builds an FAQ library on this basis.
[0202] On the contrary, if the manual detection result of the manual feedback received is that the quality of the question-and-answer pair does not meet the quality requirements, the generation system can discard the question-and-answer pair that does not meet the quality requirements to avoid including question-and-answer pairs with low quality in the FAQ. Thus, the quality reliability of the FAQ library is improved. Furthermore, when the user is performing actual applications such as retrieval and query based on the FAQ library, the accuracy and reliability of the actual application are improved.
[0203] Correspondingly, since the quality of the FAQ library is relatively high, when the FAQ is applied to scenarios such as query and retrieval, the FAQ library can support quickly and accurately answering users' questions, improving users' satisfaction and experience.
[0204] Combined with the above analysis, it can be seen that in this specification, the question-and-answer pair can at least include three dimensions of content: question, basis, and answer. Combined with the above analysis, if the generation system performs an actual quality score on the question-and-answer pair, the question-and-answer pair can also include the actual quality score. That is, the question-and-answer pair includes a score (Quality_score) for the quality of the question-and-answer pair.
[0205] In addition, the question-and-answer pair can also include the ground truth, source text, query type, confidence level, etc.
[0206] It is worth noting that knowledge data may be updated or replaced. When knowledge data is updated or replaced, the synchronous update of question-answer pairs is particularly important. In actual application scenarios, the content of knowledge data may change at any time, which requires question-answer pairs to respond to and synchronize these changes in a timely manner. Since the source of information (i.e., text source) and basis are clearly marked in the question-answer pairs, when these basis change, the generation system can quickly locate the relevant content and update it synchronously to ensure the accuracy and timeliness of the question-answer pairs.
[0207] In some embodiments, the generation system may also update the question-answer pairs based on actual application of the question-answer pairs.
[0208] Exemplarily, the generation system may obtain the user's reflow data for the question-answer pair, generate an optimized prompt corresponding to the reflow data, and optimize the question-answer pair according to the optimized prompt.
[0209] Among them, the reflux data can be understood as the feedback information of users on the question and answer pairs in actual application scenarios. For example, if the quality of the question and answer pairs is not high, such as a high error rate, the generation system can generate a corresponding prompt (optimized prompt) based on the feedback information, and optimize the question and answer pairs based on the optimized prompt to improve the quality of the question and answer pairs.
[0210] Continuing with the above example and Figure 5 When the user determines the result corresponding to the query from the question-answer pair in the FAQ library based on the retrieval enhancement generation, feedback can be provided. For example, the user feedback shows that the accuracy of the result corresponding to the query is not high. Accordingly, the generation system obtains the feedback, that is, obtains the reflow data. The generation system can analyze the reflow data to update the controllable generation strategy (such as optimizing prompt) based on the analysis results of the reflow data. So that the generation system generates the corresponding question-answer pair based on the updated controllable generation strategy.
[0211] In this embodiment, the generation system can further improve the quality of the question-answer pairs by optimizing the question-answer pairs through the reflow data, thereby improving the accuracy and reliability of the optimized question-answer pairs in actual application scenarios, thereby improving the user experience.
[0212] It should be noted that the above examples are only used to exemplarily illustrate the possible implementation manners of the generation method of this specification, and should not be construed as a limitation on the implementation manners of the generation method of this specification. Exemplarily, on the basis of the above technical concept, some of the above technical features can be combined to obtain a new embodiment; new technical features can also be added on the basis of the above examples to obtain a new embodiment; some technical features can also be reduced on the basis of the above examples to obtain a new embodiment; some of the technical features in the above examples can also be replaced with other technical features; some of the technical features and orders in the above examples can also be adjusted to obtain a new embodiment, and so on, which will not be listed one by one here.
[0213] Based on the above technical concept, this specification also provides a computer-readable non-transitory storage medium, in which at least one instruction set is stored. When the at least one instruction set is executed by a processor, the steps of the generation method described in this specification are implemented.
[0214] In some possible embodiments, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on the generation system 200, the program code is used to cause the generation system 200 to execute the steps of the generation method described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) that includes program code and can run on the generation system 200. However, the program product of this specification is not limited to this. In this specification, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the generation system 200, partially on the generation system 200, executed as an independent software package, partially on the generation system 200 and partially on a remote generation system, or entirely on a remote generation system 200.
[0215] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0216] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to embrace various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0217] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0218] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well mark out some of the devices as separate embodiments when reading this specification. That is to say, the embodiments in this specification can also be understood as an integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0219] Every patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, references, etc. (excluding any historical prosecution files associated therewith), are hereby incorporated by reference for all purposes relevant hereto, e.g., in the specification and claims of this application. However, in the event of any inconsistency or conflict between the descriptions, definitions, and / or terms of the above materials and those used in this application, the descriptions, definitions, and / or terms used in this application shall prevail.
[0220] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely exemplary and not limiting. Those skilled in the art may implement the application in this specification by adopting alternative configurations based on the embodiments in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for generating question-and-answer pairs, comprising: Constructing a knowledge base based on the obtained knowledge data, wherein the knowledge base includes a text data set at the document level and a text block data set at the unit level; Constructing knowledge graphs corresponding to the text data set and the text block data set respectively, wherein the knowledge graph is used to represent the relevance between different documents in the data set, and the data set includes the text data set and the text block data set; and Sampling the documents in the text data set and the text block data set respectively to obtain sampled documents, and generating question-and-answer pairs based on the sampled documents and the knowledge graphs corresponding to the sampled documents.
2. The method according to claim 1, wherein, The generating of the question-and-answer pairs based on the sampled documents and the knowledge graphs corresponding to the sampled documents includes: Determining a target generation method from preset generation methods according to the data type of the sampled documents, wherein the preset generation methods include a large language model generation method and a template generation method; and Using the target generation method to generate the question-and-answer pairs based on the sampled documents and the knowledge graphs corresponding to the sampled documents.
3. The method according to claim 2, wherein, The target generation method is a large language model generation method; using the target generation method to generate the question-and-answer pairs based on the sampled documents and the knowledge graphs corresponding to the sampled documents includes: Obtaining relevant documents corresponding to the sampled documents from the knowledge graph corresponding to the sampled documents; Inputting the sampled documents, the relevant documents, and a preset prompt into a large language model to obtain the question-and-answer pairs, wherein the question-and-answer pairs include questions, answers, and corresponding bases, and the preset prompt is used to guide the large language model to generate the question-and-answer pairs based on the sampled documents and the relevant documents.
4. The method according to claim 3, wherein, The preset prompt is further used to guide at least one of the style, wording, and quantity of the question-and-answer pairs.
5. The method according to any one of claims 1 to 4, wherein The quantity of the question-and-answer pairs is multiple; the method further includes: Performing similar question normalization processing on multiple question-and-answer pairs to obtain question-and-answer pairs corresponding to each type of question, so that the same type of questions correspond to one answer.
6. The method according to claim 5, wherein, A question-and-answer pair includes a question and an answer; the performing of similar question normalization processing on multiple question-and-answer pairs to obtain question-and-answer pairs corresponding to each type of question includes: Clustering similar questions into a question cluster; For any question cluster, performing quality scoring on the question-and-answer pairs within the any question cluster; and Determining the question-and-answer pair with the highest quality score as the question-and-answer pair corresponding to the any question cluster.
7. The method according to claim 6, wherein, The clustering of similar questions into a question cluster includes: For any new question, converting the new question into a new question vector; If there is a target question cluster in the existing question clusters that includes a similar question vector corresponding to the new question vector, merging the new question into the target question cluster; and If there is no question cluster in the existing question clusters that includes a similar question vector corresponding to the new question vector, creating a question cluster corresponding to the new question.
8. The method according to any one of claims 1 to 4, wherein The method further includes: Determining respectively a measured quality score and a confidence level of the question-answer pair, wherein the confidence level is used to characterize the consistency between the measured quality score and the annotated quality score; If the confidence reaches a preset confidence threshold, constructing a FAQ library based on the question and answer pairs; and If the confidence is less than the preset confidence threshold, the manual detection result of the user on the question and answer pair is obtained, and when the manual detection result indicates that the question and answer pair meets the quality requirements, the FAQ library is constructed based on the question and answer pair.
9. The method according to any one of claims 1 to 4, wherein The method further comprises: Obtain standard scripts corresponding to actual application scenarios; and The standard speech is embedded in the question-answer pair.
10. The method according to any one of claims 1 to 4, wherein, The method further comprises: Obtaining user return data for the question-answer pair; generating an optimization prompt corresponding to the reflow data; and The question-answer pair is optimized according to the optimized prompt.
11. The method according to any one of claims 1 to 4, wherein The knowledge data includes data related to a target field collected from multiple channels, and the target field is a field corresponding to the question-answer pair; The knowledge data includes: structured data, semi-structured data, and unstructured data.
12. The method according to any one of claims 1 to 4, wherein, The step of constructing a knowledge base according to the acquired knowledge data comprises: Parsing the knowledge data to obtain data in a preset text format; Converting the data in the preset text format into data in Markdown format; and The knowledge base is constructed based on the data in the Markdown format.
13. A system for generating question-answer pairs, comprising: at least one storage medium storing at least one set of instructions for generating a question-answer pair; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, the at least one instruction set is read, and the method as described in any one of claims 1 to 12 is executed according to the instructions of the at least one instruction set.
Citation Information
Patent Citations
Man-machine conversation and pre-training language model training method and system and electronic equipment
CN115587175A
Document question and answer method, device and system, electronic equipment and storage medium
CN115934905A
Intelligent questioning and answering method for government affairs
CN118153686A
Dynamic correlation enhancement retrieval generation system and method driven by intelligent knowledge graph
CN118839021A
Multi-level domain knowledge question-answering method and device based on large model
CN119202213A
Cited By
Graph-driven question and answer generation method and device based on structure perception, medium and equipment
CN120470098A
Graph-driven question-answer generation method, device, medium, and equipment based on structure perception
CN120470098B
Question and answer pair generation method and system, computer equipment and readable storage medium
CN121434347A