Data set generation method and system
By guiding the generation model to generate Q&A samples containing question information, evidence information and answer information, and combining multiple strategies to optimize the data set, the problem of insufficient quality and diversity of existing Q&A data sets is solved, and efficient and low-cost data set generation and application scenario adaptability are achieved.
Patent Information
- Application Number
- CN202510414813.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing Q&A dataset generation methods have low quality and insufficient diversity, which are difficult to meet the complex and changing needs of enterprise-level application scenarios, and manual labeling is inefficient and costly, making it difficult to quickly expand and iterate.
Through the guided generation model, it generates Q&A samples based on the knowledge base, including question information, evidence information and answer information. It adopts a combination strategy of multiple question types and question-answer styles, combines knowledge graphs and sampling strategies to generate high-quality Q&A samples, and optimizes the dataset through derived processing, repeated culling and low-quality culling.
The generated Q&A data set is of high quality and rich diversity, which can better cover practical application scenarios, improve the accuracy and efficiency of model training and evaluation, reduce manual labeling costs, and adapt to the needs of rapid iteration and expansion.
Smart Images

Figure CN120258145A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method and system for generating a dataset. Background Art
[0002] In the context of the rapid development of artificial intelligence technology, the training and deployment cycles of machine learning models have been significantly shortened. Among them, the speed and effectiveness of training and deployment of machine learning models both depend on the construction of a question-and-answer dataset.
[0003] Currently, the construction of question-and-answer datasets mainly relies on manual annotation. Manual annotation refers to the process of annotating data manually to create a question-and-answer dataset. Since the evaluation sets of manual annotation are generally generated based on the knowledge and judgment of human experts, they usually have high quality and reliability. However, manual annotation requires a large amount of human resources and time, especially when dealing with large-scale datasets, the efficiency is low. The results of manual annotation are also affected by the subjective judgment of the annotators. Different people may have different understandings of the same data, resulting in inconsistencies in the results. In addition, the scalability of manual annotation is poor. Especially in scenarios where large-scale data needs to be processed quickly or updated frequently, manual annotation often struggles to keep up with the changing demands. Manual annotation is also limited by domain dependence. Some tasks require annotators to have specific professional knowledge, increasing the difficulty and cost of annotation. Moreover, as the evaluation criteria and requirements change, the manually annotated datasets need to be re-annotated, increasing the maintenance cost. Therefore, manual annotation is not only time-consuming and laborious, subjective, but also difficult to scale, restricting the rapid iteration and deployment of large-scale systems.
[0004] With the development of technology, methods for automatically generating question-and-answer data have gradually attracted attention. By automatically generating questions and answers through algorithms, the efficiency of constructing evaluation sets can be improved. However, the current question-and-answer dataset generation solutions have obvious deficiencies. Most of the existing generated question-and-answer data are automatically generated based on large models. Due to the inevitable hallucination problem, the quality of the question-and-answer datasets generated by large models is not high. And it is also impossible to guarantee the diversity and representativeness of the question-and-answer datasets. Especially in the business-to-business (To Business, ToB) scenario, the requirements are complex and changeable. A single question-and-answer dataset is difficult to cover all possible scenarios, resulting in insufficient generalization ability in training or evaluation.
[0005] The content in the background art section is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the filing date of this disclosure, nor does it mean that it can become the prior art of this disclosure. Summary of the Invention
[0006] This specification provides a method and system for generating a dataset. By guiding a generation model to generate question-and-answer samples based on a knowledge base and requiring the generation of corresponding evidence information, the question information and answer information are made more coherent and logical, avoiding the hallucination problem and automatically generating high-quality question-and-answer samples.
[0007] In a first aspect, this specification provides a method for generating a dataset, including: obtaining a knowledge base; guiding a generation model to generate a plurality of question-and-answer samples based on the knowledge base, where each question-and-answer sample includes: question information, evidence information, and answer information, the question information is related to the content in the knowledge base, the evidence information is the information relied on by the generation model in the process of answering the question information, and the answer information is the information generated by the generation model for answering the question information; and generating the dataset based on the plurality of question-and-answer samples.
[0008] In some embodiments, guiding the generation model to generate a plurality of question-and-answer samples based on the knowledge base includes: obtaining a plurality of question types; and based on the plurality of question types, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base.
[0009] In some embodiments, based on the plurality of question types, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base includes: obtaining a plurality of question-and-answer styles, each question-and-answer style characterizing at least one of the style corresponding to the question information or the style corresponding to the answer information; and based on the plurality of question types and the plurality of question-and-answer styles, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base.
[0010] In some embodiments, based on the plurality of question types and the plurality of question-and-answer styles, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base includes: combining the plurality of question types and the plurality of question-and-answer styles to generate a plurality of guiding strategies, each guiding strategy characterizing a question type and a question-and-answer style; traversing the plurality of guiding strategies, and for each current guiding strategy during the traversal: generating a first guiding instruction based on the current guiding strategy and inputting the first guiding instruction into the generation model to guide the generation model to generate at least one question-and-answer sample corresponding to the current guiding strategy.
[0011] In some embodiments, generating the first guiding instruction based on the current guiding strategy includes: sampling a part of the content from the knowledge base as reference knowledge based on the current guiding strategy; and generating the first guiding instruction based on the current guiding strategy and the reference knowledge.
[0012] In some embodiments, the knowledge base includes at least one document, and each document is divided into multiple document blocks. Sampling at least one text knowledge from the knowledge base based on the current guiding strategy includes: if the question type represented by the current guiding strategy is a single-hop question, sampling one document or one document block from the knowledge base as the reference knowledge; if the question type represented by the current guiding strategy is a multi-hop question, sampling multiple documents or multiple document blocks from the knowledge base as the reference knowledge.
[0013] In some embodiments, sampling multiple documents from the knowledge base as the reference knowledge includes: sampling a first document from the knowledge base, sampling at least one second document associated with the first document from the knowledge base based on a first knowledge graph, and using the first document and the at least one second document as the reference knowledge, where the first knowledge graph represents the association relationship between different documents; sampling multiple document blocks from the knowledge base as the reference knowledge includes: sampling a first document block from the knowledge base, sampling at least one second document block associated with the first document block from the knowledge base based on a second knowledge graph, and using the first document block and the at least one second document block as the reference knowledge, where the second knowledge graph represents the association relationship between different document blocks.
[0014] In some embodiments, guiding the generation model to generate the multiple Q&A samples based on the knowledge base based on the multiple question types and the multiple Q&A styles includes: generating a second guiding instruction based on the multiple question types and the multiple Q&A styles, where the second guiding instruction is used to guide the generation model to generate at least one Q&A sample based on the knowledge base for different combinations of the multiple question types and the multiple Q&A styles respectively; inputting the second guiding instruction into the generation model to obtain the multiple Q&A samples.
[0015] In some embodiments, generating the second guiding instruction based on the multiple question types and the multiple Q&A styles includes: generating the second guiding instruction based on the multiple question types, the multiple Q&A styles, and a preset sampling strategy, where the second guiding instruction is used to guide the generation model to: generate multiple combinations based on the multiple question types and the multiple Q&A styles, for each combination, sample partial content from the knowledge base as the reference knowledge according to the sampling strategy, and sample at least one text knowledge from the reference knowledge.
[0016] In some embodiments, the multiple question types include M first question types, where M is an integer greater than 1, and the M first question types are obtained by the following method: obtaining a plurality of classification dimensions, and determining a plurality of candidate types corresponding to each classification dimension, where the plurality of classification dimensions include at least two of: the purpose of the question, the number of documents involved in the question, or the question type; combining the candidate types under different classification dimensions to obtain the M first question types.
[0017] In some embodiments, the data set is used to evaluate the performance of the question-answering system, and the multiple question types further include N second question types, where N is an integer greater than or equal to 1; the N second question types are question types determined based on the feedback data of the question-answering system.
[0018] In some embodiments, generating the data set based on the plurality of question-answer samples includes: performing derivation processing on at least some of the plurality of question-answer samples to obtain a plurality of derived question-answer samples, and generating an intermediate data set based on the plurality of question-answer samples and the plurality of derived question-answer samples; and removing the question-answer samples in the intermediate data set that do not meet the expectations to obtain the data set.
[0019] In some embodiments, for a target question-answer sample among the plurality of question-answer samples, the process of performing derivation processing on the target question-answer sample includes: rewriting the question information in the target question-answer sample to obtain derived question information; and generating the derived question-answer sample based on the derived question information and the answer information in the target question-answer sample.
[0020] In some embodiments, removing the question-answer samples in the intermediate data set that do not meet the expectations includes: clustering the question-answer samples in the intermediate data set based on the semantics of each question-answer sample in the intermediate data set to obtain at least two subsets; and removing some samples in each subset.
[0021] In some embodiments, removing the question-answer samples in the intermediate data set that do not meet the expectations includes: providing each question-answer sample in the intermediate data set to a scoring model, and obtaining the quality score output by the scoring model; and removing the question-answer samples in the intermediate data set whose quality scores do not meet the expectations.
[0022] In some embodiments, for the target Q&A samples in the intermediate dataset, the step of providing the target Q&A samples to a scoring model and obtaining the quality scores output by the scoring model includes: obtaining a preset quality scoring criterion, where the quality scoring criterion includes scoring description information corresponding to multiple scoring dimensions; generating a guiding instruction based on the quality scoring criterion and the Q&A samples, where the guiding instruction is used to guide the scoring model to score the target Q&A samples according to the quality scoring criterion; inputting the guiding instruction into the scoring model, and obtaining the quality scores output by the scoring model.
[0023] In some embodiments, the knowledge base is the knowledge base used by the Q&A system, and the dataset is used to evaluate the performance of the Q&A system.
[0024] In a second aspect, this specification also provides a dataset generation system, including: at least one storage medium storing at least one instruction set for generating a dataset; and at least one processor communicatively connected to the at least one storage medium, where when the dataset generation system runs, the at least one processor reads the at least one instruction set and executes the dataset generation method according to any one of the first aspect as instructed by the at least one instruction set.
[0025] As can be seen from the above technical solutions, in the method and system for generating a dataset provided in this specification, a knowledge base is obtained; a generation model is guided to generate multiple Q&A samples based on the knowledge base, where each Q&A sample includes: question information, evidence information, and answer information, the question information is related to the content in the knowledge base, the evidence information is the information relied on by the generation model in the process of answering the question information, and the answer information is the information generated by the generation model for answering the question information; and the dataset is generated based on the multiple Q&A samples. In the embodiments of this specification, by guiding the generation model to generate Q&A samples based on the knowledge base and requiring the generation of corresponding evidence information, the question information and the answer information are made more coherent and logical, avoiding the hallucination problem and automatically generating high-quality Q&A samples.
[0026] Other functions of the method and system for generating a dataset provided in this specification will be partially listed in the following description. The creative aspects of the method and system for generating a dataset provided in this specification can be fully explained through practice or the use of the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0028] Figure 1 Fig. shows a schematic diagram of an application scenario for generating a data set provided according to an embodiment of this specification;
[0029] Figure 2 Fig. shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0030] Figure 3 Fig. shows a flowchart of a method for generating a data set provided according to an embodiment of this specification; and
[0031] Figure 4 Fig. shows a schematic diagram of a process for generating a data set provided according to an embodiment of this specification. Detailed implementation manners
[0032] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various local modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of this specification, the general principles defined here can be applied to other embodiments and applications. Therefore, this specification is not limited to the shown embodiments, but is the broadest scope consistent with the claims.
[0033] The terms used here are only for the purpose of describing specific example embodiments and are not restrictive. For example, unless otherwise clearly stated in the context, the singular forms "a", "an" and "the" used here can also include the plural forms. When used in this specification, the terms "comprising", "including" and / or "containing" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups, or the addition of other features, integers, steps, operations, elements, components and / or groups in the system / method.
[0034] Considering the following description, these features of this specification and other features, as well as the operations and functions of the related elements of the structure, and the combination and manufacturing economy of the components can be significantly improved. Referring to the accompanying drawings, all of these form a part of this specification. However, it should be clearly understood that the accompanying drawings are only for the purpose of illustration and description and are not intended to limit the scope of this specification. It should also be understood that the accompanying drawings are not drawn to scale.
[0035] The flowcharts used in this specification illustrate the operations implemented by a system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0036] For ease of description, the terms that will appear later in this specification are first explained.
[0037] Retrieval Augmented Generation (RAG): Retrieval Augmented Generation is a technique that combines retrieval and generation, aiming to enhance the generation ability of large language models by dynamically retrieving relevant information. Such methods usually rely on external knowledge bases or document collections. When processing natural language tasks, they can first retrieve the context or facts related to the input from the knowledge base, and then use the large language model to generate accurate answers or texts based on the retrieval results.
[0038] In this specification, the large language model (Large Language Model, LLM) can also be abbreviated as the large model. A large language model is a natural language processing model based on deep learning technology, with the number of parameters usually reaching billions to hundreds of billions or even higher, and having powerful language understanding and generation capabilities. The large language model can adopt the Transformer architecture or its variants (such as GPT, BERT, etc.). This architecture uses the attention mechanism to globally model sequence data, can efficiently handle long-distance dependencies, and thus performs well in natural language tasks. The large language model learns the statistical features and semantic correlations of language by pre-training on a large-scale corpus, enabling it to have excellent generalization ability. The core capabilities of the large language model include but are not limited to: understanding context semantics, generating coherent and grammatically correct texts, performing logical reasoning, and handling multi-task scenarios. Its usage methods usually include two modes: direct inference and fine-tuning. In the direct inference mode, the user guides the large language model to generate specific outputs by designing prompt words or guiding instructions (Prompt). The prompt words can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation ability of the large language model. In the fine-tuning mode, the large language model is further trained on a small-scale dataset in a specific domain to optimize its performance on specific tasks. The powerful generalization ability and flexibility of the large language model make it an important tool in the field of artificial intelligence technology, providing an efficient and accurate solution for automated text generation and understanding.
[0039] In some embodiments, large language models can also have the ability to understand and generate data of other modalities (such as vision, audio, etc.). In this case, large language models can also be referred to as Multimodal Large Language Models (MLLMs). By integrating various types of inputs and outputs such as text, images, and sounds, MLLMs provide a richer and more natural interaction experience. The core advantage of MLLMs lies in their ability to process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze an image and generate descriptive text, or generate a corresponding image based on a text description. This cross-modal understanding and generation ability makes MLLMs have broad application prospects in multiple fields.
[0040] It should be noted that the key technologies of large language models can be referred to the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, publication time: March 11, 2025, publication link: https: / / doi.org / 10.48550 / arXiv.2303.18223), and this specification will not elaborate here.
[0041] The application scenarios of this specification will be introduced below.
[0042] The solution provided in this specification can be used to generate a dataset containing multiple question-and-answer samples. Each question-and-answer sample includes question information, evidence information, and answer information. This dataset can be applied to all stages of the life cycle of a model or system. For example, in the training stage, the dataset generated based on the generation method provided in this specification can be used to train a model or system; in the evaluation and inference stage, the dataset can be used to comprehensively evaluate the performance of a model or system.
[0043] In some embodiments, the generated dataset can be used as a training set for training a model or system. The question-and-answer samples generated by the method in the embodiments of this specification can be used as labeled training samples, with the question information as the input and the answer information as the correct labeled result to help the model or system conduct effective training. In addition, based on the extensive content in the knowledge base, the generation model can generate question-and-answer samples covering different fields, thereby ensuring the diversity and representativeness of the dataset and improving the training effect of the model or system. The generated question-and-answer data in this specification also includes evidence information, and the answers generated based on the evidence information are more reasonable and logical, providing richer and more complex training data for the model.
[0044] In some embodiments, the generated dataset can also be used as an evaluation set to evaluate the performance of a model or system. An evaluation set typically consists of a series of questions and their corresponding correct answers. The Q&A samples generated according to the embodiments of this specification can be used as labeled evaluation samples, with the question information as the input and the answer information as the labeled correct result, to evaluate the performance of the model's output for the question information. It can be used to comprehensively evaluate the accuracy and capabilities of a model or system in different question types and scenarios. For example, the generated dataset can be used to measure the performance of a RAG system in a Q&A task.
[0045] It should be noted that the above application scenarios of generating the dataset are just some examples among multiple usage scenarios. The method for generating the dataset provided in this specification can be applied not only to the scenarios listed above but also to other scenarios that require a dataset. Those skilled in the art should understand that when the method for generating the dataset provided in this specification is applied to other usage scenarios, its implementation manner and technical effects are similar. In this embodiment, the type of Q&A samples in the dataset is not limited and can be various Q&A data types such as images, videos, and audios.
[0046] Figure 1 FIG. shows a schematic diagram of an application scenario of a method for generating a dataset according to an embodiment of this specification. As Figure 1 shown, the application scenario 100 may include a dataset generation system 110 (hereinafter referred to as the generation system 110).
[0047] The generation system 110 generates multiple Q&A samples by combining a generation model with the content in a knowledge base, and generates the dataset based on these multiple Q&A samples. Each Q&A sample includes: question information, evidence information, and answer information. The question information is related to the content in the knowledge base, the evidence information is the information relied on by the generation model in the process of answering the question information, and the answer information is the information generated by the generation model for answering the question information.
[0048] In some embodiments, the generation model can be a trained large model. In the generation system 110, the generation model can execute different generation strategies based on different guiding instructions to generate different Q&A samples.
[0049] In some embodiments, the generation model can be deployed inside the generation system 110 to directly perform calculations and processing inside the generation system 110. This deployment method can improve the response speed and processing efficiency, and can also better control the data flow and the optimization of the generation model, facilitating the generation system 110 to adjust and upgrade the generation model.
[0050] In some embodiments, the generation model can also be deployed outside the generation system 110. The generation system 110 generates Q&A samples by invoking the external generation model. In this case, the generation system 110 can send requests and guiding instructions to an external server and receive the generated Q&A samples. This deployment method can utilize cloud computing resources or third-party high-performance computing platforms, reduce the burden on the local hardware of the generation system 110, and facilitate integration with other systems. For example, the generation model can be provided as an API service for multiple different generation systems 110 to call. This deployment method provides greater flexibility and scalability for the generation system 110, and the performance and resource usage of the model can be adjusted according to actual needs.
[0051] In some embodiments, the generation system 110 can store data and instructions for implementing the data set generation method provided in this specification and can execute or be used to execute the data and instructions. In some embodiments, the generation system 110 can include a hardware device with data information processing capabilities and the necessary programs for driving the hardware device to work.
[0052] It should be noted that the generation system 110 can correspond to a single device or a device cluster, and this specification does not limit this. When the generation system 110 corresponds to a single device, the data set generation method can be fully executed on this device. When the generation system 110 corresponds to a device cluster, the data set generation method can be executed in cooperation on multiple devices corresponding to the device cluster, and this specification does not limit this.
[0053] It should be noted that all user data obtained in this specification has been authorized by the users and does not involve user privacy.
[0054] Figure 2 The hardware structure diagram of a computing system 200 provided according to an embodiment of this specification is shown. The computing system 200 can be used as Figure 1 the generation system 110 in and execute the data set generation method described in this specification.
[0055] As Figure 2 shown, the computing system 200 can include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 can also include a communication port 250 and an internal communication bus 210. The computing system 200 can also include I / O components 260.
[0056] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250, and the I / O components 260, etc.
[0057] The I / O component 260 supports input / output between the computing system 200 and other components.
[0058] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and the network 140. The communication port 250 can be a wired communication port or a wireless communication port.
[0059] The storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. The storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, etc.
[0060] At least one processor 220 may be communicatively coupled to at least one storage medium 230. When the computing system 200 is running, at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the method for generating the data set provided in this specification. The processor 220 may execute the steps included in the method for generating the data set. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field-programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0061] For purposes of illustration only, the computing system 200 in the drawings shows only one processor 220. However, it should be noted that the computing system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 220 of the computing system 200 executes step A and step B, it should be understood that step A and step B may also be executed jointly or separately by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0062] Figure 3 FIG. 4 shows a flowchart of a method P300 for generating a data set provided according to an embodiment of this specification. As before, the computing system 200 may execute the method P300 for generating the data set of this specification. Specifically, the processor 220 in the computing system 200 may read the instruction set stored in its local storage medium and then execute the method P300 for generating the data set of this specification according to the provisions of the instruction set. As Figure 3 shown, the method P300 may include steps S310-S330.
[0063] S310: Obtain a knowledge base.
[0064] In some embodiments, the content in the knowledge base may include a wide range of content, covering information in multiple topics or fields. The content of this knowledge base can be obtained through various channels, such as database retrieval or collecting structured and unstructured data from public resources. This diverse data source enables the generated Q&A samples to have higher diversity and be able to cover a wider range of field requirements.
[0065] In some embodiments, the content in the knowledge base may focus on content in a specific field or scope, providing relatively professional knowledge. For example, the knowledge base may contain detailed data in a certain professional field, such as professional knowledge in the fields of medicine, law, finance, technology, etc. Utilizing this professional knowledge, the generated Q&A samples can target specific field application scenarios, thereby better training or evaluating models and systems focused on that field. This customized knowledge base can improve the pertinence and effectiveness of the data set and also enhance the reasoning ability and performance of the model or system in that field.
[0066] In some embodiments, the knowledge base may also include the feedback data of the online Q&A system. The feedback data refers to the user interaction data collected by the Q&A system during actual application. The Q&A samples generated based on this feedback data can not only be used to evaluate the actual performance of the online system, but also help developers identify potential problems in the system and optimize the system accordingly.
[0067] As described above, the dataset generated in the embodiments of this specification can be used to train a model or system, and can also be used to evaluate the performance of a model or system. In some embodiments, in the scenario for evaluation, the knowledge base can be the knowledge base used by a question-and-answer system, and the dataset is used to evaluate the performance of the question-and-answer system. The question-and-answer data generated based on the knowledge base used by the question-and-answer system is more targeted and effective, and can ensure that the evaluation process more truly reflects the performance of the question-and-answer system in actual applications. The dataset generated based on the same knowledge base as that used by the question-and-answer system can enhance the reliability and repeatability of the evaluation process, making the evaluation results more stable and credible. Further, based on the above embodiments, developers can also conduct a comparative analysis on different versions of the question-and-answer system, so as to identify performance differences and perform refined optimization according to the feedback.
[0068] In some embodiments, after obtaining the original documents in the knowledge base, it is also necessary to preprocess the original documents in the knowledge base. In the embodiments of this specification, the original documents in the knowledge base are first parsed into a unified text format and converted into a markup language (Markdown) format, so as to facilitate the extraction of structured information such as titles, paragraphs, and sentences. These structured information not only provide a unified data format for subsequent links, but also lay a foundation for in-depth analysis and utilization of document content.
[0069] In some embodiments, there are two ways to process documents: the processing method based on the whole document and the processing method based on document chunks (Chunks). These two methods have their own advantages and disadvantages and are applicable to different scenarios and requirements. The processing method based on the whole document can comprehensively cover the content of the document, and the questions and answers generated based on this type of document have high integrity and consistency. It is applicable to question-and-answer samples that require a comprehensive understanding of the document content, such as academic papers and long reports. The question-and-answer samples generated by the processing method based on document chunks are relatively concise and have higher diversity. This type of document chunk is applicable to generating some detailed questions or when the long document exceeds the input limit of the LLM. In the embodiments of this specification, the chunking method of the document chunks of the RAG system can be reused, and its performance meets the expectations, and it can be aligned with the existing RAG system in the application scenario of evaluating the RAG system.
[0070] In some embodiments, in order to better express the relationship between different documents and document chunks. The computing system 200 can also create a knowledge graph in the knowledge base, which is divided into the document level and the document chunk level. The knowledge graph can clearly show the relevance between documents or document chunks. Therefore, the construction of the knowledge graph helps LLMs to more effectively generate multi-hop questions.
[0071] S320: Guide the generative model to generate multiple question and answer samples based on the knowledge base, wherein each question and answer sample includes: question information, evidence information and answer information, the question information is related to the content in the knowledge base, the evidence information is the information based on which the generative model answers the question information, and the answer information is the information generated by the generative model to answer the question information.
[0072] If question-answering samples are directly generated based on a large model, the quality of the generated question-answering dataset is not high enough due to the inevitable hallucination problem. In order to meet the requirements of rapid evaluation in multiple scenarios, this manual provides a solution for automatically generating question-answering samples based on Query-Evidence-Answer (QEA).
[0073] Specifically, in the process of generating question information (Query) and answer information (Answer), a key step is introduced in the implementation of this specification: between generating question information and answer information, the generation model is required to generate corresponding evidence information (Evidence). The generation process of the question and answer samples in this specification conforms to the "thinking chain" habit of LLM, and can also make the generated questions and answers more coherent and logical, thereby improving the overall quality of the question and answer samples. At the same time, the generation process is more in line with the "white box" design concept, and the evidence generation process provides a convenient way to check the quality of question and answer samples, which is convenient for automatic or manual inspection of the quality of generated content. By verifying the accuracy and relevance of the evidence, technicians can also more effectively evaluate and optimize the generated question and answer samples to ensure the reliability and practicality of their content.
[0074] The basic idea of automatic generation of question and answer samples provided in this manual inherits and innovates technologies such as self-instruction (Self-Instruct), self-alignment (Self-Align) and self-question and answer (Self-QA). It makes full use of the generation ability of LLM and can automatically generate high-quality questions and their corresponding answers from the text. The core is to create questions that are closely related to the text content in the knowledge base through the deep understanding and generation ability of LLM, and then generate accurate and detailed answers. This process not only significantly improves the quality of the data, but also greatly increases the diversity and coverage of question and answer pairs, making the generated question and answer samples richer and more comprehensive.
[0075] In order to ensure that the data set is highly diverse and has high coverage, a guiding strategy for question types is added during the generation of question and answer samples. In some embodiments, the computing system 200 obtains multiple question types; and based on the multiple question types, guides the generation model to generate the multiple question and answer samples based on the knowledge base.
[0076] Based on the above control of the question types, it can ensure that the generated questions are more diverse, making the question-and-answer data types in the dataset more comprehensive. Furthermore, it enables more comprehensive training based on this dataset during the training process and discovers more potential problems based on this dataset during the evaluation process. The question types can be divided into two categories: preset question types (the first question type) and real return data mining question types (the second question type).
[0077] In some implementations, the multiple question types include M first question types, where M is an integer greater than 1, and the M first question types are obtained in the following manner: Obtain multiple classification dimensions and determine multiple candidate types corresponding to each classification dimension, where the multiple classification dimensions include at least two of: the purpose of the question, the number of documents involved in the question, or the question type; Combine the candidate types under different classification dimensions to obtain the M first question types.
[0078] In some embodiments, the classification dimension of the purpose of the question may correspond to three candidate types. Retrieval questions: These questions mainly focus on finding specific data or facts from existing information. Summary questions: These questions require refining a large amount of information to summarize the core ideas or key points. Solution questions: These questions focus on seeking strategies or methods to solve specific problems or challenges.
[0079] In some embodiments, the classification dimension of the number of documents involved in the question may correspond to three candidate types. Single-hop questions: These questions can be answered by referring to only one document or document block. Multi-hop questions: These questions require spanning multiple documents or document blocks and performing continuous reasoning or searching to answer. Out-of-Distribution (OOD) questions: These questions refer to questions that are outside the preset document scope and require the knowledge of the LLM itself or external knowledge to answer.
[0080] In some embodiments, the classification dimension of the question type may correspond to two candidate types. True / False questions: These questions require the responder to judge the truth or falsehood of a statement. Short-answer questions: These questions require the responder to answer a specific question in concise language.
[0081] The candidate types under the above multiple different classification dimensions can be combined randomly with restrictions. By combining the candidate types under the above different classification dimensions, the M first question types can be obtained. For example, the first question type can be retrieval type; it can also be retrieval type + single-hop question; it can also be retrieval type + single-hop question + true / false question. By combining different candidate types, it can ensure the diversity and comprehensiveness of the question types to cover different evaluation scenarios and requirements.
[0082] In some implementations, the data set is used to evaluate the performance of the question-answering system. The multiple types of questions further include N second types of questions, where N is an integer greater than or equal to 1; the N second types of questions are question types determined based on the feedback data of the question-answering system.
[0083] In some embodiments, the question-answering system has the ability to output answer information based on the question information. The question-answering system can be a system based on the RAG architecture, or can also be a system based on a neural network or an LLM model architecture. The data set generated in the embodiments of this specification can be used to evaluate the performance of the question-answering system. It is worth noting that the question-answering system mentioned in the embodiments of this specification is not limited to a hardware system or a software system, and can also refer to a model with the ability to output answer information based on the question information.
[0084] To further optimize the data set and improve the performance of the question-answering system, this specification can also identify common question types by analyzing the feedback data, and accordingly generate question-and-answer samples specifically. The feedback data refers to the real user interaction data obtained from the online system, which includes various types of questions actually asked by users. Through the analysis of the feedback data, the computing system 200 can identify various question types frequently asked by users or the question types of wrong examples generated by the question-answering system, so as to generate specific second types of questions. The above embodiments ensure that the question-and-answer samples can closely fit the actual application scenario, reflect the real problems encountered by users in daily use, and thus improve the actual performance of the question-answering system.
[0085] In practical applications, the analysis of the real feedback data in this specification is divided into two aspects: active summary and passive feedback. The second types of questions can be summarized from these two aspects.
[0086] Active summary: The types of knowledge involved in each application scenario may be different, so specific question types closely related to the scenario may appear in different scenarios. Therefore, by performing clustering analysis on the feedback data in the specified application scenario, the common question types in this scenario can be summarized. These question types can be used as the N second types of questions. The method of active summary not only helps to deeply understand the needs of users in a specific scenario, but also can generate more representative question types, thus significantly improving the effectiveness of the data set and the accuracy of the evaluation.
[0087] Passive feedback: It is also possible to conduct in-depth analysis on the bad cases generated by the question-and-answer system, summarize the types of questions therein, and then identify the deficiencies of the system. These types of questions can also be used as the N second types of questions. By analyzing these types of questions, it is possible to evaluate the question-and-answer system in a targeted manner to ensure that it can better handle various complex and changing scenarios. In addition, by expanding this type of bad-case question-and-answer dataset, more potential user question scenarios can be covered, further enhancing the system's response capabilities in the face of various questions and ensuring its comprehensiveness and accuracy.
[0088] The second types of questions based on real user needs can effectively evaluate the actual performance of the question-and-answer system, especially the accuracy and reliability of the question-and-answer system in a real application environment. And the second types of questions can also more accurately simulate the actual question-asking methods and needs of users, not only improving the authenticity of the evaluation results but also ensuring that the system can effectively respond to diverse user questions in a real environment.
[0089] The above two types of questions (the first type of questions and the second type of questions) can be used alone or in combination in the same application scenario. For example, the dataset can contain only question-and-answer samples generated based on the first type of questions or the second type of questions, or can contain question-and-answer samples generated based on both types of questions at the same time. Through this flexible combination, the dataset can cover a wider range of evaluation scenarios. Those skilled in the art can flexibly combine these two types of questions according to the requirements of different application scenarios. The dataset generated in this way can not only more comprehensively evaluate the performance of the question-and-answer system but also ensure that the system can effectively adapt to diverse scenarios and tasks.
[0090] In some embodiments, the computing system 200 can also obtain multiple question-and-answer styles, each question-and-answer style characterizing at least one of the style corresponding to the question information or the style corresponding to the answer information; and based on the multiple types of questions and the multiple question-and-answer styles, guide the generation model to generate the multiple question-and-answer samples based on the knowledge base.
[0091] According to different actual application scenarios, the computing system 200 can also generate questions and answers that conform to a specific style. Common question-and-answer styles can include the following categories: question length, question style, and answer elaboration level.
[0092] Question length: The number of words in the question can be divided into different categories according to requirements. For example, the question length can be less than 25 words, 25 to 50 words, or more than 50 words.
[0093] Question style: The expression style of questions can also be distinguished, mainly including colloquial and written styles. Colloquial questions are suitable for more natural and daily conversation scenarios, while written questions are suitable for formal and professional application scenarios.
[0094] Answer detail level: According to the complexity of the question and the need for the answer, the answer can be divided into three detail levels: brief, moderate, and as detailed as possible. Brief answers are suitable for simple questions, moderate answers are suitable for common applications, and answers as detailed as possible are suitable for complex questions to provide more comprehensive information.
[0095] The above-mentioned multiple Q&A styles can be randomly combined with limitations. For example, the Q&A style can be "question length less than 25 characters"; it can also be "question length less than 25 characters + colloquial style"; it can also be "question length less than 25 characters + colloquial style + detailed answer". Based on the control of the Q&A style, the computing system 200 can flexibly adjust the length, style of the question, and the detail level of the answer when generating Q&A, so as to ensure that the generated Q&A data meets the requirements of specific applications in terms of diversity and practicality.
[0096] After obtaining multiple question types and multiple Q&A styles, this specification provides two embodiments for guiding the generation model to generate Q&A samples: The computing system 200 can generate multiple guiding strategies, and then guide the generation model to generate Q&A samples based on a single guiding strategy; the computing system 200 can also directly guide the generation model to generate Q&A samples for different combinations of the multiple question types and the multiple Q&A styles.
[0097] In some embodiments, the computing system 200 can generate multiple guiding strategies, and then guide the generation model to generate Q&A samples based on a single guiding strategy. The computing system 200 combines the multiple question types and the multiple Q&A styles to generate multiple guiding strategies, and each guiding strategy represents a question type and a Q&A style; traverse the multiple guiding strategies, and for each current guiding strategy during the traversal: generate a first guiding instruction based on the current guiding strategy, and input the first guiding instruction into the generation model to guide the generation model to generate at least one Q&A sample corresponding to the current guiding strategy.
[0098] In some embodiments, the guiding strategy can be realized by randomly combining the multiple question types and Q&A styles with limitations. For example, the combination can be [question type is retrieval type, Q&A style is question length less than 25 characters + colloquial style]; it can also be [question type is retrieval type + single-hop question, Q&A style is question length less than 25 characters], etc.
[0099] In some embodiments, different first guiding instructions can be generated based on different guiding strategies. To ensure the comprehensiveness of the generated dataset, various guiding strategies can be traversed, and corresponding first guiding instructions can be generated for each current guiding strategy during the traversal. The first guiding instructions will serve as prompts for the generation model to guide the generation model to generate at least one Q&A sample corresponding to the current guiding strategy. By organically combining different types of questions and styles, the generated Q&A samples are more diverse in structure, content, and expression, thereby enabling the generated Q&A samples to better cover different application scenarios and requirements.
[0100] In some embodiments, targeted guiding strategies can also be generated based on different application scenarios. In this case, the multiple generated guiding strategies may only include the same type of question or Q&A style. For example, the question type in multiple guiding strategies is retrieval + single-hop questions. Such targeted guiding strategies can ensure that the dataset generated in a specific application scenario better meets the system requirements, thereby improving the performance and performance of the Q&A system in this scenario. By flexibly adjusting the guiding strategies, the computing system 200 can generate diverse and high-quality Q&A data in various tasks and scenarios, thereby improving the adaptability and practical application effect of the dataset.
[0101] In some embodiments, to ensure the accuracy of the generated Q&A samples. This specification provides different sampling strategies based on different guiding strategies. The computing system 200 samples a part of the content from the knowledge base as reference knowledge based on the current guiding strategy; and generates the first guiding instruction based on the current guiding strategy and the reference knowledge.
[0102] The reference knowledge sampled from the knowledge base will provide the necessary background and context for generating Q&A samples, ensuring that the generated Q&A samples can better reflect real application scenarios and requirements. The first guiding instruction generated based on the current guiding strategy and the reference knowledge will serve as a prompt for the generation model to guide the model to generate Q&A samples matching the current guiding strategy. By combining specific guiding strategies and sampled reference knowledge, the generation model can effectively adjust the question type and answer style during the Q&A generation process, thereby improving the overall performance of the Q&A system.
[0103] In some embodiments, the sampling strategy varies based on the type of questions characterized in the current guiding strategy. Specifically, the knowledge base includes at least one document, and each document is divided into multiple document chunks. Based on the current guiding strategy, sampling at least one text knowledge from the knowledge base includes: if the type of question characterized by the current guiding strategy is a single-hop question, sampling one document or one document chunk from the knowledge base as the reference knowledge; if the type of question characterized by the current guiding strategy is a multi-hop question, sampling multiple documents or multiple document chunks from the knowledge base as the reference knowledge.
[0104] For single-hop questions, the sampling strategy generally includes directly extracting a single information source related to the question from the knowledge base, usually only involving the information of a single document or document chunk. The sampled document or document chunk can directly generate a single-hop question without the need to reason across multiple documents or document chunks. Methods such as random sampling, iterative sampling, and weight-based sampling can be used to sample the single information source.
[0105] For multi-hop questions, it usually involves the information association between multiple documents or document chunks, and continuous reasoning or searching across multiple documents or information sources is required to generate multi-hop questions. Therefore, in the embodiments of this specification, a sampling method based on a knowledge graph is also provided to handle multi-hop questions.
[0106] If multiple documents need to be sampled, the computing system 200 samples a first document from the knowledge base, samples at least one second document associated with the first document from the knowledge base based on the first knowledge graph, and uses the first document and the at least one second document as the reference knowledge, where the first knowledge graph represents the association relationship between different documents. If multiple document chunks need to be sampled, the computing system 200 samples a first document chunk from the knowledge base, samples at least one second document chunk associated with the first document chunk from the knowledge base based on the second knowledge graph, and uses the first document chunk and the at least one second document chunk as the reference knowledge, where the second knowledge graph represents the association relationship between different document chunks.
[0107] Based on the knowledge graph in the above embodiments, the association relationship between different documents or document chunks can be effectively identified and utilized, helping the computing system 200 obtain answers from multiple information sources. The computing system 200 can efficiently integrate the information in multiple relevant documents or document chunks based on different knowledge graphs, realize reasoning across multiple information sources, and ensure the generation of more accurate multi-hop questions.
[0108] In some embodiments, the computing system 200 may also directly guide the generation model to generate Q&A samples for different combinations of the multiple question types and the multiple Q&A styles. Specifically, the computing system 200 generates a second guiding instruction based on the multiple question types and the multiple Q&A styles, and the second guiding instruction is used to guide the generation model to generate at least one Q&A sample based on the knowledge base for different combinations of the multiple question types and the multiple Q&A styles; the second guiding instruction is input into the generation model to obtain the multiple Q&A samples.
[0109] Among them, the generation model can combine multiple question types and multiple Q&A styles based on the second guiding instruction, and then generate diverse Q&A samples. In addition, the generation model can also generate the above guiding strategy based on the second guiding instruction according to the generation method of the foregoing guiding strategy, and further generate multiple Q&A samples based on these guiding strategies. The difference here is that the guiding strategy itself is automatically generated by the generation model according to the required question types and Q&A styles, rather than being preset in advance. The generation model can directly generate multiple guiding strategies, and generate multiple Q&A samples based on the multiple guiding strategies, and these samples can cover various question styles, complexities, and presentation forms. This way of adaptively generating guiding strategies based on the generation model enhances the flexibility and adaptability of the Q&A system, and can better meet the changing application scenarios and user needs.
[0110] In the current embodiment, the computing system 200 also combines a preset sampling strategy to further optimize the generation process based on multiple question types, Q&A styles, and the sampling strategy. Specifically, the computing system 200 generates the second guiding instruction based on the multiple question types, the multiple Q&A styles, and the preset sampling strategy, where the second guiding instruction is used to guide the generation model: generate multiple combinations based on the multiple question types and the multiple Q&A styles, and for each combination, sample some content in the knowledge base as reference knowledge according to the sampling strategy, and sample at least one text knowledge from the reference knowledge. The sampling strategy here is similar to the foregoing, so it will not be elaborated. The difference is that the foregoing sampling strategy (single-hop questions and multi-hop questions) can be provided to the generation model through the second guiding instruction, and the generation model samples according to the sampling strategy for each combination.
[0111] S330: Generate the data set based on the multiple Q&A samples.
[0112] In some embodiments, the Q&A samples in the generated dataset can come from the same application scenario, ensuring the applicability and consistency of the Q&A samples in a specific scenario. These Q&A samples can also be generated based on the same guiding strategy, ensuring the consistency in style and structure of the Q&A samples in the dataset. The dataset can also include Q&A samples under multiple different guiding strategies or application scenarios, thereby enhancing the diversity and comprehensiveness of the dataset. Based on different usage scenarios, the dataset can meet different requirements. Those skilled in the art can flexibly generate various different datasets based on the dataset generation method in the embodiments of this specification, so as to provide more precise support in the model training, evaluation, or optimization process.
[0113] Please refer to Figure 4 , this specification also provides an automatic generation link for Q&A samples. This link includes two key modules: a generation module and a post-processing module. Each module plays an important role in the process of generating, optimizing, and controlling the quality of Q&A samples to ensure that the generated Q&A samples are both accurate and reliable.
[0114] Among them, in the above implementation, the embodiment of the generation module is mainly included. The core task of this module is to extract relevant information from the knowledge base through a generation model, and generate multiple Q&A samples based on a combination of multiple question types, multiple Q&A styles, and sampling strategies. This process can ensure that the Q&A samples cover diverse question styles and structures to better adapt to various application scenarios and requirements.
[0115] The post-processing module focuses on further optimizing the content and structure of the dataset after it is generated to improve the accuracy and quality of the Q&A. The post-processing module mainly includes derivative processing, duplicate removal processing, and low-quality removal processing. Among them, derivative processing aims to simulate various variants and complex expressions that may occur when actual users ask questions, enhancing the diversity and authenticity of the dataset. Duplicate removal processing and low-quality removal processing are used to remove those redundant samples, samples with low quality, samples that do not meet the actual application requirements or have no practical significance, ensuring the quality of the final dataset.
[0116] The Q&A data obtained based on the generation module is usually relatively standardized. However, in actual application scenarios, users' questions may not be as regular, complete, or easy to understand as expected. To make the generated dataset closer to the actual scenario and cover as many question styles as possible, the post-processing module can first perform derivative processing on the question information in the Q&A samples. Specifically, the computing system 200 performs derivative processing on at least some of the multiple Q&A samples to obtain multiple derivative Q&A samples, and generates an intermediate dataset based on the multiple Q&A samples and the multiple derivative Q&A samples; and removes the Q&A samples in the intermediate dataset that do not meet the expectations to obtain the dataset.
[0117] In some embodiments, the process of derivative processing includes: the computing system 200 rewriting the question information in the target Q&A sample to obtain derivative question information; and generating the derivative Q&A sample based on the derivative question information and the answer information in the target Q&A sample.
[0118] Among them, the process of rewriting the question information aims to simulate different user question-asking ways, such as adjusting the grammar structure, replacing synonyms or adjusting the expression of the question, so as to better reflect the diverse question-asking habits of users in actual scenarios, to help the Q&A system cope with possible language diversity and expression variations, and to enhance the coverage of the dataset. And, the computing system 200 will retain the answer information in the target Q&A sample and generate a new Q&A sample in combination with the derivative question information. The derivative Q&A sample generated in this way can not only cover the variants of the original question. In the evaluation scenario, it can also verify the answer accuracy and consistency of the Q&A system under different question formulations without changing the answer. Through the above process of derivative processing, the dataset becomes richer, and it can better simulate the diverse user questions in actual applications, improving the adaptability and robustness of the Q&A system in the actual environment. At the same time, the Q&A samples generated in this way can more comprehensively cover different types of questions and expressions, further enhancing the training and evaluation effects of the Q&A system.
[0119] In the process of generating the above large number of Q&A samples, it is inevitable to generate some samples with repeated meanings. To avoid duplicate dirty data in the evaluation process and ensure the validity of the dataset, duplicate elimination processing can be performed on the Q&A samples. The duplicate elimination processing includes clustering processing and duplicate sample elimination processing. The purpose of clustering processing is to cluster the samples of the same category together. The purpose of duplicate sample elimination processing is to eliminate redundant Q&A samples in the dataset and select the qualified Q&A samples as valid data to be added to the dataset. Specifically, in some embodiments, the duplicate elimination processing includes: the computing system 200 clustering the Q&A samples in the intermediate dataset based on the semantics of each Q&A sample in the intermediate dataset to obtain at least two subsets; and eliminating some samples in each subset.
[0120] Through clustering processing, the computing system 200 can identify Q&A samples with similar or repeated semantics and classify them into the same category, thus avoiding the repeated existence of the same type of questions. Using the clustering method, similar question information or similar answer information is aggregated together to obtain at least two subsets. In each subset, those samples with a high degree of repetition or low quality compared with other samples in the subset are eliminated, and only the representative and informative Q&A samples are retained in each subset. Ensure that the valid Q&A samples retained in the dataset are highly diverse, information-rich and non-repetitive. Through the duplicate elimination processing in the above embodiments, the repeated or invalid Q&A samples in the dataset can be effectively removed, thus ensuring the quality and reliability of the dataset.
[0121] Due to the hallucination problem of the LLM, there is still an inevitable probability of wrong examples in the generated Q&A samples. To improve the quality of the Q&A samples and reduce the impact of wrong examples, the post-processing module also includes a low-quality elimination process, which automatically controls the quality of the Q&A samples through the low-quality elimination process, thereby improving the quality of the entire dataset. In some embodiments, the low-quality elimination process includes: the computing system 200 provides each Q&A sample in the intermediate dataset to the scoring model and obtains the quality score output by the scoring model; and eliminates the Q&A samples in the intermediate dataset whose quality scores do not meet the expectations.
[0122] In some embodiments, the basic model can be trained with a large amount of data, and the basic model can be fine-tuned in the direction of the quality scoring preference to continuously improve the basic model's judgment ability for quality scoring, resulting in a scoring model with the ability to generate quality scores based on Q&A samples. The above method of training the scoring model is applicable to scenarios with a certain amount of supervised data in stock.
[0123] In some embodiments, the scoring model can also be guided to generate scores to evaluate the quality of Q&A samples by means of guiding instructions. By designing specific guiding instructions, the scoring model can be effectively guided to generate scores consistent with the quality scoring criteria. By adjusting the content and structure of the guiding instructions, the scoring results of the scoring model can be significantly affected, making it more in line with the actual needs. This prompt-based method is applicable to zero-shot or few-shot learning scenarios, in which the scoring model needs to reason and evaluate without or with only a small amount of training data.
[0124] Specifically, the computing system 200 obtains a preset quality scoring standard, and the quality scoring standard includes scoring description information corresponding to multiple scoring dimensions; generates a guiding instruction based on the quality scoring standard and the Q&A sample, and the guiding instruction is used to guide the scoring model to score the target Q&A sample according to the quality scoring standard; inputs the guiding instruction into the scoring model and obtains the quality score output by the scoring model.
[0125] Please refer to Table 1. This specification provides a preset quality scoring criterion for generating guiding instructions, which are used to guide a scoring model to evaluate the quality score of the generated question-and-answer samples. Table 1 lists the scoring dimensions, corresponding sub-dimensions, descriptions and weights of each sub-dimension. The scoring model can score the question-and-answer samples according to different sub-dimensions. Each sub-dimension represents a specific quality indicator. For example, the question quality dimension includes sub-dimensions such as integrity, consistency and rationality. The score of each sub-dimension reflects the performance degree of this specific indicator in the sample. The scoring model can calculate the weighted sum based on the scores of each sub-dimension and the corresponding weights, so as to generate the quality score of the question-and-answer samples.
[0126] Table 1
[0127]
[0128] The computing system 200 can determine whether to eliminate a question-and-answer sample based on whether the quality score of the question-and-answer sample falls within the expected confidence interval or quality score threshold interval. If the quality score exceeds the predetermined range or fails to reach the set threshold, it is determined that the quality of the sample does not meet the expectation, and the question-and-answer sample is eliminated.
[0129] In the above embodiments, the scoring method based on guiding instructions enables the scoring model to still accurately evaluate the quality of question-and-answer samples in the absence of a large amount of labeled data. Especially for new fields or scenarios with scarce data, it provides an effective solution. Through the above weighted scoring mechanism, the scoring model can comprehensively evaluate the quality of each question-and-answer sample according to the importance of different quality indicators, ensuring that the final score is more objective and comprehensive. The computing system 200 can effectively identify high-quality and low-quality question-and-answer samples based on the quality scores, thereby optimizing the generated intermediate data set and improving its practicality and accuracy in the training and evaluation processes.
[0130] It is worth noting that in practical applications, those skilled in the art can select some processing operations in the post-processing module based on actual needs to post-process the data set generated by the generation model. For example, before performing duplicate elimination processing, the computing system 200 can directly perform duplicate elimination processing based on removing question-and-answer samples with high duplication or redundancy without performing derivation processing. For another example, the computing system 200 can also directly perform low-quality elimination processing on the data set generated by the generation model to eliminate those question-and-answer samples with poor quality. For another example, the computing system 200 can also perform duplicate elimination processing on the data set generated by the generation model first and then perform low-quality elimination processing, so as to more precisely optimize the data set. Those skilled in the art can flexibly select appropriate post-processing methods based on different scenarios and fields to ensure that the finally obtained optimized data set meets specific goals and requirements in terms of quality and diversity.
[0131] In summary, in the method and system for generating a dataset provided in this specification, a knowledge base is obtained; a generation model is guided to generate a plurality of question-and-answer samples based on the knowledge base, where each question-and-answer sample includes: question information, evidence information, and answer information, the question information is related to the content in the knowledge base, the evidence information is the information relied on by the generation model in the process of answering the question information, and the answer information is the information generated by the generation model for answering the question information; and the dataset is generated based on the plurality of question-and-answer samples. In the embodiments of this specification, by guiding the generation model to generate question-and-answer samples based on the knowledge base and requiring the generation of corresponding evidence information, the question information and the answer information are made more coherent and logical, the hallucination problem is avoided, and high-quality question-and-answer samples are automatically generated. Moreover, the diversity and coverage of the dataset are enhanced based on the combination of different question types and question-and-answer styles. Through the optimization and quality control of the question-and-answer samples, the generated dataset not only better conforms to the user's question-asking habits in the actual application scenario, can further address the hallucination problem that may occur in the LLM generation process, but also ensures that the generated question-and-answer dataset has higher accuracy, coherence, and practicality, thereby enhancing the value and effect of the evaluation set in actual applications.
[0132] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for generating a data set. When the at least one instruction set is executed by a processor, the at least one instruction set directs the processor to perform the steps of the data set generation method P300 described in this specification. In some possible implementation manners, various aspects of this specification may also be implemented in the form of a program product, which includes program code. When the program product runs on a computing system 200, the program code is used to cause the computing system 200 to perform the steps of the data set generation method P300 described in this specification. The program product for implementing the above method may adopt a portable compact disc read-only memory (CD-ROM) including program code and may run on the computing system 200. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system. The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium may include a data signal propagated as part of a carrier wave in a baseband, where the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification may be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the computing system 200, partially on the computing system 200, executed as a stand-alone software package, partially on the computing system 200 and partially on a remote computing system, or entirely on a remote computing system.
[0133] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0134] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0135] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0136] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well mark out some of the devices as separate embodiments for understanding when reading this specification. That is to say, the embodiments in this specification may also be understood as an integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0137] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except those that are inconsistent with or conflict with this document or that have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0138] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for generating a dataset, the method comprising: Obtaining a knowledge base; Guiding a generation model to generate a plurality of question-and-answer samples based on the knowledge base, wherein each question-and-answer sample includes: question information, evidence information, and answer information, the question information is related to the content in the knowledge base, the evidence information is the information relied on by the generation model in the process of answering the question information, and the answer information is the information generated by the generation model for answering the question information; and Generating the dataset based on the plurality of question-and-answer samples.
2. The method according to claim 1, wherein The guiding the generation model to generate a plurality of question-and-answer samples based on the knowledge base includes: Obtaining a variety of question types; and Based on the variety of question types, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base.
3. The method according to claim 2, wherein The based on the variety of question types, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base includes: Obtaining a variety of question-and-answer styles, each question-and-answer style characterizing at least one of the style corresponding to the question information or the style corresponding to the answer information; and Based on the variety of question types and the variety of question-and-answer styles, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base.
4. The method according to claim 3, wherein The based on the variety of question types and the variety of question-and-answer styles, guiding the generation model to generate the plurality of question-and-answer samples based on the knowledge base includes: Combining the variety of question types and the variety of question-and-answer styles to generate a variety of guiding strategies, each guiding strategy characterizing a question type and a question-and-answer style; Traversing the variety of guiding strategies, for each current guiding strategy in the traversal: Generating a first guiding instruction based on the current guiding strategy and inputting the first guiding instruction into the generation model to guide the generation model to generate at least one question-and-answer sample corresponding to the current guiding strategy.
5. The method according to claim 4, wherein The generating the first guiding instruction based on the current guiding strategy includes: Sampling a part of the content from the knowledge base as reference knowledge based on the current guiding strategy; and Generating the first guiding instruction based on the current guiding strategy and the reference knowledge.
6. The method according to claim 5, wherein the knowledge base includes at least one document, each document is divided into a plurality of document blocks, and the sampling at least one text knowledge from the knowledge base based on the current guiding strategy includes: If the question type characterized by the current guiding strategy is a single-hop question, sampling one document or one document block from the knowledge base as the reference knowledge; If the question type characterized by the current guiding strategy is a multi-hop question, sampling a plurality of documents or a plurality of document blocks from the knowledge base as the reference knowledge.
7. The method according to claim 6, wherein the sampling a plurality of documents from the knowledge base as the reference knowledge includes: Sample a first document from the knowledge base, sample at least one second document associated with the first document from the knowledge base based on a first knowledge graph, and use the first document and the at least one second document as the reference knowledge, where the first knowledge graph represents the association relationship between different documents; Sampling multiple document blocks from the knowledge base as the reference knowledge includes: Sample a first document block from the knowledge base, sample at least one second document block associated with the first document block from the knowledge base based on a second knowledge graph, and use the first document block and the at least one second document block as the reference knowledge, where the second knowledge graph represents the association relationship between different document blocks.
8. The method according to claim 3, wherein Based on the multiple question types and the multiple Q&A styles, guiding the generation model to generate the multiple Q&A samples based on the knowledge base includes: Generate a second guiding instruction based on the multiple question types and the multiple Q&A styles, where the second guiding instruction is used to guide the generation model to generate at least one Q&A sample based on the knowledge base for different combinations of the multiple question types and the multiple Q&A styles respectively; Input the second guiding instruction into the generation model to obtain the multiple Q&A samples.
9. The method according to claim 8, wherein, Generating the second guiding instruction based on the multiple question types and the multiple Q&A styles includes: Generate the second guiding instruction based on the multiple question types, the multiple Q&A styles, and a preset sampling strategy, where the second guiding instruction is used to guide the generation model to: Generate multiple combinations based on the multiple question types and the multiple Q&A styles, For each combination, sample a part of the content in the knowledge base as the reference knowledge according to the sampling strategy, and sample at least one text knowledge from the reference knowledge.
10. The method according to claim 2, wherein, The multiple question types include M first question types, where M is an integer greater than 1, and the M first question types are obtained through the following method: Obtain multiple classification dimensions, and determine multiple candidate types corresponding to each classification dimension, where the multiple classification dimensions include at least two of: the purpose of the question, the number of documents involved in the question, or the question type; Combine the candidate types under different classification dimensions to obtain the M first question types.
11. The method according to claim 10, wherein, The data set is used to evaluate the performance of the Q&A system, and the multiple question types further include N second question types, where N is an integer greater than or equal to 1; The N second question types are question types determined based on the feedback data of the Q&A system.
12. The method according to claim 1, wherein Generating the data set based on the multiple Q&A samples includes: Perform derivation processing on at least some of the multiple Q&A samples to obtain multiple derived Q&A samples, and generate an intermediate data set based on the multiple Q&A samples and the multiple derived Q&A samples; and Remove the Q&A samples that do not meet the expectations in the intermediate data set to obtain the data set.
13. The method according to claim 12, wherein For a target Q&A sample among the multiple Q&A samples, the process of deriving the target Q&A sample includes: Rewriting the question information in the target Q&A sample to obtain derived question information; and Generating the derived Q&A sample based on the derived question information and the answer information in the target Q&A sample.
14. The method according to claim 12, wherein, The removing of the Q&A samples that do not meet the expectations in the intermediate dataset includes: Clustering the Q&A samples in the intermediate dataset based on the semantics of each Q&A sample in the intermediate dataset to obtain at least two subsets; and Removing some samples in each subset.
15. The method according to claim 12, wherein, The removing of the Q&A samples that do not meet the expectations in the intermediate dataset includes: Providing each Q&A sample in the intermediate dataset to a scoring model and obtaining the quality score output by the scoring model; and Removing the Q&A samples in the intermediate dataset whose quality scores do not meet the expectations.
16. The method according to claim 15, wherein, For a target Q&A sample in the intermediate dataset, the providing of the target Q&A sample to the scoring model and obtaining the quality score output by the scoring model includes: Obtaining a preset quality scoring criterion, where the quality scoring criterion includes scoring description information corresponding to multiple scoring dimensions; Generating a guiding instruction based on the quality scoring criterion and the Q&A sample, where the guiding instruction is used to guide the scoring model to score the target Q&A sample according to the quality scoring criterion; Inputting the guiding instruction into the scoring model and obtaining the quality score output by the scoring model.
17. The method according to claim 1, wherein The knowledge base is the knowledge base used by the Q&A system, and the dataset is used to evaluate the performance of the Q&A system.
18. A system for generating a dataset, comprising: At least one storage medium storing at least one instruction set for generating a dataset; And At least one processor communicatively connected to the at least one storage medium, wherein when the dataset generation system runs, the at least one processor reads the at least one instruction set and executes the method for generating a dataset according to any one of claims 1-17 according to the indication of the at least one instruction set.