Data generation method and device based on LLM
By building a knowledge base-based pipeline task and using LLM to generate evaluation test datasets, the problems of low efficiency and low data authenticity in existing technologies are solved, high-quality dataset generation and adaptation are achieved, and the effect of evaluation testing is improved.
Patent Information
- Application Number
- CN202510895614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies are inefficient in constructing datasets for evaluating LLM Agents, and the authenticity of data samples is not high, which cannot effectively reflect the knowledge domain of the knowledge base.
By constructing a knowledge base-based pipeline task, including multiple subtasks such as generating the initial data set and data sample adaptation, using LLM to generate evaluation test data sets, and adopting prompt word templates and knowledge graph technology, we ensure that the data samples are adapted to the knowledge domain.
It realizes the automatic generation of high-quality evaluation test data sets, improves the adaptability and authenticity of data samples, and can accurately reflect the knowledge domain requirements of the knowledge base.
Smart Images

Figure CN120806088A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the specification belongs to the field of artificial intelligence, and particularly relates to a data generation method and device based on an LLM. BACKGROUND
[0002] With the rapid development of RAG (Retrieval-Augmented Generation) technology, LLM (Large Language Model) Agents have been deployed on a large scale in enterprise-level applications based on RAG architecture. In current industry practices, when it is necessary to evaluate and test LLM Agents, a data set can be constructed to simulate the question and answer process in a real business scenario, and then based on the data samples in the data set, the knowledge retrieval accuracy and generation logic rationality of the LLM Agent can be verified.
[0003] In related technologies, when constructing a data set for evaluating and testing an LLM Agent, an artificial construction mode of "expert definition + manual annotation" is usually adopted; for example, a domain expert can first design a question template, and then combine business documents to manually annotate standard answers to construct data samples for evaluating and testing the LLM Agent.
[0004] However, the artificial construction method for generating a data set for evaluating and testing an LLM Agent not only has low efficiency, but also may have the problem that the authenticity of the data samples in the data set is not high enough. SUMMARY
[0005] The present specification proposes a data generation method based on an LLM, comprising:
[0006] acquiring a knowledge base for generating a data set, and constructing a pipeline task for generating the data set based on the knowledge base; wherein the data set is used for evaluating and testing a target service related to an LLM; the pipeline task comprises a plurality of subtasks executed in sequence; the plurality of subtasks comprise a first subtask for generating an initial data set based on the knowledge base; and a second subtask for adapting data samples in the initial data set to a knowledge field to which the knowledge base belongs;
[0007] in response to a request for executing the pipeline task, constructing a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and inputting the first prompt word to an LLM to perform inference calculation based on the first prompt word by the LLM, to generate an initial data set based on the knowledge base;
[0008] In response to the first subtask being completed, a second prompt word is constructed for the second subtask based on a preset prompt word template for the second subtask, and the second prompt word is input into the LLM to perform inference calculation by the LLM based on the second prompt word, so as to adapt the data samples in the initial data set to the knowledge field to which the knowledge base belongs, to obtain a target data set for evaluating and testing the target service.
[0009] Optionally, the data samples in the data set include sample pairs composed of question samples and answer samples; and the target service includes a service of optimizing and adjusting an answer output by the LLM by referencing knowledge in the knowledge base.
[0010] Optionally, the knowledge base contains a plurality of knowledge documents.
[0011] A first prompt word is constructed for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and the first prompt word is input into the LLM to perform inference calculation by the LLM based on the first prompt word, so as to generate an initial data set based on the knowledge base, including:
[0012] The knowledge documents contained in the knowledge base are sequentially designated as target knowledge documents, and a first prompt word is constructed for the first subtask based on a preset prompt word template for the first subtask and the target knowledge document designated from the knowledge base; wherein the first prompt word is used to instruct the LLM to generate at least one initial data sample based on knowledge materials retrieved from the target knowledge document;
[0013] The first prompt word is input into the LLM to perform inference calculation by the LLM based on the first prompt word, retrieve knowledge materials from the target knowledge document, and generate at least one initial data sample based on the retrieved knowledge materials;
[0014] The initial data samples generated by the LLM based on the knowledge materials retrieved from each knowledge document contained in the knowledge base are summarized to obtain an initial data set.
[0015] Optionally, the knowledge base contains a plurality of knowledge documents.
[0016] A first prompt word is constructed for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and the first prompt word is input into the LLM to perform inference calculation by the LLM based on the first prompt word, so as to generate an initial data set based on the knowledge base, including:
[0017] constructing knowledge documents contained in the knowledge base into a global knowledge graph; wherein nodes in the knowledge graph represent knowledge materials extracted from each knowledge document contained in the knowledge base; edges between nodes in the knowledge graph represent relationships between knowledge materials extracted from the knowledge base;
[0018] constructing a first prompt for the first subtask based on a prompt word template preset for the first subtask and the knowledge graph; wherein the first prompt is used to instruct the LLM to generate a plurality of initial data samples based on knowledge materials retrieved from the global knowledge graph;
[0019] inputting the first prompt into the LLM to perform inference calculation based on the first prompt by the LLM, retrieving knowledge materials from the global knowledge graph, and generating a plurality of initial data samples based on the retrieved knowledge materials;
[0020] constructing an initial data set based on the plurality of initial data samples generated by the LLM.
[0021] Optionally, a second prompt is constructed for the second subtask based on a prompt word template preset for the second subtask, and the second prompt is input into the LLM to perform inference calculation by the LLM based on the second prompt, and the data samples in the initial data set are adapted to the knowledge field to which the knowledge base belongs, including:
[0022] constructing a second prompt for the second subtask based on a prompt word template preset for the second subtask; wherein the second prompt is used to instruct the LLM to adapt the problem samples contained in the data samples in the initial data set to the knowledge field to which the knowledge base belongs;
[0023] inputting the second prompt into the LLM to perform inference calculation by the LLM based on the second prompt, and adapting the problem samples contained in the data samples in the initial data set to the knowledge field to which the knowledge base belongs.
[0024] Optionally, adapting the problem samples contained in the data samples in the initial data set to the knowledge field to which the knowledge base belongs includes:
[0025] converting the problem samples contained in the data samples in the initial data set into a data structure adapted to the knowledge field to which the knowledge base belongs; wherein the data structure includes a data structure conforming to the questioning manner of the user in the knowledge field.
[0026] Optionally, the data structure adapted to the knowledge domain to which the knowledge base belongs includes a data structure adapted to the application scenario to which the knowledge base belongs; wherein the data structure includes a data structure that conforms to the way users ask questions in the application scenario.
[0027] Optionally, the pipeline task further includes a third subtask of removing semantically similar duplicate data samples contained in the target data set;
[0028] The method further comprises:
[0029] Obtaining a data sample output by the LLM after being adapted to the knowledge domain to which the knowledge base belongs;
[0030] Calculating semantic similarity between the data sample output by the LLM and the data sample in the target data set;
[0031] Determining whether the semantic similarity reaches a preset threshold;
[0032] If yes, deleting the data sample output by the LLM as a duplicate data sample with semantic similarity to the data sample in the target data set;
[0033] If not, the data samples output by the LLM are stored in the target data set.
[0034] Optionally, the semantic similarity is represented by a ROUGE-L score.
[0035] Optionally, the pipeline task further includes a fourth subtask of converting question samples contained in the data samples in the target data set into colloquial question samples with the same semantics;
[0036] The method further comprises:
[0037] Constructing a third prompt word for the fourth subtask based on a preset prompt word template for the fourth subtask; wherein the third prompt word is used to instruct the LLM to convert the question sample contained in the data sample in the target data set into a colloquial question sample with the same semantics;
[0038] The third prompt word is input into the LLM, so that the LLM performs inference calculation based on the third prompt word, and converts the question samples contained in the data samples in the target data set into spoken question samples with the same semantics.
[0039] Optionally, the target service includes a service provided by an LLM agent built based on RAG technology.
[0040] This specification also proposes a data generation device based on LLM, including:
[0041] a construction module configured to obtain a knowledge base for generating a data set, and construct a pipeline task for generating the data set based on the knowledge base, wherein the data set is used for evaluation testing of a target service related to the LLM, the pipeline task comprises a plurality of sub-tasks executed in sequence, the plurality of sub-tasks comprises a first sub-task for generating an initial data set based on the knowledge base, and a second sub-task for adapting data samples in the initial data set to a knowledge field to which the knowledge base belongs;
[0042] a generation module configured to, in response to a request for executing the pipeline task, construct a first prompt for the first sub-task based on a preset prompt template for the first sub-task and the knowledge base, input the first prompt into the LLM, and perform inference calculation by the LLM based on the first prompt, to generate the initial data set based on the knowledge base;
[0043] an adaptation module configured to, in response to completion of execution of the first sub-task, construct a second prompt for the second sub-task based on a preset prompt template for the second sub-task, input the second prompt into the LLM, and perform inference calculation by the LLM based on the second prompt, to adapt the data samples in the initial data set to the knowledge field to which the knowledge base belongs, to obtain a target data set for evaluation testing of the target service.
[0044] In the above embodiment, on the one hand, by converting the task of generating the evaluation data set based on the knowledge base into the form of the pipeline task composed of a plurality of sub-tasks, not only can the evaluation data set be automatically generated without human intervention by executing the pipeline task, but also each generation link of the evaluation data set can be modularized, so that each link of the evaluation data set can be extended or optimized independently of other links, thereby greatly enhancing the adaptability and customizability in generating the evaluation data set.
[0045] On the other hand, by introducing the sub-task of generating the initial evaluation data set based on the knowledge base and the sub-task of adapting the data samples in the initial evaluation data set to the knowledge field to which the knowledge base belongs in the above pipeline task, it can be ensured that the data samples in the generated evaluation data set can be adapted to the knowledge field to which the knowledge base belongs, so that the data samples in the finally generated evaluation data set not only have universal applicability, but also accurately reflect the actual needs and context of the knowledge field to which the knowledge base belongs, thereby significantly improving the authenticity of the generated data samples. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present specification, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0047] Figure 1 is a flow chart of a data generation method based on LLM shown in an embodiment of the present specification;
[0048] Figure 2 is a schematic diagram of automatically generating evaluation data sets based on pipeline tasks shown in an embodiment of the present specification;
[0049] Figure 3 is a schematic diagram of a preset prompt word template for a first subtask shown in an embodiment of the present specification;
[0050] Figure 4 is a schematic diagram of another preset prompt word template for a first subtask shown in an embodiment of the present specification;
[0051] Figure 5 is a schematic diagram of a preset prompt word template for a second subtask shown in an embodiment of the present specification;
[0052] Figure 6 is a schematic diagram of a preset prompt word template for a fourth subtask shown in an embodiment of the present specification;
[0053] Figure 7 is a schematic structural diagram of an electronic device shown in an embodiment of the present specification;
[0054] Figure 8 is a block diagram of a data generation device based on LLM shown in an embodiment of the present specification. DETAILED DESCRIPTION
[0055] In order for those skilled in the art to better understand the technical solutions in the present specification, the technical solutions in the embodiments of the present specification will be described clearly and completely in conjunction with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only some embodiments of the present specification, not all embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present specification.
[0056] The present specification aims to propose a technical solution for automatically generating an evaluation dataset that matches the knowledge domain of a knowledge base through a pipeline task in a scenario of performing evaluation testing on a target service related to an LLM.
[0057] Please refer to Figure 1 , Figure 1 A flowchart of an LLM-based data generation method shown in the present specification; the method includes the following execution process:
[0058] Step 102, obtaining a knowledge base for generating a dataset, and constructing a pipeline task for generating the dataset based on the knowledge base; wherein the dataset is used for evaluation testing on a target service related to an LLM; the pipeline task includes a plurality of subtasks executed in sequence; the plurality of subtasks include a first subtask for generating an initial dataset based on the knowledge base; and a second subtask for adapting data samples in the initial dataset to the knowledge domain of the knowledge base;
[0059] The above target service can specifically include a service related to an LLM; wherein the service related to an LLM can specifically include a service engine that uses an LLM as the underlying service engine.
[0060] For example, in an embodiment shown, an LLM Agent corresponding to the LLM can be constructed based on the RAG technology; wherein the RAG technology is a technology that combines information retrieval and generative LLMs to enhance the performance of LLMs in knowledge-intensive tasks. The LLM Agent is a service program that uses LLM as a service engine, can perform complex tasks (such as reasoning, planning, tool invocation, etc.), and has a certain degree of autonomy. Its core capabilities rely on the language understanding and generation capabilities of LLM, but it can also call external tools (such as search engines, APIs, databases) to supplement the capabilities of LLM. In this case, the above target service can specifically include the services provided by the LLM Agent.
[0061] It should be noted that the specific type of the above target service is not particularly limited in the present specification, and in actual applications, it can include any form of service that uses an LLM as the underlying service engine.
[0062] In an embodiment shown, the above target service can specifically include a service that optimizes and adjusts the output of the LLM by referencing the knowledge in the knowledge base. Based on this target service, the output content of the LLM can be optimized into more professional output content with the help of the knowledge in the knowledge base of the professional field.
[0063] For example, taking the classic LLM for question and answer as an example, the above target service can be a service that optimizes and adjusts the answer output by the LLM by referencing the knowledge in the knowledge base. In this scenario, based on the target service, the answer content output by the LLM can be optimized into a more professional answer with the help of the knowledge in the knowledge base of the professional field.
[0064] In actual applications, when the target service needs to be evaluated and tested, the question and answer process in the real knowledge field can usually be simulated by constructing an evaluation data set. For example, the data samples in the constructed evaluation data set can specifically include sample pairs composed of question samples and answer samples. Then, based on the data samples in the evaluation data set, the knowledge retrieval accuracy and the rationality of the generation logic of the target service can be verified.
[0065] When constructing the evaluation data set for evaluating and testing the target service, the knowledge base for generating the evaluation data set can be obtained first, and then the evaluation data set can be automatically constructed based on the knowledge in the knowledge base.
[0066] It should be noted that the knowledge stored in the knowledge base can specifically include knowledge related to the knowledge field that needs to be evaluated. In actual applications, after determining the knowledge field that needs to be evaluated, the knowledge related to the knowledge field can be collected, and then the knowledge base can be constructed based on the collected knowledge.
[0067] The specific form of the knowledge stored in the knowledge base is not particularly limited in the present specification.
[0068] For example, in one embodiment shown, the knowledge in the knowledge base can specifically be in the form of knowledge documents, and in this case, the knowledge base can specifically be a knowledge base composed of several knowledge documents.
[0069] After obtaining the knowledge base for generating the evaluation data set, in order to automatically generate the data set in batches, a pipeline task can also be created; wherein the pipeline task can be specifically used to automatically generate an evaluation data set that is adapted to the knowledge field to which the knowledge base belongs according to the knowledge in the knowledge base.
[0070] Please refer to Figure 2 , Figure 2 A schematic diagram of automatically generating an evaluation data set based on a pipeline task is shown in the present specification.
[0071] As Figure 2As shown, the pipeline task can include a plurality of sub-tasks executed in sequence. The plurality of sub-tasks can include a first sub-task and a second sub-task. The first sub-task can be configured to generate an initial data set based on knowledge in the knowledge base. The second sub-task can be configured to adapt data samples in the initial data set to a knowledge domain to which the knowledge base belongs, to obtain a target data set for evaluating the target service.
[0072] It should be noted that the first sub-task and the second sub-task can be referred to as basic tasks included in the pipeline task. In actual applications, the sub-tasks included in the pipeline task can be flexibly extended based on specific requirements on the basis of the first sub-task and the second sub-task.
[0073] For example, please continue to refer to Figure 2 In an embodiment, the pipeline task can include a third sub-task and a fourth sub-task in addition to the first sub-task and the second sub-task.
[0074] The third sub-task can be configured to remove duplicate data samples with similar semantics in the target data set generated by executing the second sub-task. The fourth sub-task can be configured to convert problem samples in the target data set into colloquial problem samples with the same semantics.
[0075] At step 104, in response to a request to execute the pipeline task, a first prompt word is constructed for the first sub-task based on a prompt word template preset for the first sub-task and the knowledge base, and the first prompt word is input to the LLM to perform inference calculation based on the first prompt word by the LLM, to generate an initial data set based on the knowledge base;
[0076] After the pipeline task is created, a request to execute the pipeline task can be received, and in response to the received request to execute the pipeline task, each sub-task included in the pipeline task can be executed in a preset execution order.
[0077] It should be noted that in actual applications, the inference capability of the LLM itself can be used to execute the sub-tasks included in the pipeline task. For example, a prompt word can be constructed for each sub-task included in the pipeline task by using a preset prompt word template, and the constructed prompt word can be input to the LLM to perform inference calculation based on the prompt word, to execute each sub-task.
[0078] In this case, after receiving a request to perform the pipeline task, a prompt word template preset for the first subtask can be obtained in response to the received request, and then a first prompt word can be constructed for the first subtask based on the prompt word template preset for the first subtask and the knowledge base, and the first prompt word can be input into the LLM to perform inference calculation by the LLM based on the first prompt word, and generate an initial data set based on the knowledge in the knowledge base.
[0079] It should be noted that the generation method used when generating the initial data set based on the knowledge in the knowledge base will not be particularly limited in this specification.
[0080] In an embodiment shown, taking the knowledge in the knowledge base as a knowledge document as an example, the initial data set can be generated by processing the knowledge documents in the knowledge base one by one.
[0081] In this case, when performing the first subtask, the knowledge documents contained in the knowledge base can be sequentially designated as target knowledge documents, and a first prompt word can be constructed for the first subtask based on the prompt word template preset for the first subtask and the target knowledge document designated from the knowledge base; wherein the first prompt word is used to instruct the LLM to generate at least one initial data sample based on the knowledge material retrieved from the target knowledge document.
[0082] For example, each knowledge document can be designated to generate at most N initial data samples. In actual applications, in order to control the number of initial data samples that can be generated for each knowledge document, the value of N can be set to a small value (such as 1-2).
[0083] Then, the first prompt word can be input into the LLM to perform inference calculation by the LLM based on the first prompt word, retrieve knowledge material from the target knowledge document, and generate at least one initial data sample based on the knowledge material retrieved from the target knowledge document.
[0084] For example, taking the above data sample as a sample pair composed of a question sample and an answer sample, at least one question sample and at least one answer sample corresponding to the question sample can be constructed based on the knowledge material retrieved from the target knowledge document.
[0085] After each knowledge document contained in the knowledge base is processed, the initial data samples generated by the LLM based on the knowledge material retrieved from each knowledge document contained in the knowledge base can be summarized to obtain an initial data set.
[0086] By adopting the manner of processing the knowledge documents in the knowledge base one by one, the initial data set is generated, so that the knowledge coverage of each initial data sample generated will focus on the specified target knowledge document, so as to ensure that each initial data sample generated can find the source of the knowledge reference, facilitating backtracking.
[0087] It should be noted that if the initial data set is generated by adopting the manner of processing the knowledge documents in the knowledge base one by one, the data structure and specific content of the prompt word template preset for the first subtask will not be specifically limited in the present specification, and can be flexibly configured based on specific needs in actual application.
[0088] For example, please refer to Figure 3 , Figure 3 The present specification shows a schematic diagram of a prompt word template preset for the first subtask; taking the above-mentioned data sample as an example of a sample pair composed of a question sample and an answer sample, if the initial data set is generated by adopting the manner of processing the knowledge documents in the knowledge base one by one, the prompt word template preset for the first subtask can refer to the form shown in Figure 3 .
[0089] In an embodiment shown, still taking the knowledge in the knowledge base as a knowledge document for example, the initial data set can also be generated by adopting the manner of processing the knowledge documents in the knowledge base globally.
[0090] In this case, when performing the first subtask, the knowledge documents contained in the knowledge base can be constructed into a global knowledge graph; wherein the nodes in the knowledge graph represent the knowledge materials extracted from the knowledge documents contained in the knowledge base. The edges between the nodes in the knowledge graph represent the relationships between the knowledge materials extracted from the knowledge base.
[0091] For example, in actual application, the GraphRAG technology can be used to extract the knowledge materials and the relationships between the knowledge materials from all the knowledge documents contained in the knowledge base, and then based on the extracted knowledge materials and the relationships between the knowledge materials, a global knowledge graph is constructed.
[0092] After the knowledge documents contained in the knowledge base are constructed into a global knowledge graph, the first prompt for the first subtask can be constructed based on the prompt word template preset for the first subtask and the knowledge graph; wherein the first prompt is used to instruct the LLM to generate a plurality of initial data samples based on the knowledge materials retrieved from the global knowledge graph; for example, the specific number of the plurality of initial data samples can be the maximum number of data samples that the initial data set can contain, which is preset in advance.
[0093] After the first prompt word is constructed for the first subtask, the first prompt word can be input to the LLM to perform inference calculation based on the first prompt word, retrieve knowledge materials from the global knowledge graph, and generate a plurality of initial data samples based on the retrieved knowledge materials, and then an initial data set can be constructed based on the plurality of initial data samples generated by the LLM.
[0094] By adopting the global processing manner for the knowledge documents in the knowledge base to generate the initial data set, the knowledge coverage of each initial data sample generated will be focused on all the knowledge documents in the knowledge base, so that each initial data sample generated can be a comprehensive question across the knowledge documents.
[0095] It should be noted that if the global processing manner for the knowledge documents in the knowledge base is adopted to generate the initial data set, the data structure and specific content of the prompt word template preset for the first subtask are no longer specifically limited in the present specification, and can be flexibly configured based on specific requirements in actual application.
[0096] For example, please refer to Figure 4 , Figure 4 Another schematic diagram of the prompt word template preset for the first subtask shown in the present specification; still taking the above-mentioned data sample as an example of a sample pair composed of a question sample and an answer sample, if the global processing manner for the knowledge documents in the knowledge base is adopted to generate the initial data set, the prompt word template preset for the first subtask can refer to the form shown in Figure 4 .
[0097] Step 106, in response to the completion of the execution of the first subtask, a second prompt word is constructed for the second subtask based on the prompt word template preset for the second subtask, and the second prompt word is input to the LLM to perform inference calculation based on the second prompt word by the LLM, and the data samples in the initial data set are adapted to the knowledge field to which the knowledge base belongs to obtain a target data set for evaluating and testing the target service.
[0098] After the first subtask included in the above-mentioned pipeline task is executed, the second subtask included in the above-mentioned pipeline task can be executed, the prompt word template preset for the second subtask is obtained, and a second prompt word is constructed for the second subtask based on the prompt word template preset for the second subtask, and then the second prompt word is input to the LLM to perform inference calculation based on the second prompt word by the LLM, and the data samples in the initial data set generated by executing the first subtask are adapted to the knowledge field to which the knowledge base belongs.
[0099] In an embodiment shown, taking the sample pair composed of the question sample and the answer sample in the above data sample as an example, since the answer sample is usually the knowledge contained in the knowledge base with relatively high standardization, and the question sample is usually the question generated by the LLM based on the inference ability of the LLM on the knowledge contained in the knowledge base, the standardization of the question sample is usually lower than that of the answer sample. Therefore, when adapting the data sample in the initial data set to the knowledge field to which the knowledge base belongs, only the question sample contained in the data sample can be adapted to the knowledge field to which the knowledge base belongs, and the answer sample contained in the data sample can not be adapted to the knowledge field to which the knowledge base belongs.
[0100] In this case, when performing the second subtask included in the pipeline task, a second prompt can still be constructed for the second subtask based on the prompt word template preset for the second subtask; wherein the second prompt can be used to instruct the LLM to adapt the question sample contained in the data sample in the initial data set to the knowledge field to which the knowledge base belongs. Then, the second prompt can be input to the LLM to perform inference calculation based on the second prompt by the LLM to adapt the question sample contained in the data sample in the initial data set to the knowledge field to which the knowledge base belongs.
[0101] It should be noted that adapting the data sample in the initial data set to the knowledge field to which the knowledge base belongs usually means converting the data sample in the initial data set into a data sample with higher specialization in the knowledge field to which the knowledge base belongs. The conversion method used when converting the data sample in the initial data set into a data sample with higher specialization in the knowledge field to which the knowledge base belongs is not particularly limited in the present specification, and can be selected flexibly based on specific needs in actual application.
[0102] For example, the prompt can be edited flexibly based on specific needs, and relevant instructions can be added to the prompt to specify the specific conversion method.
[0103] In an embodiment shown, taking the adaptation of the question sample contained in the data sample in the initial data set to the knowledge field to which the knowledge base belongs as an example, the conversion method can specifically include converting the question sample contained in the data sample in the initial data set into a data structure adapted to the knowledge field to which the knowledge base belongs.
[0104] In this case, when the problem samples contained in the data samples in the initial data set are adapted to the knowledge field to which the knowledge base belongs, the problem samples contained in the data samples in the initial data set can be specifically converted into a data structure adapted to the knowledge field; the data structure can specifically include a data structure conforming to the questioning manner of the user in the knowledge field to which the knowledge base belongs.
[0105] It should be noted that the questioning manner of the user in the knowledge field to which the knowledge base belongs usually has specific question and answer patterns or structural characteristics, and therefore the data structure conforming to the questioning manner of the user in the knowledge field to which the knowledge base belongs can usually include any form of data structure capable of embodying such specific question and answer patterns or structural characteristics.
[0106] In this way, the problem samples contained in the data samples in the initial data set can be converted into a data structure conforming to the questioning manner of the user in the knowledge field to which the knowledge base belongs, so that the problem samples are closer to the questioning manner of the user in the knowledge field, which helps to improve the authenticity of the data samples in the constructed data set.
[0107] In an embodiment shown, since the application scenario to which the knowledge base belongs can usually represent the knowledge field to which the knowledge base belongs, when the problem samples contained in the data samples in the initial data set are adapted to the knowledge field to which the knowledge base belongs, the problem samples contained in the data samples in the initial data set can also be converted into a data structure adapted to the application scenario to which the knowledge base belongs; the data structure can specifically include a data structure conforming to the questioning manner of the user in the application scenario to which the knowledge base belongs.
[0108] It should be noted that the questioning manner of the user in the application scenario to which the knowledge base belongs usually has specific question and answer patterns or structural characteristics, and therefore the data structure conforming to the questioning manner of the user in the application scenario to which the knowledge base belongs can usually include any form of data structure capable of embodying such specific question and answer patterns or structural characteristics.
[0109] For example, in one example, in many application scenarios, the questioning initiated by the user usually conforms to a three-part data structure. In the three-part data structure, the first part usually corresponds to “what to do”, the second part usually corresponds to “what steps have been completed”, and the third part usually corresponds to “what problem then occurs”.
[0110] See Figure 5 , Figure 5A schematic diagram of a prompt template preset for a second subtask is shown in the present specification; in this case, instructions related to the three-part data structure can be added to the above-mentioned second prompt template to guide the LLM to convert the question samples contained in the data samples in the initial data set into the three-part data structure. At this time, the above-mentioned second prompt template can refer to the form shown in Figure 5 . In this way, the question samples contained in the data samples in the initial data set can be converted into a data structure that conforms to the user's way of asking questions in the application scenario to which the knowledge base belongs, so that these question samples are closer to the user's way of asking questions in the real application scenario, which helps to improve the authenticity of the data samples in the constructed data set.
[0111] When the data samples in the initial data set are adapted to the knowledge domain to which the above-mentioned knowledge base belongs, the data samples in the initial data set have been converted into data samples with a higher degree of specialization within the knowledge domain to which the knowledge base belongs, and at this time the data set can be used as a target data set for evaluating and testing the above-mentioned target service.
[0112] In actual applications, if the knowledge in the above-mentioned knowledge base is not carefully selected and cleaned, it may result in a high degree of coupling between the knowledge in the knowledge base, and there may be semantic repetition between some knowledge. For example, taking the knowledge in the knowledge base as knowledge documents, there may be repetition between these knowledge documents in the knowledge base, or there may be paragraphs with similar semantics.
[0113] In this case, there may be repeated or some fragments with very similar semantics in the above-mentioned target data set constructed by the LLM based on the knowledge in the knowledge base, which may result in uneven distribution of data samples in the target data set and poor diversity of data samples.
[0114] In one embodiment shown, please continue to refer to Figure 2 , in order to ensure the uniform distribution and diversity of the data samples in the target data set, the above-mentioned pipeline task can specifically include a third subtask in addition to the above-mentioned first and second subtasks. The third subtask can specifically be used to remove repeated data samples with similar semantics contained in the above-mentioned target data set generated by executing the above-mentioned second subtask.
[0115] In this case, after obtaining the target data set for evaluating and testing the above-mentioned target service by executing the above-mentioned second subtask, the above-mentioned third subtask can be further executed to remove repeated data samples with similar semantics contained in the above-mentioned target data set, so as to further optimize the data samples in the target data set.
[0116] It should be noted that, unlike the first subtask and the second subtask described above, the third subtask can no longer need to use the inference capability of the LLM to perform, but can discover and remove the semantically similar duplicate data samples contained in the target data set by performing semantic similarity calculation.
[0117] In one embodiment shown, when performing the third subtask, the data sample output by the LLM after adapting to the knowledge field to which the knowledge base belongs in the process of performing the second subtask, that is, the data sample obtained after the LLM adapts the data sample in the initial data set to the knowledge field to which the knowledge base belongs, can be obtained; Then, the semantic similarity between the data sample output by the LLM and the data sample in the target data set can be calculated, and it is determined whether the semantic similarity reaches a preset threshold;
[0118] If the semantic similarity between the data sample output by the LLM and the data sample in the target data set reaches the preset threshold, at this time there is a high semantic similarity between the data sample output by the LLM and the data sample in the target data set, the data sample output by the LLM can be deleted as a semantically similar duplicate data sample in the target data set;
[0119] On the contrary, if the semantic similarity between the data sample output by the LLM and the data sample in the target data set does not reach the preset threshold, at this time there is a low semantic similarity between the data sample output by the LLM and the data sample in the target data set, the data sample output by the LLM can be stored in the target data set.
[0120] For example, in actual application, the data sample in the target data set can be archived by default, and after obtaining the data sample output by the LLM after adapting to the knowledge field to which the knowledge base belongs, the semantic similarity between the data sample and the data sample in the target data set that has been archived can be calculated, at this time if the semantic similarity between the data sample output by the LLM and the data sample in the target data set that has been archived reaches the preset threshold, the data sample output by the LLM can also be archived to add the data sample to the target data set; On the contrary, the data sample output by the LLM is no longer archived and directly deleted.
[0121] In this way, it can be ensured that the data samples finally added to the target data set are all data samples with low semantic similarity, so that semantically similar duplicate data samples in the target data set can be removed.
[0122] It should be noted that the calculation method used when calculating the semantic similarity between the data sample of the LLM output and the data sample in the target data set, that is, the representation method of the semantic similarity, is not limited in the specification, and can be selected flexibly based on specific needs in actual application.
[0123] For example, in the embodiment shown, the semantic similarity can be represented by a ROUGE-L score. ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence) is an index based on LCS (Longest Common Subsequence) to measure the similarity between a candidate text and a set of reference texts. LCS refers to the longest common subsequence of two text sequences. For example, assuming there are two text sequences X and Y, if there is a sequence Z that is a subsequence of both X and Y, then Z is called a common subsequence of X and Y. Among them, the longest common subsequence is called the longest common subsequence. In actual application, the above-mentioned subsequence usually refers to a new sequence obtained by deleting some elements from the original sequence without changing the order of the remaining elements.
[0124] In this case, the ROUGE-L score between the data sample of the LLM output after being adapted to the knowledge domain to which the knowledge base belongs and the data sample in the target data set that has been archived can be calculated, and it is determined whether the calculated ROUGE-L score reaches a preset threshold; if so, the data sample of the LLM output can be deleted as a repeated data sample that is semantically similar to the data sample in the target data set; if not, the data sample of the LLM output can be added to the target data set.
[0125] It should be noted that the specific calculation process of calculating the ROUGE-L score between data samples is not described in detail in the specification, and those skilled in the art can refer to the description in related technologies in actual application.
[0126] Of course, in addition to using the ROUGE-L score to represent the semantic similarity, other representation methods can also be used to represent the semantic similarity in actual application, such as using the vector distance between vectors corresponding to texts to represent the semantic similarity, which will not be enumerated one by one in the specification.
[0127] In actual applications, when users ask questions to the LLM, the questions asked to the LLM may be some colloquial questions due to the lack of knowledge of the professional knowledge field. Based on this, in order to make the data samples in the target data set finally generated for the evaluation test of the target service more in line with the needs of users and more in line with real application scenarios, the data samples in the above target data set can also be rewritten in a colloquial manner to adapt to the questioning manner of ordinary users.
[0128] In an embodiment shown, please continue to refer to Figure 2 The above pipeline task can further include a fourth subtask. The fourth subtask can be specifically used to rewrite the data samples in the target data set in a colloquial manner.
[0129] For example, taking the above data sample as an example, the data sample is composed of a question sample and an answer sample, the fourth subtask can be used to convert the question sample contained in the data sample in the target data set into a colloquial question sample with the same semantics.
[0130] The fourth subtask can still be executed by using the inference ability of the LLM itself.
[0131] In this case, after the third subtask is executed, the prompt word template preset for the fourth subtask can be obtained, and the third prompt word for the fourth subtask can be constructed based on the prompt word template preset for the fourth subtask. The third prompt word is input into the LLM to perform inference calculation based on the third prompt word by the LLM, and the data samples in the target data set are rewritten in a colloquial manner, and the question sample contained in the data sample in the target data set is converted into a colloquial question sample with the same semantics.
[0132] It should be noted that the data structure and specific content of the prompt word template preset for the fourth subtask will not be limited in this specification, and in actual applications, it can be flexibly configured based on specific needs.
[0133] For example, please refer to Figure 6 , Figure 6 The schematic diagram of the prompt word template preset for the fourth subtask shown in this specification; the prompt word template preset for the fourth subtask can be in the form shown in Figure 6 .
[0134] When the fourth subtask is executed, the data samples in the target data set are not only adapted to the knowledge field to which the knowledge base belongs, and the repeated data samples with similar semantics are removed, but also are rewritten in a spoken language. At this time, the target data set can be used as a test set for evaluating the target service, and the target service can be evaluated.
[0135] In the above technical solution, on the one hand, by converting the task of generating the evaluation data set based on the knowledge base into the form of a pipeline task composed of multiple subtasks, the evaluation data set can be automatically generated without human intervention by executing the pipeline task, and each generation link of the evaluation data set can be modularized, so that each link of the evaluation data set can be extended or optimized independently of other links, thereby greatly enhancing the adaptability and customizability of the evaluation data set.
[0136] For example, as shown in Figure 2 the third subtask and the fourth subtask can be flexibly extended based on the first subtask and the second subtask to further optimize the data samples in the target data set obtained by executing the second subtask.
[0137] On the other hand, by introducing the subtasks of generating the initial evaluation data set based on the knowledge base and adapting the data samples in the initial evaluation data set to the knowledge field to which the knowledge base belongs in the pipeline task, it can be ensured that the data samples in the generated evaluation data set can be adapted to the knowledge field to which the knowledge base belongs, so that the data samples in the finally generated evaluation data set not only have universal applicability, but also accurately reflect the actual needs and context of the knowledge field to which the knowledge base belongs, thereby significantly improving the authenticity of the generated data samples.
[0138] For example, in actual application, by executing the subtask of adapting the data samples in the initial data set to the knowledge field to which the knowledge base belongs, the problem samples contained in the data samples in the initial data set can be converted into a data structure conforming to the questioning manner of the user in the knowledge field to which the knowledge base belongs, so that the problem samples are closer to the questioning manner of the user in the knowledge field, which helps to improve the authenticity of the data samples in the constructed data set.
[0139] Corresponding to the embodiments of the foregoing method, the present specification also provides embodiments of an apparatus, an electronic device, and a storage medium.
[0140] Figure 7 FIG. 1 is a schematic structural diagram of an electronic device according to an example embodiment. Please refer toFigure 7 At the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710, and of course can also include other required hardware. One or more embodiments of the present specification can be implemented in a software manner, such as reading a corresponding computer program from the non-volatile memory 710 into the memory 708 by the processor 702 and then running. Of course, in addition to the software implementation, one or more embodiments of the present specification do not exclude other implementation manners, such as a logic device or a combination of software and hardware, and the like, that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.
[0141] As shown in Figure 8 , Figure 8 is a block diagram of an LLM-based data generation device according to an exemplary embodiment of the present specification, which can run in an electronic device as shown in Figure 7 to implement the technical solutions of the present specification. The device 80 includes:
[0142] A construction module 801 acquires a knowledge base for generating a data set and constructs a pipeline task for generating the data set based on the knowledge base; wherein the data set is used to evaluate and test a target service related to an LLM; the pipeline task includes a plurality of subtasks executed in sequence; the plurality of subtasks include a first subtask for generating an initial data set based on the knowledge base; and a second subtask for adapting data samples in the initial data set to a knowledge field to which the knowledge base belongs;
[0143] A generation module 802, in response to a request for executing the pipeline task, constructs a first prompt for the first subtask based on a prompt word template preset for the first subtask and the knowledge base, and inputs the first prompt to an LLM to perform inference calculation based on the first prompt by the LLM, to generate an initial data set based on the knowledge base;
[0144] An optimization module 803, in response to completion of execution of the first subtask, constructs a second prompt for the second subtask based on a prompt word template preset for the second subtask, and inputs the second prompt to an LLM to perform inference calculation based on the second prompt by the LLM, to adapt data samples in the initial data set to a knowledge field to which the knowledge base belongs, to obtain a target data set for evaluating and testing the target service.
[0145] Correspondingly, the specification also provides an electronic device, comprising a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps in all the method processes described previously.
[0146] Correspondingly, the specification also provides a computer-readable storage medium having stored thereon executable computer program instructions; wherein the instructions, when executed by a processor, implement the steps in all the method processes described previously.
[0147] Correspondingly, the specification also provides a computer program product having stored thereon executable computer program instructions; wherein the computer program instructions, when executed by a processor, implement the steps in all the method processes described previously.
[0148] In the 1990s, it was relatively easy to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has evolved, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flows into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming the PLD, rather than by ordering a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in the above-mentioned hardware description languages and programmed into an integrated circuit, a hardware circuit implementing the logical method flow can be easily obtained.
[0149] The controller can be implemented in any suitable way, for example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. The skilled person will also appreciate that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the controller in the form of logic gates, switches, an application specific integrated circuit, a programmable logic controller and an embedded microcontroller, etc. to perform the same functions by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein to perform the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means to perform the various functions can be considered to be both a software module implementing the method and a structure within a hardware component.
[0150] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present application does not rule out that with the development of future computer technology, computers implementing the functions of the above embodiments can be personal computers, laptop computers, vehicle human-computer interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or combinations of any of these devices.
[0151] Although the method operations of the embodiments of the present disclosure are described in a particular, sequential order, one or more of the method operations can be omitted, or the method operations can be performed in an order other than the described order. Additionally, one or more of the method operations can be performed concurrently, or with partial concurrence. Furthermore, one or more of the method operations can be performed by different entities, or over different time periods. Additionally, the term "comprising" is used throughout to mean including, but not limited to, the elements listed after the term. Further, the term "or" as used herein means any one member of a logical disjunction (i.e., OR), and therefore, is synonymous with the term "and / or" as those terms are used in the art. The term "another" means an additional or extra, and thus, can mean one or more than one. The terms "comprise(s)," "comprising," "include(s)," and "including," as well as variations thereof, are intended to be open-ended, and include the presence of subjects, integers, steps, or components that are not recited in the disclosure. Further, the terms "first," "second," and the like, are used merely as labels, and are not intended to impose numerical requirements on their objects. The use of the term "about" in relation to a reference number is intended to encompass the reference number ±10% of the reference number.
[0152] For ease of description, the above apparatuses are described in functional modules for separate description. Of course, when implementing one or more of the present disclosure, the functions of the modules can be implemented in one or more software and / or hardware, or the modules implementing the same function can be implemented by a combination of sub-modules or sub-units. The apparatus embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division, and actual implementation can have another division manner. For example, multiple units or components can be combined or integrated into another system, or some features can be omitted or not implemented. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0153] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The term "means for" can be replaced with the term "configuration for". Figure 1 The term "module" can be replaced with the term "logic".
[0154] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 Figure 1 of the block or blocks.
[0156] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0157] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory. The memory is an example of computer-readable media.
[0158] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage, graphene storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0159] Those skilled in the art will appreciate that the one or more embodiments described herein can be provided as a method, a system or a computer program product. Accordingly, the one or more embodiments described herein can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the one or more embodiments described herein can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable code.
[0160] The one or more embodiments described herein can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The one or more embodiments described herein can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.
[0161] The various embodiments described in this specification are described in the context of progressive embodiments, with each embodiment building on the previous one. The same or similar parts between embodiments are cross-referenced as appropriate. Each embodiment focuses on the differences between that embodiment and the previous one. In particular, the system embodiments are described relatively simply, as they are substantially similar to the method embodiments. In the description of the specification, the use of the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the particular feature, structure, material or characteristic being described is included in at least one embodiment or example of the specification. Illustrative descriptions of the above terms do not necessarily refer to the same embodiment or example in this specification. Moreover, the particular features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. Furthermore, different embodiments or examples described in this specification can be combined and combined with features of other embodiments or examples, where such combinations do not contradict each other.
[0162] The above description merely provides examples of the one or more embodiments described in this specification and does not limit the one or more embodiments described in this specification. The one or more embodiments described in this specification can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the one or more embodiments described in this specification should be included in the scope of the claims.
Claims
1. A data generation method based on LLM, comprising: Obtaining a knowledge base for generating a data set, and constructing a pipeline task for generating the data set based on the knowledge base; wherein the data set is used to evaluate and test a target service related to the LLM; the pipeline task includes multiple subtasks executed in sequence; the multiple subtasks include a first subtask for generating an initial data set based on the knowledge base; and a second subtask for adapting data samples in the initial data set to the knowledge domain to which the knowledge base belongs; In response to a request to execute the pipeline task, construct a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and input the first prompt word into the LLM so that the LLM performs inference calculation based on the first prompt word and generates an initial data set based on the knowledge base; In response to the completion of the first subtask, a second prompt word is constructed for the second subtask based on the prompt word template preset for the second subtask, and the second prompt word is input into the LLM, so that the LLM performs inference calculation based on the second prompt word, and adapts the data samples in the initial data set to the knowledge domain to which the knowledge base belongs, so as to obtain a target data set for evaluating and testing the target service.
2. According to the method as claimed in claim 1, the data samples in the data set include sample pairs consisting of question samples and answer samples; the target service includes a service for optimizing and adjusting the answers output by the LLM by referencing the knowledge in the knowledge base.
3. The method according to claim 2, wherein the knowledge base comprises a plurality of knowledge documents; Constructing a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and inputting the first prompt word into an LLM so that the LLM performs inference calculation based on the first prompt word and generates an initial data set based on the knowledge base, including: The knowledge documents contained in the knowledge base are sequentially designated as target knowledge documents, and a first prompt word is constructed for the first subtask based on a prompt word template preset for the first subtask and the target knowledge document designated from the knowledge base; wherein the first prompt word is used to instruct the LLM to generate at least one initial data sample based on the knowledge material retrieved from the target knowledge document; Inputting the first prompt word into the LLM, so that the LLM performs inference calculation based on the first prompt word, retrieves knowledge materials from the target knowledge document, and generates at least one initial data sample based on the retrieved knowledge materials; The initial data samples generated by the LLM based on the knowledge materials retrieved from the various knowledge documents contained in the knowledge base are aggregated to obtain an initial data set.
4. The method according to claim 2, wherein the knowledge base comprises a plurality of knowledge documents; Constructing a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and inputting the first prompt word into an LLM so that the LLM performs inference calculation based on the first prompt word and generates an initial data set based on the knowledge base, including: The knowledge documents contained in the knowledge base are constructed into a global knowledge graph; wherein the nodes in the knowledge graph represent the knowledge materials extracted from the respective knowledge documents contained in the knowledge base; and the edges between the nodes in the knowledge graph represent the relationships between the knowledge materials extracted from the knowledge base; Constructing a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge graph; wherein the first prompt word is used to instruct the LLM to generate a number of initial data samples based on knowledge materials retrieved from the global knowledge graph; Inputting the first prompt word into the LLM, so that the LLM performs inference calculation based on the first prompt word, retrieves knowledge materials from the global knowledge graph, and generates a number of initial data samples based on the retrieved knowledge materials; An initial data set is constructed based on the initial data samples generated by the LLM.
5. The method according to claim 2, Constructing a second prompt word for the second subtask based on a preset prompt word template for the second subtask, and inputting the second prompt word into the LLM, so that the LLM performs inference calculation based on the second prompt word to adapt the data samples in the initial data set to the knowledge domain to which the knowledge base belongs, including: Constructing a second prompt word for the second subtask based on a preset prompt word template for the second subtask; wherein the second prompt word is used to instruct the LLM to adapt the question samples included in the data samples in the initial data set to the knowledge domain to which the knowledge base belongs; The second prompt word is input into the LLM, so that the LLM performs inference calculation based on the second prompt word, and adapts the question samples contained in the data samples in the initial data set to the knowledge domain to which the knowledge base belongs.
6. The method according to claim 5, wherein the problem samples contained in the data samples in the initial data set are adapted to the knowledge domain to which the knowledge base belongs, comprising: The question samples contained in the data samples in the initial data set are converted into a data structure adapted to the knowledge domain to which the knowledge base belongs; wherein the data structure includes a data structure that conforms to the user's questioning method in the knowledge domain.
7. The method according to claim 6, wherein the data structure adapted to the knowledge domain to which the knowledge base belongs comprises a data structure adapted to the application scenario to which the knowledge base belongs; wherein, The data structure includes a data structure that conforms to the user's questioning method in the application scenario.
8. The method according to claim 2, wherein the pipeline task further comprises a third subtask of removing semantically similar duplicate data samples contained in the target data set; The method further comprises: Obtaining a data sample output by the LLM after being adapted to the knowledge domain to which the knowledge base belongs; Calculating semantic similarity between the data sample output by the LLM and the data sample in the target data set; Determining whether the semantic similarity reaches a preset threshold; If yes, deleting the data sample output by the LLM as a duplicate data sample with semantic similarity to the data sample in the target data set; If not, the data samples output by the LLM are stored in the target data set.
9. The method according to claim 8, wherein the semantic similarity is represented by a ROUGE-L score.
10. The method according to claim 8, wherein the pipeline task further comprises a fourth subtask of converting question samples contained in the data samples in the target dataset into colloquial question samples with the same semantics; The method further comprises: Constructing a third prompt word for the fourth subtask based on a preset prompt word template for the fourth subtask; wherein the third prompt word is used to instruct the LLM to convert the question sample contained in the data sample in the target data set into a colloquial question sample with the same semantics; The third prompt word is input into the LLM, so that the LLM performs inference calculation based on the third prompt word, and converts the question samples contained in the data samples in the target data set into spoken question samples with the same semantics.
11. The method according to claim 1, wherein the target service comprises a service provided by an LLM agent constructed based on RAG technology.
12. A data generation device based on LLM, comprising: A construction module is configured to obtain a knowledge base for generating a data set and construct a pipeline task for generating the data set based on the knowledge base; wherein the data set is used to evaluate and test a target service related to the LLM; the pipeline task includes a plurality of subtasks executed in sequence; the plurality of subtasks includes a first subtask for generating an initial data set based on the knowledge base; and a second subtask for adapting data samples in the initial data set to the knowledge domain to which the knowledge base belongs; a generation module, in response to a request to execute the pipeline task, constructing a first prompt word for the first subtask based on a preset prompt word template for the first subtask and the knowledge base, and inputting the first prompt word into the LLM so that the LLM performs inference calculation based on the first prompt word and generates an initial data set based on the knowledge base; The adaptation module, in response to the completion of the first subtask, constructs a second prompt word for the second subtask based on the prompt word template preset for the second subtask, and inputs the second prompt word into the LLM, so that the LLM performs inference calculation based on the second prompt word, and adapts the data samples in the initial data set to the knowledge domain to which the knowledge base belongs, so as to obtain a target data set for evaluating and testing the target service.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.