Large Model Data Generation Method, Apparatus, Device, Medium and Product

By determining the inference hop number and constructing data generation agents in large language models, the problems of high cost and poor quality of generated data in the prior art are solved, and efficient and accurate data generation is achieved.

CN118656473BActive Publication Date: 2025-07-18JINAN INSPUR DATA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411124488.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-07-18
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

In the prior art, two models are required to generate data for large language models, resulting in high training costs, difficult and poor generation quality.

Method used

By determining the inference hop number of problems to be generated, the document block to be extracted is determined in the pre-constructed document block diagram model, the data generation agent is constructed, and the prompt word is input to the preset large language model to generate data.

Benefits of technology

It improves the accuracy and relevance of data generation, realizes automated and intelligent data processing, improves the efficiency and flexibility of data generation, and ensures the consistency and reliability of generated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118656473B_ABST
    Figure CN118656473B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a large model data generation method, apparatus, device, medium and product, including: determining the inference hop count corresponding to the problem to be generated; determining the document blocks to be extracted in a pre-constructed document block graph model according to the inference hop count; constructing a data generation agent based on the document blocks to be extracted, where the data generation agent includes prompt words; inputting the prompt words into a preset large language model so that the preset large language model generates generated data corresponding to the prompt words. Through the embodiment of the present invention, by determining the inference hop count of the problem to be generated, the relevant document blocks in the pre-constructed document block graph model are accurately located, and then a data generation agent including prompt words is constructed, and the prompt words are input into the preset large language model to generate high-quality generated data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and particularly relates to a method, device, equipment, medium and product for generating large model data. Background Art

[0002] A large language model is an artificial intelligence model based on a deep learning network, which is oriented to the field of natural language processing. Because the number of model parameters is generally large, it is called a large language model, and the number of its parameters generally exceeds 1 billion. Large language models have achieved good results in multiple tasks and have become a hot topic in current research. In recent years, data generation technology has been widely used in various fields. For example, in the field of database testing, the generated data is often used as the database content.

[0003] In the related art, to solve the above problems, generally two models are required. And before generating evaluation question and answer pair data in different fields, both models need to be retrained, which is costly and difficult to train, easily resulting in poor generation quality. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method, device, equipment, medium and product for generating large model data. The specific technical solutions are as follows:

[0005] In the first aspect of the present invention, a method for generating large model data is first provided. The method includes: determining the inference hop count corresponding to the question to be generated;

[0006] Determining the document block to be extracted in the pre-constructed document block graph model according to the inference hop count;

[0007] Constructing a data generation agent according to the document block to be extracted, where the data generation agent includes prompt words;

[0008] Inputting the prompt words into a preset large language model so that the preset large language model generates generated data corresponding to the prompt words.

[0009] Optionally, the constructing a data generation agent according to the document block to be extracted includes:

[0010] Constructing a data generation agent according to the document block to be extracted and a prompt word template, where the prompt word template includes instructions, generation rules, input specifications, output specifications, and output examples.

[0011] Optionally, the constructing a data generation agent according to the document block to be extracted and a prompt word template includes:

[0012] Constructing a data generation agent according to the document block to be extracted and a prompt word template;

[0013] Determine a prompt based on a pre-constructed prompt template, where the prompt is used to define the number of inference hops.

[0014] Optionally, before the step of constructing a data generation agent according to the document block to be extracted and the prompt template, the method includes:

[0015] Set the prompt template according to a preset rule.

[0016] Optionally, setting the prompt template according to a preset rule includes:

[0017] Set the prompt template to generate at least two questions.

[0018] Optionally, setting the prompt template according to a preset rule includes:

[0019] Set the data generation theme corresponding to the prompt template.

[0020] Optionally, setting the prompt template according to a preset rule includes:

[0021] Set the data generation format corresponding to the prompt template, where the data generation format is determined based on an output example.

[0022] Optionally, before the step of determining the number of inference hops corresponding to the question to be generated, the method includes:

[0023] Structurally process a preset document according to a preset data model to generate a standard document;

[0024] Perform a chunking process on the standard document to obtain document chunks;

[0025] Perform a filtering process on the document chunks to obtain target document chunks;

[0026] Construct a document chunk graph model based on the target document chunks.

[0027] Optionally, the structurally processing a preset document according to a preset data model to generate a standard document includes:

[0028] Construct a preset data model according to the title information and chapter information;

[0029] Structurally process the preset document according to the title information and the chapter information to generate a standard document, where the standard document includes at least one piece of chapter information and / or title information.

[0030] Optionally, the performing a chunking process on the standard document to obtain document chunks includes:

[0031] Perform a hierarchical chunking process on the standard document to obtain a number of document chunks.

[0032] Optionally, the hierarchical chunking of the standard document is performed to obtain several document chunks, including:

[0033] Chunk the standard document according to the title information and / or section information to obtain the first title document chunk;

[0034] Determine whether the first title document chunk is greater than a preset value;

[0035] If so, continue to chunk the first title document chunk according to the title information and / or the section information to obtain the second title document chunk;

[0036] If the second title document chunk is greater than the preset value, chunk the second title document chunk according to the preset semantic similarity rule to obtain several document chunks.

[0037] Optionally, if the second title document chunk is greater than the preset value, chunk the second title document chunk according to the preset semantic similarity rule to obtain several document chunks, including:

[0038] If the second title document chunk is greater than the preset value, chunk the second title document chunk according to the preset semantic similarity rule to obtain semantic document chunks, and configure structural theme information in the semantic document chunks to obtain several document chunks;

[0039] The structural theme information includes the title information and / or the section information.

[0040] Optionally, several of the target document blocks include at least one parent level and at least one sub level. The constructing a document block graph model based on the target document blocks includes:

[0041] Construct a document block graph model according to the hierarchical relationship between the target document block corresponding to the parent level and the target document block corresponding to the sub level.

[0042] Optionally, after the step of inputting the prompt into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt, the method includes:

[0043] Determine whether the generated data meets a preset evaluation criterion based on the generated data;

[0044] If not, correct the generation rule and output specification corresponding to the prompt until the generated data meets the preset evaluation criterion.

[0045] Optionally, after the step of inputting the prompt into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt, the method further includes:

[0046] Retrieve based on the generated data to obtain a retrieval result, and store the retrieval result in a preset database.

[0047] Optionally, after the step of retrieving based on the generated data to obtain a retrieval result and storing the retrieval result in a preset database, the method includes:

[0048] Verify the generated data based on the retrieval result;

[0049] If the generated data is in a verified success state, use the generated data as the target data set.

[0050] In a second aspect of the implementation of the present invention, there is also provided a large model data generation device, which is applied to a storage control chip in a device monitoring system. An on-chip cache pool is installed on the storage control chip. The device includes:

[0051] A first determination module, configured to determine the number of inference hops corresponding to the problem to be generated;

[0052] A second determination module, configured to determine the document block to be extracted in a pre-constructed document block graph model according to the number of inference hops;

[0053] An agent construction module, configured to construct a data generation agent according to the document block to be extracted, and the data generation agent includes a prompt;

[0054] A data generation module, configured to input the prompt into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt.

[0055] In a third aspect of the implementation of the present invention, there is also provided a communication device, including: a transceiver, a memory, a processor, and a program stored on the memory and executable on the processor;

[0056] The processor is configured to read the program in the memory to implement the large model data generation method according to any one of the first aspects.

[0057] In a fourth aspect of the implementation of the present invention, there is also provided a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computer, the computer is enabled to implement the large model data generation method according to any one of the first aspects.

[0058] In the fourth aspect of the implementation of the present invention, a computer program product is further provided, including a computer program / instructions, which when executed by a processor, implement the large model data generation method as described in any one of the first aspect.

[0059] The large model data generation method provided by the embodiments of the present invention includes: determining the inference hop count corresponding to the problem to be generated; determining the document blocks to be extracted in the pre-constructed document block graph model according to the inference hop count; constructing a data generation agent according to the document blocks to be extracted, where the data generation agent includes prompt words; inputting the prompt words into a preset large language model so that the preset large language model generates the generated data corresponding to the prompt words. By determining the inference hop count of the problem to be generated, the relevant document blocks in the pre-constructed document block graph model are accurately located, and then a data generation agent including prompt words is constructed, and the prompt words are input into the preset large language model to generate high-quality generated data. This process not only improves the accuracy and relevance of data generation, but also realizes automated and intelligent data processing through the construction of the agent and the application of the preset large language model, greatly improving the efficiency and flexibility of data generation, and at the same time ensuring the consistency and reliability of the generated data. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0061] Figure 1 The step flow of the large model data generation method provided by the embodiments of the present invention Figure 1 ;

[0062] Figure 2 is the step flow of the large model data generation method provided by the embodiments of the present invention Figure 2 ;

[0063] Figure 3 is the step flow of the large model data generation method provided by the embodiments of the present invention Figure 3 ;

[0064] Figure 4 is the block diagram of the large model data generation device provided by the embodiments of the present invention;

[0065] Figure 5 is a schematic diagram of a communication device provided by the embodiments of the present invention;

[0066] Figure 6 is a schematic diagram of an exemplary large model data generation process provided by the embodiments of the present invention;

[0067] Figure 7It is an exemplary hierarchical block diagram provided by an embodiment of the present invention;

[0068] Figure 8 It is an exemplary data block extraction diagram provided by an embodiment of the present invention. Detailed implementation manners

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each implementation manner of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each implementation manner of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following implementation manners, the technical solutions claimed in the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation manners of the present invention. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.

[0070] Referring to Figure 1 , a step flow of a large model data generation method provided by an embodiment of the present invention is shown Figure 1 , the method may include:

[0071] Step 101, determining the inference hop count corresponding to the problem to be generated;

[0072] It should be noted that in the embodiments of the present application, data generation is performed based on the inference hop count. One hop is one inference step. For example, when calculating "1 + 2 * 3", it is necessary to first calculate "2 * 3" and then calculate "1 + 6", so it is a 2-hop inference problem.

[0073] Step 102, determining the document blocks to be extracted in a pre-constructed document block graph model according to the inference hop count;

[0074] Therefore, before generating data, it is necessary to obtain different data blocks according to the hop count of the problem to be generated. As Figure 8 shown, when generating a 1-hop problem, only one data block needs to be extracted. When generating a 2-hop problem, one data block and its parent node data block need to be extracted.

[0075] It should be noted that in the embodiments of the present application, the pre-constructed document block graph model is constructed based on the hierarchical relationship between data blocks. Specifically, reference can be made to the content described later Figure 2 .

[0076] In the embodiments of the present application, by generating data based on multi-hop reasoning, complex multi-hop reasoning problems can be generated.

[0077] Step 103: Construct a data generation agent according to the document block to be extracted, where the data generation agent includes prompt words.

[0078] Further, the constructing a data generation agent according to the document block to be extracted includes:

[0079] Construct a data generation agent according to the document block to be extracted and a prompt word template, where the prompt word template includes instructions, generation rules, input specifications, output specifications, and output examples.

[0080] It should be noted that in the embodiments of the present application, reference may be made to Figure 6 As shown, based on the previously determined document blocks to be extracted, a data generation agent can be constructed, and the data generation agent can include instructions, generation rules, input specifications, output specifications, and output examples.

[0081] Essentially, the data generation agent is a prompt word generation template based on an agent, and the purpose is to construct an initial prompt word template.

[0082] Construction of the data generation agent:

[0083] Document blocks to be extracted: First, according to the previously determined document blocks to be extracted, these blocks contain the necessary information and data for constructing the data generation agent.

[0084] Agent components: The data generation agent includes several key components, such as instructions, generation rules, input specifications, output specifications, and output examples. These components together define how the agent processes and generates content.

[0085] Functions of the agent:

[0086] Instructions: Define the basic operations and objectives of the agent, and guide how the agent executes tasks.

[0087] Generation rules: Specify the specific rules and logic for content generation, ensuring that the generated content meets specific standards and requirements.

[0088] Input specifications: Define the input formats and types accepted by the agent, ensuring the correctness and usability of the input data.

[0089] Output specifications: Define the content formats and standards of the agent's output, ensuring that the output content is clear, accurate, and easy to understand.

[0090] Output examples: Provide some actual output examples to help users and developers better understand the expected form and quality of the output.

[0091] For example, referring to the following example, an agent-based prompt generation template can be filled with appropriate content as needed. It should be noted that the comments in the template are not included in the prompt template.

[0092] {

[0093] # Instruction Clear instruction

[0094] You are an expert in {{domain}}. Generate #M questions that require #N-hop(s) of reasoning to solve based on the given content.

[0095] ## Rules Generation rules

[0096] 1. Questions must be directly related to and answerable using the provided content.

[0097] 2. Each question should require exactly #N hops of reasoning. A hop is defined as a step from one intermediate result to another. Example: For "1 + 2 * 3", we have 2 hops:

[0098] - Hop 1: 2 * 3 = 6

[0099] - Hop 2: 1 + 6 = 7 # Generate the content of how many hops and clearly define what a hop is

[0100] 3. Ask specific questions about the content, avoiding general queries like "What is the key information in the given paragraph?".

[0101] 4. Ensure questions are evenly distributed across the provided content chunks. When generating multiple questions, the questions should be in different content areas

[0102] 5. [Optional] Ensure that each question is related to at least one of the given categories. # Optional, generate content according to the specified topics

[0103] ## Content The provided data chunks

[0104] [Multiple content chunks for #N-hop(s) question generation]

[0105] 1. Content from chunk 1

[0106] 2. Content from chunk 2 ...

[0107] ## Content Categories Optional, content topic classification

[0108] [List of content categories]

[0109] ## Output Format Required output format

[0110] For each question, provide:

[0111] 1. Question text

[0112] 2. Answer

[0113] 3. Detailed explanation of the #N-hop(s) reasoning process

[0114] 4. [Optional] Category

[0115] ## Generation Process Generation process

[0116] Initial Generation

[0117] Generate #M questions adhering to the rules above.

[0118] Self-Review, self-propose modification suggestions, the core of the agent

[0119] Critically evaluate your initial generation:

[0120] }

[0121] Furthermore, constructing a data generation agent according to the to-be-extracted document block and the prompt template includes:

[0122] Construct a data generation agent according to the to-be-extracted document block and the prompt template;

[0123] Determine a prompt based on a pre-constructed prompt template, where the prompt is used to define the number of reasoning hops.

[0124] It should be noted that in the embodiments of the present application, the composition of the prompt template includes instructions, rules (requirements), input, output specifications, and output examples.

[0125] The core of this template is that the prompt clearly gives the definition of one hop, enabling the large model to generate complex questions with multi-hop reasoning. Without a clear definition, the questions generated by the large model are very simple questions that do not contain complex reasoning processes. Moreover, the agent mechanism enables the large model to generate data according to the process of "initial generation - reflection - optimization", and this self-reflection mechanism can improve the accuracy and quality of the generated data.

[0126] Furthermore, before the step of constructing a data generation agent according to the to-be-extracted document block and the prompt template, the method includes:

[0127] Set the prompt template according to preset rules.

[0128] Furthermore, setting the prompt template according to preset rules includes:

[0129] Set the prompt template to generate at least two questions.

[0130] Furthermore, setting the prompt template according to preset rules includes:

[0131] Set the data generation theme corresponding to the prompt template.

[0132] Further, the setting of the prompt template according to the preset rules includes:

[0133] Set the data generation format corresponding to the prompt template, and the data generation format is determined based on the output example.

[0134] It should be noted that in the embodiments of the present application, in order to improve the accuracy and relevance of data generation, the embodiments of the present application optimize the prompt template. Specifically, it may include the following contents:

[0135] Generate multiple questions: If only one question is generated, the Q&A data generated by the large model is often a conceptual question, such as "What is XX?". When multiple questions are generated, the generation quality will be improved, such as "Which technology can be used to achieve XX?".

[0136] Generate according to the theme: Giving a definite generation theme is equivalent to limiting the scope of data generation, which can avoid generating irrelevant content and improve the generation accuracy.

[0137] Control the format of the generated data according to the example: Different types of data can be generated according to needs, such as Q&A pairs, short answers, judgments, etc. Giving specific examples generates more accurately than just requiring a specific format without providing examples. Among them, the format of the output data example can be in JSON format, or other formats can be generated according to needs. Specifically, the following content can be referred to, that is, to generate a data example, the output format can be controlled using the example:

[0139] {

[0140] "Question": "Which native runtime complies with the Open Container Initiative (OCI) standard? ",

[0141] "Option_A": "runC",

[0142] "Option_B": "runV",

[0143] "Option_C": "kata-containers",

[0144] "Option_D": "gvisor",

[0145] "Answer": "A"

[0146] },

[0147] {

[0148] "Question": "Which API object is the recommended way to run scalable, stateless applications on a cluster? "​

[0149] "Option_A": "ReplicaSet",

[0150] "Option_B": "Deployment",

[0151] "Option_C": "DaemonSet",

[0152] "Option_D": "Pod",

[0153] "Answer": "B"

[0154] },

[0155] {

[0156] "Question": "The user schedules a CronJob to run once an hour. When it's time for this CronJob to run, what happens in the cluster?",

[0157] "Option_A": "The Kubelet watches the API server for CronJob objects. When it's time to run the Job, it runs the Pod directly.",

[0158] "Option_B": "The Kube-scheduler watches the API server for CronJob objects, which is why it's called the kube-scheduler.",

[0159] "Option_C": "The CronJob controller component creates a Pod and waits for it to finish running.",

[0160] "Option_D": "The CronJob controller component creates a Job. Then the Job controller creates a Pod and waits for it to finish running.",

[0161] "Answer": "D"

[0162] }

[0164] It should be noted that in the embodiments of the present application, a self-feedback intelligent agent mechanism is used to generate data, improving the quality of the generated data.

[0165] Step 104, input the prompt into a preset large language model so that the preset large language model generates generated data corresponding to the prompt. ​

[0166] It should be noted that in the embodiments of the present application, each part in the prompt template in the data generation agent can be assembled into a prompt, and then the prompt is input into a preset large language model so that the preset large language model generates generated data corresponding to the prompt.

[0167] Specifically, the process of assembling a set of prompts involves integrating each component in the prompt template to form a clear and specific set of instructions, which will guide the large language model to generate the required content. The following is an exemplary and detailed explanation of how to assemble prompts according to the composition of the prompt template:

[0168] Instruction:

[0169] Define the goal: First, clarify the goal and task of the generated content. The instruction should be concise and clear, directly pointing to the core purpose of the generated content.

[0170] Example: "Generate an article about the application of artificial intelligence in the medical field."

[0171] Rule (requirement):

[0172] Set standards: List the specific rules and requirements for content generation, such as word count limits, style requirements, format specifications, etc.

[0173] Example: "The article length should be between 800 and 1000 words, and it should adopt a formal scientific and technological article style."

[0174] Input:

[0175] Provide background information: Provide the necessary background information or data to help the model better understand the context and requirements.

[0176] Example: "The input includes the latest development of AI medical technology, relevant case studies, and expert opinions."

[0177] Output specification:

[0178] Define the format: Define the specific format and structure of the output, such as title, paragraph distribution, citation format, etc.

[0179] Example: "The article should include an introduction, a body (divided into three parts: technology development, case analysis, and future outlook), and a conclusion."

[0180] Output sample:

[0181] Provide a reference: Provide one or more output samples to help the model understand the expected output form and quality.

[0182] Example: "Reference sample: [link or text]"

[0183] The process of assembling the prompt words is to integrate the above-mentioned various parts into a coherent set of instructions.

[0184] The large model data generation method provided by the embodiments of the present invention includes: determining the number of reasoning hops corresponding to the problem to be generated; determining the document blocks to be extracted in a pre-constructed document block graph model according to the number of reasoning hops; constructing a data generation agent according to the document blocks to be extracted, the data generation agent including prompt words; and inputting the prompt words into a preset large language model so that the preset large language model generates generated data corresponding to the prompt words. By determining the number of reasoning hops of the problem to be generated in the embodiments of the present invention, the relevant document blocks in the pre-constructed document block graph model are accurately located, and then a data generation agent including prompt words is constructed, and the prompt words are input into the preset large language model to generate high-quality generated data. This process not only improves the accuracy and relevance of data generation, but also realizes automated and intelligent data processing through the construction of the agent and the application of the preset large language model, greatly improving the efficiency and flexibility of data generation, and at the same time ensuring the consistency and reliability of the generated data.

[0185] Refer to Figure 2 , which shows the step flow of the large model data generation method provided by the embodiments of the present invention Figure 2 , the method may include:

[0186] Step 201, perform structured processing on a preset document according to a preset data model to generate a standard document;

[0187] Further, the performing structured processing on a preset document according to a preset data model to generate a standard document includes:

[0188] Construct a preset data model according to the title information and chapter information;

[0189] Perform structured processing on the preset document according to the title information and the chapter information to generate a standard document, the standard document including at least one piece of chapter information and / or title information.

[0190] It should be noted that in the embodiments of the present application, a unified preset data model can be constructed and the document can be subjected to structured processing.

[0191] Specifically, the preset data model can be expressed as {Title, Header 1, Header 2, Header 3, Header 4, Header 5}, where Title is the title information and Header is the chapter information. Therefore, the preset document is subjected to structured processing based on the title information and / or chapter information to generate a standard document.

[0192] Since the Markdown format can well support the above data model, various preset documents, including PPTs, web pages, Word documents, PDFs, etc., can be automatically segmented into a Markdown structure. Then, after preprocessing, all documents will be organized into structured documents.

[0193] Therefore, the present invention is different from other methods that directly generate data using large models. The present invention generates data based on document facts, which can greatly improve the accuracy and relevance of data generation. And by using a unified data model to process different types of files, they can be processed into unified structured data.

[0194] Step 202: Perform chunking on the standard document to obtain document chunks.

[0195] Further, the performing chunking on the standard document to obtain document chunks includes:

[0196] Perform hierarchical chunking on the standard document to obtain a number of document chunks.

[0197] It should be noted that after processing the preset document, a standard document can be obtained, and then the standard document can be chunked to obtain document chunks.

[0198] Specifically, in the embodiments of the present application, it is implemented through hierarchical data chunking.

[0199] Further, the performing hierarchical chunking on the standard document to obtain a number of document chunks includes:

[0200] Chunk the standard document according to the title information and / or chapter information to obtain first-title document chunks.

[0201] Determine whether the first-title document chunks are greater than a preset value.

[0202] If so, continue to chunk the first-title document chunks according to the title information and / or the chapter information to obtain second-title document chunks.

[0203] If the second-title document chunks are greater than the preset value, chunk the second-title document chunks according to the preset semantic similarity rule to obtain a number of document chunks.

[0204] Further, the if the second-title document chunks are greater than the preset value, chunk the second-title document chunks according to the preset semantic similarity rule to obtain a number of document chunks includes:

[0205] If the second title document block is larger than the preset value, perform chunking processing on the second title document block according to the preset semantic similarity rule to obtain semantic document chunks, and configure structural topic information in the semantic document chunks to obtain a number of document chunks;

[0206] The structural topic information includes the title information and / or the chapter information.

[0207] It should be noted that in the embodiments of the present application, the hierarchical data chunking includes 2 steps. Specifically, the first level is the structure-priority division.

[0208] Among them, the structure-priority division includes: first, divide according to the document structure, that is, divide according to the title information and / or chapter information in the document structure. For example, first chunk by the Header 1 (usually a chapter) level, and it can be divided into 5 chapters. If the chunk size of each chapter is less than the set value, the whole chunk is retained. If the chunk size of each chapter exceeds the set value, then recursively divide according to the next level Header 2 (usually a sub-chapter).

[0209] The second level is the semantic similarity division. By analogy with the above steps, if the chunk size still exceeds the set value after dividing to the last chapter, then divide according to semantic similarity to ensure that the size of each chunk is less than the set value. At the same time, in each chunk, structured topic information such as Title, Header 1, and Header 2 will also be added to ensure that each chunk is knowledge-complete.

[0210] Specifically, the above steps can be referred to Figure 7 , and the hierarchical chunking example can be seen, such as Figure 7 Example: The document is divided into 2 chunks, and each chunk contains structured title, chapter 1, and sub-chapter topic information. This processing will ensure that each chunk contains complete knowledge points and corresponding topics, rather than simply truncating like the fixed-length division.

[0211] In addition, in the embodiments of the present application, through the hierarchical data chunking method, the knowledge point structure of the document chunks can be retained.

[0212] Step 203, perform filtering processing on the document chunks to obtain target document chunks;

[0213] It should be noted that in the embodiments of the present application, after obtaining the document chunks, it can be known that at this time, the document chunks will contain some unimportant information, resulting in the inability to generate a high-quality professional data set. Therefore, the document chunks can be filtered to obtain target document chunks.

[0214] Specifically, the large model can be used first to judge the professionalism of the chunk information and remove the chunks with weak professionalism. For example, if the question is to know the version content of Kubernetes, but the document only involves the version number in the version content, then generating Q&A pairs based on this information will generate information about the version number, which is quite different from the professional information of Kubernetes. Therefore, such document chunks can be filtered first to ensure the quality of the remaining target document chunks.

[0215] Step 204: Construct a document chunk graph model based on the target document chunks.

[0216] Furthermore, several of the target document chunks include at least one parent level and at least one child level. The constructing of the document chunk graph model based on the target document chunks includes:

[0217] Construct a document chunk graph model according to the hierarchical relationship between the target document chunks corresponding to the parent level and the target document chunks corresponding to the child level.

[0218] It should be noted that in the embodiments of the present application, in order to reflect the association relationship between each data chunk, therefore, it is realized by constructing a document chunk graph model for the target document chunks. Professional data that is more complex and requires multi-hop reasoning can be generated based on the association relationship between the document chunks.

[0219] The specific graph construction method can be to point the chunks at the child level to the chunks at the parent level according to the Header level.

[0220] Step 205: Determine the number of reasoning hops corresponding to the question to be generated.

[0221] It should be noted that in the embodiments of the present application, data generation is performed based on the number of reasoning hops. One hop is one reasoning step. For example, when calculating "1 + 2 * 3", it is necessary to calculate "2 * 3" first and then calculate "1 + 6", so it is a 2-hop reasoning problem.

[0222] Step 206: Determine the document chunks to be extracted in the pre-constructed document chunk graph model according to the number of reasoning hops.

[0223] Therefore, before generating data, different data chunks need to be obtained according to the number of hops of the question to be generated. As Figure 8 shown, when generating a 1-hop question, only one data chunk needs to be extracted. When generating a 2-hop question, one data chunk and its parent node data chunk need to be extracted.

[0224] Step 207: Construct a data generation agent according to the document chunks to be extracted. The data generation agent includes prompt words.

[0225] Step 208: Input the prompt into a preset large language model so that the preset large language model generates generated data corresponding to the prompt.

[0226] It should be noted that in the embodiments of the present application, steps 207-208 are as described in the previous discussion and will not be elaborated here.

[0227] In the embodiments of the present invention, by determining the inference hop count of the problem to be generated, relevant document blocks in the pre-constructed document block graph model are accurately located. Then, a data generation intelligent agent containing the prompt is constructed, and the prompt is input into a preset large language model to generate high-quality generated data. This process not only improves the accuracy and relevance of data generation but also realizes automated and intelligent data processing through the construction of the intelligent agent and the application of the preset large language model, greatly enhancing the efficiency and flexibility of data generation while ensuring the consistency and reliability of the generated data.

[0228] Refer to Figure 3 which shows the step flow of the large model data generation method provided by the embodiments of the present invention Figure 3 The method may include:

[0229] Step 301: Determine the inference hop count corresponding to the problem to be generated;

[0230] Step 302: Determine the document blocks to be extracted in the pre-constructed document block graph model according to the inference hop count;

[0231] Step 303: Construct a data generation intelligent agent according to the document blocks to be extracted, and the data generation intelligent agent includes the prompt;

[0232] Step 304: Input the prompt into a preset large language model so that the preset large language model generates generated data corresponding to the prompt;

[0233] It should be noted that in the embodiments of the present application, steps 301-304 are as described in the previous discussion and will not be elaborated here.

[0234] Step 305: Determine whether the generated data meets the preset evaluation criteria based on the generated data;

[0235] Step 306: If not, correct the generation rule and output specification corresponding to the prompt until the generated data meets the preset evaluation criteria.

[0236] Based on the data generated by the large model, evaluate whether the generated data meets the requirements, and continuously iterate the prompt. Generally, only the rule and output parts need to be modified until the requirements are met.

[0237] Step 307: Retrieve based on the generated data, obtain the retrieval results, and store the retrieval results in a preset database.

[0238] Step 308: Verify the generated data based on the retrieval results;

[0239] Use the generated data as a search term to retrieve relevant content from knowledge bases such as domain databases and the Internet (optional). The specific operations are as follows:

[0240] First, select a search engine or database: According to the domain in which the data is generated, select a suitable search engine (such as Baidu, Bing) or domain-specific database (such as PubMed, IEEE Xplore).

[0241] Second, use the API for keyword search: Utilize the API of the search engine (such as Baidu Smart Cloud Search API, Bing Search API) to retrieve the data generated by the large model as keywords.

[0242] Third, obtain the most relevant content: Retrieve the first few most relevant items from the search results, usually the first 5 to 10 items.

[0243] Fourth, collect the retrieval results: Save all relevant retrieval results, including web links, documents, papers, etc., and use them as the accurate content for subsequent verification.

[0244] Step 309: If the generated data is in a verified successful state, use the generated data as the target data set.

[0245] Based on the content retrieved in steps 307 - 308, verify the data generated by the large model to check whether the generated data is trustworthy. The specific operations are as follows:

[0246] First, input the retrieved content and the generated data into the large model together: Select another large model (different from the large model that generated the data), and input the most relevant content obtained in step seven and the generated data into this model.

[0247] Second, the large model makes a content determination: Request the large model to determine the generated data to confirm its accuracy and trustworthiness. The large model will judge whether the generated data conforms to the retrieved information based on the input retrieved content and generated data.

[0248] Third, filter out untrustworthy data: If the large model determines that the generated data is inaccurate, discard the generated content. If the large model determines that the data is accurate, retain the generated content.

[0249] Fourth, organize reliable data: Organize the generated data determined to be accurate by the large model into the final dataset, only retaining the content determined to be accurate, and prepare for the next step of human expert review.

[0250] Fifth, record the verification process: Record the verification process of each piece of data, including the retrieved results used, the determination results of the large model and its judgments, for traceability and further optimization.

[0251] In addition, in the embodiments of the present application, human experts are used to review the dataset after verification in step eight again to filter out the content that does not meet the requirements and generate the final dataset. Since a large amount of datasets have been generated using the large model, only a small amount of manual review is required in this step, so the amount of human participation in the overall process is small.

[0252] In addition, in the embodiments of the present application, an automatic verification will be performed after the data is generated. The data source for verification can be any data source to ensure the accuracy of the generated data.

[0253] In addition, in the embodiments of the present application, the present invention introduces a caching system, which can be cached in the local system during the generation of network requests and the process of calling the large language model. This method only needs to be re-run in case of interruption (such as caused by network errors) or errors. It saves repeated access to the network and large model requests, thus saving time.

[0254] In the embodiments of the present invention, by determining the inference hop count of the problem to be generated, the relevant document blocks in the pre-constructed document block graph model are accurately located, and then a data generation intelligent agent containing prompt words is constructed, and the prompt words are input into the preset large language model to generate high-quality generated data. This process not only improves the accuracy and relevance of data generation, but also realizes automated and intelligent data processing through the construction of the intelligent agent and the application of the preset large language model, greatly improving the efficiency and flexibility of data generation, and at the same time ensuring the consistency and reliability of the generated data.

[0255] In addition, the embodiments of the present invention can clean out high-quality training data: A large amount of high-quality data is required during the training and fine-tuning of the large model. The quality of the existing data is mixed, and the present invention can clean out high-quality data. The embodiments of the present invention can automatically generate data in a specific format: During the fine-tuning of the large model, data in a specific format is generally required, such as in the form of question-and-answer pairs. The present invention can generate any form of data according to needs. The present invention can be used to generate question-and-answer data pairs and can also be used as an evaluation dataset. The embodiments of the present invention can generate accurate data: The present invention generates content based on the domain database and uses the method of automatic verification and a small amount of manual review after data generation to ensure the accuracy of data generation.

[0256] Refer toFigure 4 , showing a schematic structural diagram of a large model data generation device provided by an embodiment of the present invention. The device includes:

[0257] A first determination module 401, configured to determine the inference hop count corresponding to the problem to be generated;

[0258] A second determination module 402, configured to determine the document blocks to be extracted in a pre-constructed document block graph model according to the inference hop count;

[0259] An agent construction module 403, configured to construct a data generation agent according to the document blocks to be extracted. The data generation agent includes prompt words;

[0260] A data generation module 404, configured to input the prompt words into a preset large language model, so that the preset large language model generates generated data corresponding to the prompt words.

[0261] In the embodiment of the present invention, by determining the inference hop count of the problem to be generated, the relevant document blocks in the pre-constructed document block graph model are accurately located. Furthermore, a data generation agent including prompt words is constructed, and the prompt words are input into the preset large language model to generate high-quality generated data. This process not only improves the accuracy and relevance of data generation, but also realizes automated and intelligent data processing through the construction of the agent and the application of the preset large language model, greatly improving the efficiency and flexibility of data generation, and at the same time ensuring the consistency and reliability of the generated data.

[0262] The embodiment of the present invention also provides a communication device, as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 complete communication with each other through the communication bus 504.

[0263] The memory 503 is used to store computer programs;

[0264] When the processor 501 is configured to execute the programs stored on the memory 503, the following steps can be implemented:

[0265] Determine the inference hop count corresponding to the problem to be generated;

[0266] Determine the document blocks to be extracted in a pre-constructed document block graph model according to the inference hop count;

[0267] Construct a data generation agent according to the document blocks to be extracted. The data generation agent includes prompt words;

[0268] Input the prompt into a preset large language model so that the preset large language model generates generated data corresponding to the prompt.

[0269] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on the transmission medium. The data processed by the processor can be transmitted through a wired medium or transmitted over a wireless medium through an antenna. Further, the antenna also receives data and transmits the data to the processor. The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.

[0270] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0271] The communication interface is used for communication between the above terminal and other devices.

[0272] The memory can include a Random Access Memory (RAM), or can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0273] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0274] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the large model data generation method described in any one of the above embodiments.

[0275] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute the large model data generation method described in any one of the above embodiments.

[0276] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a Solid State Disk (SSD)).

[0277] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0278] Each embodiment in this specification is described in a related manner. For the identical and similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0279] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for generating large model data, characterized in that, The method includes: Determine the inference hop count corresponding to the problem to be generated; Determine the document blocks to be extracted in the pre-constructed document block graph model according to the inference hop count; Construct a data generation agent according to the document blocks to be extracted, where the data generation agent includes prompt words; Input the prompt words into a preset large language model so that the preset large language model generates generated data corresponding to the prompt words; The constructing a data generation agent according to the document blocks to be extracted includes: Construct a data generation agent according to the document blocks to be extracted and a prompt word template, where the prompt word template includes instructions, generation rules, input specifications, output specifications, and output examples; The determining the document blocks to be extracted in the pre-constructed document block graph model according to the inference hop count includes: Determine data chunks and the parent node data chunks corresponding to the data chunks in the pre-constructed document block graph model according to the inference hop count; Among them, the pre-constructed document block graph model is constructed based on the hierarchical relationship between data chunks; The constructing a data generation agent according to the document blocks to be extracted and a prompt word template includes: Construct a data generation agent according to the document blocks to be extracted and a prompt word template; Determine prompt words based on the pre-constructed prompt word template, where the prompt words are used to define the inference hop count; Before the step of determining the inference hop count corresponding to the problem to be generated, the method includes: Perform structured processing on a preset document according to a preset data model to generate a standard document; Perform chunking processing on the standard document to obtain document chunks, including: performing hierarchical processing of first-level structure priority division and second-level semantic similarity division on the standard document to obtain document chunks; Perform filtering processing on the document chunks to obtain target document chunks; Construct a document block graph model based on the target document chunks.

2. The method according to claim 1, wherein Before the step of constructing a data generation agent according to the document blocks to be extracted and a prompt word template, the method includes: Set the prompt word template according to preset rules.

3. The method according to claim 2, wherein The setting the prompt word template according to preset rules includes: Set the prompt word template to generate at least two questions.

4. The method according to claim 2, wherein The setting the prompt word template according to preset rules includes: Set the data generation theme corresponding to the prompt word template.

5. The method according to claim 2, wherein The setting the prompt word template according to preset rules includes: Set the data generation format corresponding to the prompt word template, where the data generation format is determined based on the output example.

6. The method according to claim 1, wherein The performing structured processing on a preset document according to a preset data model to generate a standard document includes: Construct a preset data model according to the title information and chapter information; Perform structured processing on the preset document according to the title information and the chapter information to generate a standard document, where the standard document includes at least one piece of chapter information and / or title information.

7. The method according to claim 1, characterized in that The performing chunking processing on the standard document to obtain document chunks includes: Perform hierarchical chunking processing on the standard document to obtain a number of document chunks.

8. The method according to claim 1, characterized in that, The performing hierarchical chunking processing on the standard document to obtain a number of document chunks includes: Perform chunking processing on the standard document according to the title information and / or chapter information to obtain first-title document chunks; Determine whether the first title document block is greater than a preset value; If so, continue to perform chunking processing on the first title document block according to the title information and / or the section information to obtain a second title document block; If the second title document block is greater than the preset value, perform chunking processing on the second title document block according to a preset semantic similarity rule to obtain a plurality of document chunks.

9. The method according to claim 8, wherein The "if the second title document block is greater than the preset value, perform chunking processing on the second title document block according to a preset semantic similarity rule to obtain a plurality of document chunks" includes: If the second title document block is greater than the preset value, perform chunking processing on the second title document block according to a preset semantic similarity rule to obtain semantic document chunks, and configure structural theme information in the semantic document chunks to obtain a plurality of document chunks; The structural theme information includes the title information and / or the section information.

10. The method according to claim 1, wherein A plurality of the target document blocks include at least one parent level and at least one sub level. The constructing a document block graph model based on the target document blocks includes: Constructing a document block graph model according to the hierarchical relationship between the target document block corresponding to the parent level and the target document block corresponding to the sub level.

11. The method according to claim 1, wherein After the step of inputting the prompt word into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt word, the method includes: Determining whether the generated data meets a preset evaluation criterion based on the generated data; If not, correct the generation rule and output specification corresponding to the prompt word until the generated data meets the preset evaluation criterion.

12. The method according to claim 11, wherein After the step of inputting the prompt word into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt word, the method further includes: Performing a search based on the generated data, obtaining a search result, and storing the search result in a preset database.

13. The method according to claim 12, wherein After the step of performing a search based on the generated data, obtaining a search result, and storing the search result in a preset database, the method includes: Verifying the generated data based on the search result; If the generated data is in a verified successful state, use the generated data as a target data set.

14. A large language model data generation device, characterized in that, The device includes: A first determination module, configured to determine an inference hop count corresponding to a question to be generated; A second determination module, configured to determine a document block to be extracted in a pre-constructed document block graph model according to the inference hop count; An agent construction module, configured to construct a data generation agent according to the document block to be extracted, where the data generation agent includes a prompt word; A data generation module, configured to input the prompt word into a preset large language model to enable the preset large language model to generate generated data corresponding to the prompt word; The constructing a data generation agent according to the document block to be extracted includes: Constructing a data generation agent according to the document block to be extracted and a prompt word template, where the prompt word template includes an instruction, a generation rule, an input specification, an output specification, and an output example; Determining the document block to be extracted in the pre-constructed document block graph model according to the inferred hop count includes: Determining data chunks and the corresponding parent node data chunks of the data chunks in the pre-constructed document block graph model according to the inferred hop count; Among them, the pre-constructed document block graph model is constructed based on the hierarchical relationship between data chunks; Constructing a data generation agent according to the document block to be extracted and the prompt template includes: Constructing a data generation agent according to the document block to be extracted and the prompt template; Determining a prompt word based on the pre-constructed prompt template, where the prompt word is used to define the inferred hop count; Before the step of determining the inferred hop count corresponding to the question to be generated, the device is further configured to: Structurally process a preset document according to a preset data model to generate a standard document; perform chunking processing on the standard document to obtain document chunks, including: performing hierarchical processing of first-level structure priority division and second-level semantic similarity division on the standard document to obtain document chunks; performing filtering processing on the document chunks to obtain target document chunks; constructing a document block graph model based on the target document chunks.

15. A communication device, characterized in that, Including: A transceiver, a memory, a processor, and a program stored on the memory and executable on the processor; The processor is configured to read the program in the memory to implement the large model data generation method described in any one of claims 1-13.

16. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the large model data generation method described in any one of claims 1-13.

17. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the large model data generation method described in any one of claims 1-13.

Citation Information

Patent Citations

  • Dialogue method and system based on document retrieval enhanced machine language model

    CN117807199A