Generation method and device of instruction data set, equipment and storage medium

By generating multiple question-and-answer types of questions and SQL sample data, filling the template based on preset rules, calculating the scores of the initial instructions, and determining the proportion distribution of the problem types, it solves the problem that it is difficult to quickly build high-quality fine-tuned instruction data sets in the existing technology, and realizes the high accuracy of the model when generating SQL statements and the diversity of the data sets.

CN119938834APending Publication Date: 2025-05-06BEIJING ZHONGJIAOXING ROAD INTERNET OF VEHICLES TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997751.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to quickly build high-quality fine-tuning instruction datasets, resulting in low accuracy of the model when generating SQL statements.

Method used

By generating multiple question-and-answer types of questions and SQL sample data based on small sample question-and-answer data, creating templates, generating card slot data based on preset rules, filling templates, calculating the scores of initial instructions, determining the proportion distribution of question types, and generating the final instruction data set.

Benefits of technology

It realizes automatic generation of high-quality fine-tuning instruction data sets based on small sample data, which improves the accuracy of the model when generating SQL statements and the diversity and coverage of the data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938834A_ABST
    Figure CN119938834A_ABST
Patent Text Reader

Abstract

The invention discloses an instruction data set generation method and device, equipment and a storage medium. Comprising the steps of generating question and SQL pair sample data of various question and answer types based on collected small sample question and answer data; generating a template for the sample data based on the question and the SQL, wherein the template comprises card slots corresponding to various data types; generating card slot data based on a preset rule, and filling a card slot in the template according to the card slot data to obtain a plurality of expanded initial instructions; the method comprises the steps of obtaining an initial instruction, calculating a score of the initial instruction, determining proportion distribution of each problem type based on the score of the initial instruction, extracting an instruction of a preset proportion type from the initial instruction according to the proportion distribution, and generating a final instruction data set. According to the method, a large number of high-quality fine-tuning data samples can be generated by utilizing limited small sample data, so that the accuracy of the model during SQL statement generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, equipment and storage medium for generating an instruction data set. Background Art

[0002] In the field of natural language processing (NLP), with the rapid development of artificial intelligence technology, large language models (LLMs) have shown excellent performance. Through deep learning and pre-training technology, these models can understand and generate natural language and handle complex language tasks. Instruction following, as a key link to support LLM training, is crucial to improving the accuracy and execution efficiency of the model. In order to improve the accuracy of instruction following, instruction fine-tuning technology can be used to train the model. Instruction fine-tuning is a technology that performs supervised training on LLMs through instruction datasets containing instructions and expected outputs, thereby improving the accuracy and execution efficiency of the model.

[0003] The traditional way of constructing instruction datasets mainly relies on manual annotation, which has many problems such as high cost, low efficiency, and uneven data quality. With the continuous improvement of the capabilities of large-scale language models, the traditional method based on manual annotation can no longer meet the needs of large-scale, high-quality instruction data. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device and storage medium for generating an instruction data set, so as to at least solve the technical problem in the related art that it is difficult to quickly construct a fine-tuning instruction data set.

[0005] According to one aspect of an embodiment of the present application, a method for generating an instruction data set is provided, comprising:

[0006] Generate sample data of questions and SQL pairs of various question-answer types based on the collected small sample question-answer data;

[0007] Generate a template for the sample data based on the question and SQL, the template including slots corresponding to multiple data types;

[0008] Generate card slot data based on preset rules, fill the card slots in the template according to the card slot data, and obtain a plurality of expanded initial instructions;

[0009] The score of the initial instruction is calculated, and the proportion distribution of each question type is determined based on the score of the initial instruction. According to the proportion distribution, instructions of a preset proportion type are extracted from the initial instruction to generate a final instruction data set.

[0010] In one embodiment, the generation of sample data of questions and SQL pairs of various question and answer types based on the collected small sample question and answer data includes:

[0011] Classifying the small sample question and answer data based on the question and answer topic type;

[0012] Classify SQL query statements based on SQL value types;

[0013] The question types and SQL value types are arranged and combined to generate sample data of questions and SQL pairs of the plurality of question-and-answer types.

[0014] In one embodiment, generating a template for sample data based on the question and SQL includes:

[0015] Decomposing the problem and SQL pair sample data to obtain multiple components, wherein the components include time, conditions, and problems;

[0016] A slot is set for each component to obtain a template containing multiple slots.

[0017] In one embodiment, the generating the card slot data based on a preset rule and filling the card slot in the template according to the card slot data includes:

[0018] Based on LLM prompt engineering technology, synonym expansion is performed on each question-answer type to generate question slot data;

[0019] Randomly generate time slot data based on the preset time format;

[0020] Randomly extract conditional fields from database fields to generate conditional slot data;

[0021] According to the problem slot data, time slot data and condition slot data, the corresponding slots in the template are filled.

[0022] In one embodiment, after obtaining the expanded multiple initial instructions, the method further includes:

[0023] One or more standardization processes of replacement, merging, and marking are performed on the words in the initial instructions to generate standardized initial instructions.

[0024] In one implementation, calculating the score of the initial instruction includes:

[0025] Input the question in the initial instruction into a preset model to obtain an SQL query statement output by the model;

[0026] Calculate the error rate of the initial instruction based on the SQL query statement corresponding to the question in the initial instruction and the SQL query statement output by the model;

[0027] A score of the initial instruction is obtained according to the error rate of the initial instruction.

[0028] In one implementation, determining the proportion distribution of each question type based on the score of the initial instruction includes:

[0029] Normalizing the scores of the initial instructions to obtain normalized instruction scores;

[0030] The instruction scores of different question types are summarized to obtain the scores of each question type;

[0031] According to the scores of the various question types, the proportion distribution of the various question types is obtained.

[0032] According to another aspect of an embodiment of the present application, a device for generating an instruction data set is provided, comprising:

[0033] The question-answer classification module is used to generate sample data of questions and SQL pairs of various question-answer types based on the collected small sample question-answer data;

[0034] A template making module, used to generate a template for sample data based on the question and SQL, wherein the template includes slots corresponding to multiple data types;

[0035] An instruction expansion module, used to generate card slot data based on preset rules, fill the card slots in the template according to the card slot data, and obtain a plurality of expanded initial instructions;

[0036] An instruction data set generation module is used to calculate the score of the initial instruction, determine the proportion distribution of each question type based on the score of the initial instruction, and extract instructions of a preset proportion type from the initial instruction according to the proportion distribution to generate a final instruction data set.

[0037] According to another aspect of an embodiment of the present application, there is further provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method for generating the instruction data set through the computer program.

[0038] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for generating an instruction data set when running.

[0039] The technical solution provided by the embodiments of the present application may have the following beneficial effects:

[0040] The embodiment of the present application provides a method for generating an instruction data set, by creating a template based on sample data of questions and SQL pairs of various question and answer types, and generating slot data based on rules to fill in the template, so as to create a rich variety of question and answer pairs, and enhance the coverage and diversity of the data set. A large number of high-quality fine-tuning instruction data sets can be automatically generated based on small sample instruction data, improving the efficiency of constructing instruction data sets and the accuracy of the model when generating SQL statements. And the present application can generate a more balanced and representative instruction data set by calculating the score of the initial instruction and determining the distribution of the proportion of question types accordingly, which helps to train a more accurate and reliable language model. It provides strong technical support for building efficient and accurate natural language processing models. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 is a flowchart of an optional method for generating an instruction data set according to an embodiment of the present application;

[0043] Figure 2 is a schematic diagram of another method for generating an instruction data set according to an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of a device for generating an instruction data set according to an embodiment of the present application;

[0045] Figure 4 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0047] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0048] At the same time, it should be understood that for the sake of ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The technologies, methods and devices known to ordinary technicians in the relevant fields may not be discussed in detail, but in appropriate cases, the technologies, methods and devices should be regarded as part of the specification.

[0049] This application mainly focuses on the difficulty of generating NL2SQL fine-tuned large model data sets in vertical industry scenarios, especially in the scenario development stage, where collecting a sufficient amount of sample data for model training faces great challenges. To address this problem, this application proposes a set of innovative methods to generate a large number of high-quality fine-tuning data samples using limited small sample data by fine-tuning question and answer pairs, building special templates, and formulating corresponding rules. This method aims to accurately fine-tune large models to improve their accuracy in generating SQL statements and effectively reduce the illusion that may occur when the model understands natural language.

[0050] The following is combined with Figure 1-2 The method for generating the instruction data set in the embodiment of the present application is described in detail. Figure 1 As shown, the method mainly includes the following steps:

[0051] S101 generates sample data of questions and SQL pairs of various question-and-answer types based on the collected small sample question-and-answer data.

[0052] In one embodiment, first, a series of representative natural language questions are collected. These questions should cover different question and answer types to ensure that the subsequently generated data set is diverse and extensive. For each collected question, a corresponding SQL statement needs to be generated. Through the above steps, a small sample question and answer data set is constructed, which will serve as the basis for the subsequent automatic generation of more question and answer pair sample data.

[0053] Furthermore, the small sample Q&A data are classified based on the Q&A topic type.

[0054] In one implementation, based on manually collected small samples and business insight analysis, the problems are finely classified. For example, the problem categories include:

[0055] Question and answer topic types:

[0056] Sales volume topic category: covers sales volume-related issues, including sales volume, year-on-year sales volume, month-on-month sales volume, sales volume share, and combinations of sales volume and other dimensions.

[0057] The main category of ownership: issues related to ownership, including ownership, year-on-year ownership, month-on-month ownership, ownership ratio, and the combination of ownership and other dimensions.

[0058] Furthermore, the SQL query statements are classified based on the SQL value type. In one embodiment, the SQL value type classification includes:

[0059] Single-value questions: These questions have only one return value, such as "What is the sales volume in Hebei Province in 2024?".

[0060] Time series trend questions: This type of question focuses on the changing trend of data, such as "What is the sales trend of Dongfeng Group every month in 2024?".

[0061] Dimension distribution problems: This type of problem focuses on the distribution of data in a certain dimension, such as "the sales ranking of various brands in Beijing in 2024."

[0062] Furthermore, the question types and SQL value types are arranged and combined to generate sample data of question and SQL pairs of various question and answer types.

[0063] In a further embodiment of the present application, by systematically permuting and combining question types and SQL value types, it is possible to generate sample data of questions and SQL pairs of various question-answer types. For example, by combining "sales volume" with "single-value question", a question about sales volume at a specific time point and a corresponding SQL query statement are generated. The following is an example of a sales volume single-value question:

[0064] Question: What will be the sales volume in Hebei Province in 2024?

[0065] SQL: SELECT COUNT(*)AS sales FROM sales_table WHERE year=2024ANDprovince='Hebei Province'.

[0066] By automatically generating various types of question-answer pair sample data from small sample question-answer data, the reliance on manual labeling is reduced, thereby significantly reducing the cost and time of data preparation.

[0067] S102 generates a template for the sample data based on the question and the SQL, where the template includes slots corresponding to multiple data types.

[0068] In one embodiment of the present application, the sample data of the question and SQL are decomposed to obtain multiple components, including time, condition, and question. A slot is set for each component to obtain a template containing multiple slots. The slot is a variable.

[0069] Specifically, we first need to abstract the composition of the question-SQL pair and decompose it into three core components: time, screening conditions, and questions. Next, we set corresponding "slots" for each component based on the sample data. Slots are variables. For example:

[0070] Question: What will be the sales volume in Hebei Province in 2024?

[0071] SQL: SELECT COUNT(*)AS sales FROM sales_table WHERE year=2024ANDprovince='Hebei Province';

[0072] Card slot correspondence:

[0073] Question structure: [time] + [screening conditions] + [question];

[0074] SQL structure: SELECT COUNT(*)AS sales FROM sales_table

time

filter condition

[0075] This application converts complex natural language questions into structured components so that the data set can be automatically generated and expanded later. By breaking down the question into three parts: time, filter conditions, and questions, various variants can be generated more flexibly to expand the data set. This structured approach enables the model to better understand and generalize different types of queries, improving the adaptability of the model in practical applications.

[0076] S103 generates card slot data based on preset rules, fills the card slots in the template according to the card slot data, and obtains a plurality of expanded initial instructions.

[0077] In one implementation, the card slot data is generated based on a preset rule, and the card slots in the template are filled according to the card slot data.

[0078] Specifically, firstly, based on the LLM prompt engineering technology, synonym expansion is performed on each question-and-answer type of question to generate question slot data.

[0079] LLM (Large Language Model) prompt engineering technology is used to generate synonyms or approximate expressions for each question-answering type of question. Synonym expansion is performed on each question to generate multiple different question statements that are semantically equivalent but have different wording. The expanded question statements are used as question slot data, which will be used to fill the corresponding slots in the template. For example, for the question "sales trend", it can be expanded into several synonymous questions such as "sales trend", "sales distribution", "sales change" and "how is the sales trend".

[0080] Furthermore, based on a preset time format, time slot data is randomly generated.

[0081] In the implementation mode of the present application, the design of the time slot is to increase the diversity of the time dimension in the question and answer data set, so that the generated questions are closer to the actual application scenario and can cover the needs of different time granularities. First define the time granularity: determine the different time granularities that the time slot can contain, such as year, month, day, time period (from year to year, from month to month, from day to day), quarter, current year, current month, etc. The values ​​of these time slots are randomly generated based on the program to ensure the diversity and extensiveness of the data set in time. Apply the generated time values ​​to the questions and SQL templates, replace the corresponding time slots, to create specific questions and query statements.

[0082] Furthermore, condition fields are randomly extracted from the database fields to generate condition slot data. The design of the filter condition slot is to enable the question-and-answer dataset to flexibly adapt to various database query scenarios, especially when data needs to be filtered according to specific conditions. First, define the filter condition type: determine the type that the filter condition slot can contain, including single filter conditions and multiple filter conditions. Extract the values ​​of the table fields from the database and perform deduplication processing to ensure the diversity and uniqueness of the filter conditions. Generate filter conditions in various formats including field value, field value plus field name, field name plus field value, etc. For example, the content of the filter condition includes: field value (such as "Dongfeng"), field value plus field name (such as "Dongfeng brand"), field name plus field value (such as "brand is Dongfeng"). For an example of two filter conditions, it can be "Brand is Dongfeng National II".

[0083] Finally, fill the corresponding slots in the template according to the question slot data, time slot data, and condition slot data.

[0084] In the implementation steps of the present application, the question slot data, time slot data and filter condition slot data are combined and filled into a predefined template to generate a specific question and answer pair. First, combine the slot data: randomly extract a question from the question slot data set to ensure that each generated question and answer pair has a core query intent. Randomly extract a time value from the time slot data set. If the time is not extracted, use the preset default time. Randomly extract one or more filter conditions from the filter condition slot data set to increase the specificity of the query statement. Among them, the question is a required option, and at least one of the time and filter condition must be selected. If no time is extracted for filling, the default time is used.

[0085] Fill the extracted questions, time and screening conditions into the corresponding slots in the template to form a complete question-answer pair. Generate a specific question-answer pair based on the filled template, including natural language questions and corresponding SQL query statements, to obtain the initial instruction data set.

[0086] By setting corresponding templates and rules based on this application, a large number of instruction data sets can be automatically generated, which improves the quality of the data set and the performance of the model.

[0087] In one embodiment, after obtaining the expanded multiple initial instructions, the method further includes: performing one or more standardization processes of replacing, merging, and marking on the words in the initial instructions to generate standardized initial instructions.

[0088] In the implementation of this application, after completing the initial generation of slot data, synonym mapping is an important step to ensure the language quality of the dataset and improve model performance.

[0089] Specifically, we will perform precise replacement operations: identify and replace those colloquial words that do not match professional definitions. For example, in colloquial language, "year-on-year" is often used to refer to "year-on-year growth rate", so we will update "year-on-year" in the question to "year-on-year growth rate"; similarly, "month-on-month" will be accurately replaced with "month-on-month growth rate" to improve the professionalism and accuracy of the question statement.

[0090] For those words that are easy to cause confusion, we clearly distinguish them. For example, "year-on-year" and "proportion" are easy to confuse, so "proportion" is changed to "occupancy rate" to enhance the clarity of meaning.

[0091] Specially mark the terms that are unique to the field and may be confused with the general language. For example, "Beijing heavy truck" may be misunderstood by the big model as a heavy truck in Beijing City, so we mark and replace "Beijing heavy truck" with "{Beijing heavy truck}" to avoid ambiguity.

[0092] Merge synonyms in spoken language, such as uniformly replacing "FAW" and "Jiefang" with "FAW Jiefang" to enhance data consistency.

[0093] Through synonym mapping, the clarity of problem statements is ensured, reducing ambiguity and vagueness in natural language. It improves the model's ability to understand the intent of questions, enabling the model to respond more accurately to user queries. Synonym mapping processing improves the quality of the dataset, providing more accurate and consistent data for model training. It reduces interference in model training, improving the effect and final performance of model training.

[0094] Through this method, this application not only improves the quality of the dataset and the performance of the model, but also provides users with a more accurate and natural interaction experience, further promoting the development of natural language processing technology.

[0095] S104 calculates the score of the initial instruction, determines the proportion distribution of each question type based on the score of the initial instruction, and extracts instructions of a preset proportion type from the initial instruction according to the proportion distribution to generate the final instruction dataset.

[0096] In one implementation, calculating the score of the initial instruction includes: inputting the question in the initial instruction into a preset model to obtain the SQL query statement output by the model; calculating the error rate of the initial instruction according to the SQL query statement corresponding to the question in the initial instruction and the SQL query statement output by the model; obtaining the score of the initial instruction according to the error rate of the initial instruction. Among them, the preset model can adopt a large language model in the prior art.

[0097] Specifically, multiple questions of the same type can be input into the large language model to obtain the output SQL statements. Compare the SQL query statement output by the model with the correct SQL query statement corresponding to the question in the initial instruction. If they are consistent, it means the model's answer is correct. According to the comparison result, calculate the error rate of the model output. For example, if 10 questions are input and 4 are answered correctly, the error rate is 60%. Then the score of the input initial instruction is 60%. The higher the score of a question that is more difficult to answer correctly, the more valuable the question is that is prone to errors.

[0098] Furthermore, according to the score of the initial instruction, perform normalization to obtain the normalized instruction score. Summarize the instruction scores of different question types to obtain the scores of each question type; obtain the proportion distribution of each question type according to the scores of each question type.

[0099] Normalize the scores of the initial instructions so that the sum of the scores of all instructions is 1 (or 100%) to facilitate comparison and calculation of the relative importance of each instruction. Sum up the scores of instructions for different question types and calculate the total score for each question type. Based on the summarized scores, calculate the proportion of each question type in all questions, that is, the percentage of each type of score in the total score.

[0100] Furthermore, based on the distribution of the proportions of each question type, questions of corresponding proportions are extracted to obtain the final instruction data set constructed.

[0101] In order to facilitate understanding of the method of the embodiment of the present application, Figure 2 Further description, such as Figure 2 As shown, including:

[0102] Step 1: Generate sample data of question-answer type questions and SQL pairs. Based on a small sample of question-answer data, generate sample data of multiple question-answer type questions and SQL pairs.

[0103] Step 2: Generate a template. Generate a template containing slots for multiple data types based on the question and SQL.

[0104] Step 3: Generate slot data and fill in the template. Generate slot data using preset rules and fill in the slots in the template to obtain the expanded initial instructions.

[0105] Step 4: Calculate the initial instruction score. Calculate the score of the initial instruction based on the matching degree between the SQL query statement output by the model and the expected SQL query statement.

[0106] Step 5: Determine the distribution of question types. Based on the scores of the initial instructions, determine the distribution of each question type in the final data set.

[0107] Step 6: Generate the final instruction dataset. According to the distribution of question types, extract the corresponding proportion of questions to construct the final instruction dataset.

[0108] The embodiment of the present application provides a method for generating an instruction data set, by creating a template based on sample data of questions and SQL pairs of various question and answer types, and generating slot data based on rules to fill in the template, so as to create a rich variety of question and answer pairs, and enhance the coverage and diversity of the data set. A large number of high-quality fine-tuning instruction data sets can be automatically generated based on small sample instruction data, improving the efficiency of constructing instruction data sets and the accuracy of the model when generating SQL statements. And the present application can generate a more balanced and representative instruction data set by calculating the score of the initial instruction and determining the distribution of the proportion of question types accordingly, which helps to train a more accurate and reliable language model. It provides strong technical support for building efficient and accurate natural language processing models.

[0109] According to another aspect of the embodiment of the present application, there is also provided a device for generating an instruction data set for implementing the above-mentioned method for generating an instruction data set. Figure 3 As shown, the device comprises:

[0110] A question-answer classification module 301 is used to generate sample data of questions and SQL pairs of various question-answer types based on the collected small sample question-answer data;

[0111] A template making module 302 is used to generate a template for sample data based on the question and SQL, and the template includes slots corresponding to multiple data types;

[0112] The instruction expansion module 303 is used to generate card slot data based on preset rules, fill the card slots in the template according to the card slot data, and obtain a plurality of expanded initial instructions;

[0113] The instruction data set generation module 304 is used to calculate the score of the initial instruction, determine the proportion distribution of each question type based on the score of the initial instruction, extract instructions of a preset proportion type from the initial instruction according to the proportion distribution, and generate a final instruction data set.

[0114] It should be noted that the device for generating an instruction data set provided in the above embodiment only uses the division of the above functional modules as an example when executing the method for generating an instruction data set. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for generating an instruction data set provided in the above embodiment and the method for generating an instruction data set are of the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0115] According to another aspect of an embodiment of the present application, an electronic device corresponding to the method for generating an instruction data set provided in the aforementioned embodiment is also provided to execute the aforementioned method for generating an instruction data set.

[0116] Please refer to Figure 4 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 4 As shown, the electronic device includes: a processor 400, a memory 401, a bus 402 and a communication interface 403, and the processor 400, the communication interface 403 and the memory 401 are connected via the bus 402; the memory 401 stores a computer program that can be run on the processor 400, and when the processor 400 runs the computer program, it executes the method for generating an instruction data set provided in any of the aforementioned embodiments of the present application.

[0117] The memory 401 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 403 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.

[0118] The bus 402 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 401 is used to store programs, and the processor 400 executes the program after receiving the execution instruction. The method for generating the instruction data set disclosed in any implementation of the embodiment of the present application may be applied to the processor 400, or implemented by the processor 400.

[0119] The processor 400 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 400. The above processor 400 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and completes the steps of the above method in combination with its hardware.

[0120] The electronic device provided in the embodiment of the present application and the method for generating the instruction data set provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.

[0121] According to another aspect of the embodiments of the present application, a computer-readable storage medium corresponding to the method for generating an instruction data set provided in the aforementioned embodiments is also provided, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method for generating an instruction data set provided in any of the aforementioned embodiments.

[0122] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0123] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the method for generating an instruction data set provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0124] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.

Claims

1. A method for generating an instruction data set, characterized in that: include: Generate sample data of questions and SQL pairs of various question-answer types based on the collected small sample question-answer data; Generate a template for the sample data based on the question and SQL, the template including slots corresponding to multiple data types; Generate card slot data based on preset rules, fill the card slots in the template according to the card slot data, and obtain a plurality of expanded initial instructions; The score of the initial instruction is calculated, and the proportion distribution of each question type is determined based on the score of the initial instruction. According to the proportion distribution, instructions of a preset proportion type are extracted from the initial instruction to generate a final instruction data set.

2. The method according to claim 1, characterized in that Based on the collected small sample question and answer data, sample data of questions and SQL pairs of various question and answer types are generated, including: Classifying the small sample question and answer data based on the question and answer topic type; Classify SQL query statements based on SQL value types; The question types and SQL value types are arranged and combined to generate sample data of questions and SQL pairs of the plurality of question-and-answer types.

3. The method according to claim 1, characterized in that The generating a template for the sample data based on the problem and SQL includes: Decomposing the problem and SQL pair sample data to obtain multiple components, wherein the components include time, conditions, and problems; A slot is set for each component to obtain a template containing multiple slots.

4. The method according to claim 3, characterized in that: The generating the card slot data based on the preset rules and filling the card slots in the template according to the card slot data includes: Based on LLM prompt engineering technology, synonym expansion is performed on each question-answer type to generate question slot data; Randomly generate time slot data based on the preset time format; Randomly extract conditional fields from database fields to generate conditional slot data; According to the problem slot data, time slot data and condition slot data, the corresponding slots in the template are filled.

5. The method according to claim 1, characterized in that: After the expanded initial instructions, it also includes: One or more standardization processes of replacement, merging, and marking are performed on the words in the initial instructions to generate standardized initial instructions.

6. The method according to claim 1, characterized in that Calculating the score of the initial instruction includes: Input the question in the initial instruction into a preset model to obtain an SQL query statement output by the model; Calculate the error rate of the initial instruction based on the SQL query statement corresponding to the question in the initial instruction and the SQL query statement output by the model; A score of the initial instruction is obtained according to the error rate of the initial instruction.

7. The method according to claim 1, characterized in that The proportion distribution of each question type is determined based on the score of the initial instruction, including: Normalizing the scores of the initial instructions to obtain normalized instruction scores; The instruction scores of different question types are summarized to obtain the scores of each question type; According to the scores of the various question types, the proportion distribution of the various question types is obtained.

8. A device for generating an instruction data set, characterized in that: include: The question-answer classification module is used to generate sample data of questions and SQL pairs of various question-answer types based on the collected small sample question-answer data; A template making module, used to generate a template for sample data based on the question and SQL, wherein the template includes slots corresponding to multiple data types; An instruction expansion module, used to generate card slot data based on preset rules, fill the card slots in the template according to the card slot data, and obtain a plurality of expanded initial instructions; An instruction data set generation module is used to calculate the score of the initial instruction, determine the proportion distribution of each question type based on the score of the initial instruction, and extract instructions of a preset proportion type from the initial instruction according to the proportion distribution to generate a final instruction data set.

9. An electronic device, characterized in that: The invention comprises a processor and a memory storing program instructions, wherein the processor is configured to execute the method for generating an instruction data set according to any one of claims 1 to 7 when executing the program instructions.

10. A computer-readable medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions are executed by a processor to implement a method for generating an instruction data set as described in any one of claims 1 to 7.