Instruction tuning data generation method and device, and storage medium

By introducing a knowledge concept feedback mechanism and a large language model with pre-fine-tuning training, the instruction fine-tuning data is generated based on subtext blocks, which solves the problem of uneven data distribution and achieves more efficient and higher-quality instruction fine-tuning data generation.

WO2025139434A1PCT designated stage expired Publication Date: 2025-07-03BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/131763
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-11-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The prior art has the problem of uneven data distribution when generating instruction fine-tuning data, resulting in poor results on long-tail samples.

Method used

By introducing a knowledge concept feedback mechanism, a large language model with pre-fine-tuned training is used to generate knowledge concepts based on subtext blocks, and instruction fine-tuning data is generated through the data distribution feedback mechanism to ensure balanced data distribution.

Benefits of technology

It improves the efficiency and quality of automated construction of instruction fine-tuning data, solves the problem of uneven data distribution, and improves the balance of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131763_03072025_PF_FP_ABST
    Figure CN2024131763_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of natural language processing, and relates to an instruction tuning data generation method and device, and a storage medium. The method comprises: 1) acquiring a first knowledge base; 2) segmenting the first knowledge base into a plurality of text sub-blocks according to a fixed length, and sequentially inputting the text sub-blocks into a large language model, so as to generate a plurality of knowledge concepts; 3) inputting the knowledge concepts and preset related background knowledge into the large language model, so as to generate first instruction tuning data, and processing the first instruction tuning data to obtain second instruction tuning data; 4) determining whether the amount of second instruction tuning data is greater than an average value of the total amount of the second instruction tuning data, and if so, returning to step 3), otherwise, proceeding to step 5); and 5) using the second instruction tuning data as instruction tuning data of the knowledge concepts. The method is an instruction tuning data generation method in which a knowledge concept feedback mechanism is combined with a large language model, such that the construction efficiency of instruction tuning data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method, device and storage medium for generating instruction fine-tuning data Technical Field

[0001] The present invention belongs to the field of natural language processing technology, and relates to a method, device and storage medium for generating instruction fine-tuning data, and in particular to a method, device and storage medium for generating instruction fine-tuning data that combines a knowledge concept feedback mechanism and a large language model. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, the new generation of artificial intelligence technologies represented by large language models have brought revolutionary breakthroughs. Large language models have demonstrated excellent performance in various applications.

[0003] However, training large language models requires a large amount of high-quality training data, especially instruction fine-tuning data. This instruction fine-tuning data often requires manual annotation, which is very costly. Automating the generation of high-quality instruction fine-tuning data is a significant research topic. To this end, many researchers have begun researching methods for generating instruction fine-tuning data.

[0004] For example, the Chinese invention patent application number 202310827694.5 proposes a method for generating instruction fine-tuning data, the specific steps of which are as follows: S1, obtaining a first knowledge base and a first preset number of seed tasks; S2, generating prompt information, the prompt information including: the first preset number of seed tasks, the first knowledge base, and preset instruction generation requirements; S3, obtaining first instruction fine-tuning data based on the prompt information and a preset large language model; S4, processing the first instruction fine-tuning data to obtain second instruction fine-tuning data. This generation method improves the quality of generated instruction fine-tuning data and reduces the probability of generating noise data by introducing knowledge base data. However, the above-mentioned instruction fine-tuning data generation method inherits the preference characteristics of the large language model, that is, it tends to favor words with high frequency of occurrence. Therefore, the generated data is unevenly distributed, which works well for common fine-tuning instructions, but is limited in effect for long-tail samples.

[0005] Therefore, in view of the defects of the above-mentioned existing technologies, it is urgent to study a new method for generating instruction fine-tuning data.

[0006] Summary of the Invention

[0007] In response to the shortcomings of existing technical solutions, the present invention proposes a method for generating instruction fine-tuning data that combines a knowledge concept feedback mechanism and a large language model, which is applied to large language model instruction fine-tuning and improves the efficiency of constructing instruction fine-tuning data.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A method for generating instruction fine-tuning data, characterized by comprising the following steps:

[0010] 1) Obtaining a first knowledge base;

[0011] 2) dividing the first knowledge base into a plurality of sub-text blocks according to a fixed length, and inputting the plurality of sub-text blocks into the large language model in sequence to generate a plurality of knowledge concepts respectively;

[0012] 3) Inputting one of the knowledge concepts and the preset related background knowledge into the large language model to generate first instruction fine-tuning data corresponding to the knowledge concept;

[0013] 4) Determine whether the number of second instruction fine-tuning data corresponding to the knowledge concept is greater than the average number of second instruction fine-tuning data corresponding to all knowledge concepts; if so, return to step 3); if not, proceed to step 5);

[0014] 5) Processing the first instruction fine-tuning data corresponding to the knowledge concept to obtain the second instruction fine-tuning data corresponding to the knowledge concept and using the second instruction fine-tuning data as the instruction fine-tuning data of the knowledge concept.

[0015] Preferably, the large language model in step 2) is a large language model that has been pre-fine-tuned.

[0016] Preferably, the fine-tuning training includes the following steps:

[0017] Create a knowledge concept extraction training dataset;

[0018] The large language model is fine-tuned using the knowledge concept extraction training dataset.

[0019] Preferably, the fine-tuning training method is a full-parameter fine-tuning method, a LoRA fine-tuning method, or a Prefix Tuning fine-tuning method.

[0020] Preferably, the preset relevant background knowledge in step 3) is obtained from a pre-prepared database, which stores the original data involved in generating instruction fine-tuning data, and a retrieval method is used to recall text fragments related to the knowledge concept from the database as the preset relevant background knowledge.

[0021] Preferably, in step 4), determining whether the number of second instruction fine-tuning data corresponding to the knowledge concept is greater than the average number of second instruction fine-tuning data corresponding to all knowledge concepts is specifically as follows: obtaining the number N-KC of second instruction fine-tuning data corresponding to the knowledge concept, the number M of types of knowledge concepts, and the number N of second instruction fine-tuning data corresponding to all knowledge concepts, and determining whether N-KC is greater than N / M.

[0022] Preferably, in step 5), after obtaining the second instruction fine-tuning data, the number N-KC of second instruction fine-tuning data corresponding to the knowledge concept, the number M of types of knowledge concepts, and the number N of second instruction fine-tuning data corresponding to all knowledge concepts are updated.

[0023] Preferably, the large language model is ChatGLM, ChatGPT or GPT-4.

[0024] In addition, the present invention also provides a device for generating instruction fine-tuning data, characterized by comprising:

[0025] one or more processors;

[0026] a memory for storing one or more programs;

[0027] When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the above-mentioned instruction fine-tuning data generation method.

[0028] Finally, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps in the above-mentioned instruction fine-tuning data generation method are implemented.

[0029] Compared with the prior art, the method, device, and storage medium for generating instruction fine-tuning data of the present invention have one or more of the following beneficial technical effects:

[0030] 1. By introducing a knowledge concept feedback mechanism, the present invention solves the problem of uneven data distribution in the method of automatically generating instruction fine-tuning data based on a large language model, thereby improving the overall quality of automatically constructed instruction fine-tuning data.

[0031] 2. The present invention first fine-tunes the large language model for the knowledge concept extraction task, so that the large language model can extract knowledge concepts based on the given text. Then, based on the extracted knowledge concepts and related text, the large language model automatically generates instruction fine-tuning data corresponding to the knowledge concepts. During the generation process, data distribution feedback is performed based on the amount of instruction fine-tuning data corresponding to each knowledge concept, thereby guiding the data distribution generated by the large language model to be more balanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG1 is a flow chart of a method for generating instruction fine-tuning data according to the present invention. DETAILED DESCRIPTION

[0033] Before describing in detail any embodiment of the present invention, it should be understood that the present invention is not limited in its application to the details of construction and arrangement of components set forth in the following description or illustrated in the following drawings. The present invention is capable of other embodiments and can be practiced or carried out in various ways. In addition, it should be understood that the words and terms used herein are for descriptive purposes and should not be considered restrictive. As used herein, "including" or "having" and variations thereof are intended to cover the items listed below and their equivalents as well as additional items.

[0034] Furthermore, in the disclosure of the present invention, the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the element may be multiple, and the term "one" should not be understood as a limitation on the quantity.

[0035] In response to the shortcomings of existing technical solutions, the present invention proposes a method, device and storage medium for generating instruction fine-tuning data. By introducing a knowledge concept feedback mechanism, it solves the problem of uneven data distribution in the method of automatically generating instruction fine-tuning data based on a large language model, thereby improving the overall quality of automatically constructed instruction fine-tuning data.

[0036] FIG1 is a flow chart showing a method for generating instruction fine-tuning data according to the present invention. As shown in FIG1 , the method for generating instruction fine-tuning data according to the present invention comprises the following steps:

[0037] 1. Obtain the first knowledge base.

[0038] In the present invention, a first knowledge base is obtained and recorded as K.

[0039] Furthermore, similar to the prior art, in the present invention, obtaining the first knowledge base may specifically include:

[0040] 1. Acquire first knowledge data by crawling or downloading, wherein the first knowledge data includes structured knowledge data and / or unstructured knowledge data. The first knowledge data may be, for example, Wikipedia documents, common crawl data, existing knowledge graphs, etc.

[0041] 2. Divide the first knowledge data into a pre-set format to form a first knowledge base, wherein the pre-set format includes a document, a table, and other formats.

[0042] 2. Divide the first knowledge base into multiple sub-text blocks according to fixed lengths, and input the multiple sub-text blocks into the large language model in sequence to generate multiple knowledge concepts respectively.

[0043] Because the first knowledge base is too long, directly inputting it into the large language model will not easily generate knowledge concepts. Therefore, in this invention, the first knowledge base K is divided into multiple sub-text blocks of a fixed length N, denoted as C. These sub-text blocks are sequentially input into the large language model to generate multiple knowledge concepts, denoted as KC.

[0044] It should be noted that, since existing large language models are not specifically trained for generating knowledge concepts based on sub-text blocks, directly using these large language models to generate knowledge concepts will be less effective. Therefore, in this invention, when generating knowledge concepts based on sub-text blocks, the large language model used is a pre-trained large language model. Furthermore, this pre-trained large language model can be deployed on a cloud server for ease of use.

[0045] In the present invention, the specific large language model fine-tuning training includes the following steps:

[0046] 1. Create a knowledge concept extraction training dataset.

[0047] In the present invention, the knowledge concept extraction training data set can be collected in the form of manual annotation, and a sub-text block is input, and the knowledge concept corresponding to the sub-text block is output.

[0048] The format of the knowledge concept extraction training data set of the present invention is as follows:

[0049] {"input":"You are a knowledge concept extractor that directly outputs knowledge concepts based on the input text without explanation. The input text is as follows: {Blockchain is a distributed database technology that can securely store and transmit data between multiple computers. The working principle of blockchain is that it divides data into different "blocks" and connects these blocks in the form of a chain. Each block contains some data and the hash value of the previous block (a value that can confirm the uniqueness of the data), thus forming an ever-extending chain.\n\nOne of the characteristics of blockchain is that it is decentralized. This means that it does not rely on a central server, but is distributed among multiple different computers. In this way, it can resist the risk of single point failure and also makes it more difficult to tamper with the data.\n\nBlockchain also has very high security. Integrity. Since each block contains the hash value of the previous block, if someone wants to modify the data in one block, they must also modify the hash values ​​of all subsequent blocks, which is almost impossible. At the same time, blockchain uses a technology called consensus algorithm, which can verify the data transmitted by each node and ensure that all nodes have the same data. \n\nBlockchain technology is currently widely used in industries such as finance, supply chain management, healthcare, and government. For example, Bitcoin is a cryptocurrency based on blockchain technology.}\nOutput knowledge concepts:","target":"Blockchain\nDistributed database\nHash value\nChain\nDecentralization\nCentral server\nSingle point of failure\nConsensus algorithm\nBitcoin\nCryptocurrency\nFinance\nSupply chain management\nHealthcare"}

[0050] Among them, "input" represents the complete input, including prompt and sub-text block C, and "target" represents the output knowledge concept KC.

[0051] In the above example, the output knowledge concepts KC include blockchain, distributed database, hash value, chain, decentralization, central server, single point of failure, consensus algorithm, Bitcoin, cryptocurrency, finance, supply chain management, and healthcare, etc. Each knowledge concept represents a knowledge concept type.

[0052] 2. Use the knowledge concept extraction training dataset to fine-tune the large language model.

[0053] Large language models can use existing large language models in the industry, such as ChatGLM, ChatGPT, or GPT-4. ChatGLM is a bilingual Chinese-English conversational bot developed by Zhipu AI, a company focused on transforming Tsinghua University's technological achievements. ChatGPT is an AI chatbot developed by OpenAI. GPT-4 is a next-generation AI chatbot developed by OpenAI and is more advanced than ChatGPT.

[0054] During the specific fine-tuning training, the large language model is used to complete fine-tuning training on the knowledge concept extraction training dataset created above to enhance the knowledge concept generation capability of the large language model.

[0055] The fine-tuning method can use mainstream large language model fine-tuning methods, such as the full parameter fine-tuning method, the LoRA fine-tuning method, or the Prefix Tuning fine-tuning method.

[0056] 3. Input one of the knowledge concepts and the preset related background knowledge into the large language model respectively to generate first instruction fine-tuning data corresponding to the knowledge concept.

[0057] After obtaining a plurality of knowledge concepts KC, for each knowledge concept, the knowledge concept and preset related background knowledge are respectively input into the large language model to generate first instruction fine-tuning data corresponding to the knowledge concept.

[0058] The preset relevant background knowledge can be obtained from a pre-prepared database. This database stores original materials such as books, papers, and reports involved in generating instruction fine-tuning data. Using mainstream search methods, text snippets highly relevant to knowledge concepts are retrieved from this database as relevant background knowledge. Using this background knowledge, the large language model generates instruction fine-tuning data, improving data quality and reducing hallucinations.

[0059] Taking into account the existing large language models in the industry, such as ChatGLM, ChatGPT or GPT-4, which all have the ability to generate instruction fine-tuning data based on knowledge concepts and preset related background knowledge by default, therefore, in the present invention, when generating the first instruction fine-tuning data, the existing large language model in the industry can be used directly without the need to train and fine-tune it.

[0060] It should be noted that, since there are multiple sub-text blocks, different sub-text blocks may obtain the same knowledge concept. Therefore, the same knowledge concept may correspond to multiple first instruction fine-tuning data.

[0061] 4. Determine whether the number of second instruction fine-tuning data corresponding to the knowledge concept is greater than the average number of second instruction fine-tuning data corresponding to all knowledge concepts. If it is greater than the average, return to step 3; if it is not greater than the average, proceed to step 5.

[0062] Among them, in judging whether the number of second instruction fine-tuning data corresponding to the knowledge concept is greater than the average number of second instruction fine-tuning data corresponding to all knowledge concepts, the specific steps are: obtaining the number N-KC of second instruction fine-tuning data corresponding to the knowledge concept, the number M of types of knowledge concepts, and the number N of second instruction fine-tuning data corresponding to all knowledge concepts, and judging whether N-KC is greater than N / M.

[0063] Specifically, at the beginning, the number N-KC of second instruction fine-tuning data corresponding to the knowledge concept, the number M of types of knowledge concepts, and the number N of second instruction fine-tuning data corresponding to all knowledge concepts are initialized to 0. As the instruction fine-tuning data is continuously generated, N-KC, M and N are continuously updated.

[0064] Therefore, in the present invention, if N - KC is greater than N / M, it indicates that the instruction fine-tuning data corresponding to the current knowledge concept KC is greater than the average, and further data supplementation is not required. The current knowledge concept KC is skipped and the process returns to step 3 to process another knowledge concept KC. If N - KC is less than N / M, it indicates that the instruction fine-tuning data corresponding to the current knowledge concept KC is less than the average, and further data supplementation is required, and the process proceeds to step 5.

[0065] 5. Process the first instruction fine-tuning data corresponding to the knowledge concept to obtain the second instruction fine-tuning data corresponding to the knowledge concept and use the second instruction fine-tuning data as the instruction fine-tuning data of the knowledge concept.

[0066] In the present invention, the processing of the first instruction fine-tuning data corresponding to the knowledge concept mainly includes conventional deduplication, harmful information filtering, etc.

[0067] After obtaining the second instruction fine-tuning data, it is necessary to update the number N-KC of second instruction fine-tuning data corresponding to the knowledge concept, the number M of types of knowledge concepts, and the number N of second instruction fine-tuning data corresponding to all knowledge concepts.

[0068] Specifically, as described above, at the very beginning, the number N-KC of the second instruction fine-tuning data corresponding to the knowledge concept, the number M of the types of knowledge concepts, and the number N of the second instruction fine-tuning data corresponding to all knowledge concepts are initialized to 0. As the instruction fine-tuning data is continuously generated, N-KC, M, and N are continuously updated. For example, after the first first instruction fine-tuning data corresponding to the first knowledge concept is accepted, the data N-KC of the second instruction fine-tuning data corresponding to the first knowledge concept is increased by 1, and the number N of the second instruction fine-tuning data corresponding to all knowledge concepts is also increased by 1; and if the first knowledge concept has not appeared before and belongs to a new type of knowledge concept, then the number M of the types of knowledge concepts is also increased by 1. If the first knowledge concept has appeared before and does not belong to a new type of knowledge concept, then the number M of the types of knowledge concepts does not need to be updated.

[0069] It can be seen that in the present invention, by introducing a knowledge concept feedback mechanism, the problem of uneven data distribution in the method of automatically generating instruction fine-tuning data based on a large language model is solved, thereby improving the overall quality of the automatically constructed instruction fine-tuning data. Specifically, the present invention first fine-tunes the large language model for the knowledge concept extraction task, so that the large language model can extract knowledge concepts based on a given text, and then, based on the extracted knowledge concepts and related texts, enables the large language model to automatically generate instruction fine-tuning data corresponding to the knowledge concepts. During the generation process, data distribution feedback is performed based on the amount of instruction fine-tuning data corresponding to each knowledge concept, thereby guiding the data distribution generated by the large language model to be more balanced.

[0070] In addition, the present invention also provides a device for generating instruction fine-tuning data, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the instruction fine-tuning data generation method as described above.

[0071] Finally, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps in the above-mentioned instruction fine-tuning data generation method are implemented.

[0072] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art may, based on the principles of the present invention, modify or replace the technical solutions of the present invention with equivalents without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for generating instruction fine-tuning data, characterized in that, Including the following steps: 1) Obtain the first knowledge base; 2) Split the first knowledge base into multiple sub-text blocks according to a fixed length, and sequentially input the multiple sub-text blocks into a large language model to respectively generate multiple knowledge concepts; 3) Input one of the knowledge concepts and preset relevant background knowledge into the large language model respectively to generate first instruction fine-tuning data corresponding to the knowledge concept; 4) Determine whether the quantity of the second instruction fine-tuning data corresponding to the knowledge concept is greater than the average value of the quantities of the second instruction fine-tuning data corresponding to all knowledge concepts. If it is greater than the average value, return to step 3). If it is not greater than the average value, proceed to step 5); 5) Process the first instruction fine-tuning data corresponding to the knowledge concept, and after processing, obtain the second instruction fine-tuning data corresponding to the knowledge concept and use the second instruction fine-tuning data as the instruction fine-tuning data for the knowledge concept.

2. The method for generating instruction fine-tuning data according to claim 1, wherein The large language model in step 2) is a large language model that has been pre-fine-tuned and trained.

3. The instruction fine-tuning data generation method according to claim 2, wherein The fine-tuning training includes the following steps: Create a knowledge concept extraction training data set; Use the knowledge concept extraction training data set to perform fine-tuning training on the large language model.

4. The method for generating instruction fine-tuning data according to claim 3, wherein The fine-tuning training method is a full-parameter fine-tuning method, a LoRA fine-tuning method, or a Prefix Tuning fine-tuning method.

5. The instruction fine-tuning data generation method according to claim 4, wherein The preset relevant background knowledge in step 3) is obtained from a pre-prepared database. The database stores the original materials involved in the generation of instruction fine-tuning data, and uses a retrieval method to recall the text fragments related to the knowledge concept from the database as the preset relevant background knowledge.

6. The instruction fine-tuning data generation method according to claim 5, wherein In step 4), determining whether the quantity of the second instruction fine-tuning data corresponding to the knowledge concept is greater than the average value of the quantities of the second instruction fine-tuning data corresponding to all knowledge concepts is specifically: obtain the quantity N-KC of the second instruction fine-tuning data corresponding to the knowledge concept, the quantity M of the types of knowledge concepts, and the quantity N of the second instruction fine-tuning data corresponding to all knowledge concepts, and determine whether N-KC is greater than N / M.

7. The instruction fine-tuning data generation method according to claim 6, wherein In step 5), after obtaining the second instruction fine-tuning data, update the quantity N-KC of the second instruction fine-tuning data corresponding to the knowledge concept, the quantity M of the types of knowledge concepts, and the quantity N of the second instruction fine-tuning data corresponding to all knowledge concepts.

8. The method for generating instruction fine-tuning data according to claim 7, wherein The large language model is ChatGLM, ChatGPT, or GPT-4.

9. A generation device for instruction fine-tuning data, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the instruction fine-tuning data generation method according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the instruction fine-tuning data generation method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Generation method and device of instruction fine tuning data, equipment and medium

    CN116861928A

  • Method and device for training generative large language model based on knowledge base feedback

    CN117009490A

  • Method for enhancing memory ability of large model to external knowledge base

    CN117114012A

  • Multi-level visual evaluation report generation method and device based on large language model

    CN117194637A

  • Generation method and equipment of instruction fine tuning data and storage medium

    CN117763113A