Instruction fine tuning data set screening method and device and medium

By constructing a tag graph for semantic spatial modeling, the problems of low efficiency and insufficient accuracy of instruction fine-tuning data set screening in the existing technology are solved, and efficient and reliable data set screening is achieved, and high-quality and diverse data subsets can be selected in large-scale data pools.

CN120492639AActive Publication Date: 2025-08-15SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510380313.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-15
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively screen out high-quality and diverse data sets during instruction fine-tuning, and the existing methods are inefficient in sampling on big data pools, and they cannot accurately capture the semantics of complex instructions.

Method used

By constructing a label graph for semantic spatial modeling, using the relationship between labels as edge weights, the information gain value of the data subset is calculated, and the filtering target is to maximize the information gain value, and the data subset is selected from the original instruction fine-tuning dataset.

Benefits of technology

It significantly improves the sampling efficiency and screening accuracy of instruction fine-tuning data sets, can balance the diversity and quality of data selection, and achieve efficient and reliable data set screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492639A_ABST
    Figure CN120492639A_ABST
Patent Text Reader

Abstract

The invention relates to an instruction fine tuning data set screening method and device and a medium. The method comprises the steps of obtaining an original instruction fine tuning data set; the method comprises the following steps of: constructing a tag graph by taking tags as graph nodes and taking a relationship between the tags as an edge weight, and performing semantic space modeling on an original instruction fine-tuning data set; wherein the contribution of each piece of data in the original instruction fine adjustment data set to the information amount of the data set comes from the label of the corresponding piece of data; and considering the transmission of the information amount on the label graph, calculating the information gain value of the data subset, and screening out a final instruction fine-tuning data subset from the original instruction fine-tuning data set by taking the maximum information gain value of the data subset as a screening target. Compared with the prior art, the method comprehensively considers the quality and diversity of the data set in the semantic space, and data set screening is more efficient and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to a method, device and medium for screening an instruction fine-tuning data set. Background Art

[0002] With the rapid development of natural language processing technology, large language models have broad application prospects in the field of natural language understanding. However, during instruction fine-tuning, how to select high-quality and diverse subsets from large-scale data pools as training data has become a pressing issue.

[0003] Traditional data screening methods for instruction fine-tuning typically define data quality evaluation criteria and use heuristic rules to maintain data diversity. However, this rule-based approach fails to comprehensively consider the quality and diversity of a dataset. Moreover, heuristic rules are mostly implemented in the embedding space, making it difficult to accurately capture the semantics of complex instructions. Other methods based on submodular functions, while defining a method for evaluating dataset diversity based on submodular functions, require calculating the similarity in the embedding space between each element in each iteration, resulting in low sampling efficiency on large data pools and making it difficult to meet the requirements of efficient data processing.

[0004] After searching, Chinese invention patent application CN118260429A discloses a processing method for optimizing a large language model fine-tuning dataset. The method scores and labels each sample record in a first sample library based on a first sample scoring model and a first sample labeling model, then clusters all sample records in the first sample library based on the sample labels to obtain multiple first-category label record clusters. A fine-tuning dataset is constructed based on all the obtained first-category label record clusters and the first sample library with a preset data distribution indicator set as a reference. However, in the process of constructing the sample scoring model and the sample labeling model, more semantic and grammatical features cannot be introduced, resulting in insufficient accuracy and stability of the model.

[0005] Therefore, how to effectively screen out high-quality and diverse instruction data from large data sets while considering the semantic information of the data is an urgent problem that needs to be solved. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a method, device and medium for screening instruction fine-tuning data sets.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] According to a first aspect of the present invention, a method for screening an instruction fine-tuning dataset is provided, comprising:

[0009] Get the original instruction fine-tuning dataset;

[0010] Using labels as graph nodes and relationships between labels as edge weights, a label graph is constructed to perform semantic space modeling on the original instruction fine-tuning dataset. The contribution of each piece of data in the original instruction fine-tuning dataset to the information content of the dataset comes from the label of the corresponding piece of data.

[0011] Considering the transfer of information on the label graph, the information gain value of the data subset is calculated. The data subset with the maximum information gain value is used as the screening target, and the final instruction fine-tuning data subset is screened out from the original instruction fine-tuning data set.

[0012] Preferably, the edge weights between the tags are characterized by semantic similarity. When the semantic similarity between the tags is lower than a preset threshold, the edge weights between the tags are reset to 0.

[0013] Preferably, in the original instruction fine-tuning data set, the i-th data is represented as:

[0014]

[0015] Where: Indicates M rounds of conversation data, is the query of the jth round of dialogue in the i-th data, r i j is the reply of the i-th data in the j-th round of dialogue, L i Indicates the label list corresponding to the i-th data, s i Indicates the score of the i-th data.

[0016] Preferably, the information gain value of the data subset is calculated as follows:

[0017]

[0018] E i =s i v i

[0019] Where: D is the data subset; Φ is the information gain function, which satisfies the monotonically increasing and decreasing increasing speed; A is the transfer matrix, and the element a in the transfer matrix A is pq represents the amount of information transferred between the pth label and the qth label, which is calculated based on the edge weight between the labels; E i is the information amount of the i-th data, s i is the score of the i-th data, v i is the label vector of the i-th data.

[0020] Preferably, the information gain function Φ is a function that increases monotonically but at a decreasing rate of increase.

[0021] Preferably, the information gain function Φ is expressed as follows:

[0022] Φ(x)=x a ,0<a<1

[0023] Or, the information gain function Φ is expressed as follows:

[0024] Φ(x)=-e -x

[0025] Where: x is the input data.

[0026] Preferably, the information volume v of the i-th data i , the calculation expression is:

[0027]

[0028] Where: is the amount of information corresponding to the kth word in the i-th data, σ is the activation function, L i is the label list corresponding to the i-th data, l k is the label corresponding to the kth word in the i-th data, and K is the number of words in the i-th data.

[0029] Preferably, the element a in the transfer matrix A pq It represents the amount of information transferred between the p-th label and the q-th label, which is calculated based on the edge weight between the labels. The calculation expression is:

[0030]

[0031] Where: α is the transfer parameter; ω p is the similarity of the p-th tag itself, which can be set to a constant value of 1.

[0032] According to a second aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, any one of the methods described above is implemented.

[0033] According to a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, any one of the methods described above is implemented.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) The present invention constructs a label graph to perform semantic space modeling on the original instruction fine-tuning dataset, quantifies the quality and diversity of the dataset in the semantic space, and maximizes the information gain of the current label graph during each screening. There is no need to calculate the similarity of the embedding space between each element in each iterative screening, which significantly improves the sampling efficiency on large-scale instruction fine-tuning datasets. The information gain of the dataset is obtained in the semantic space, which avoids the problem in the prior art that the heuristic rules in the embedding space cannot accurately capture the semantics of complex instructions, thereby improving the accuracy of the instruction fine-tuning dataset screening.

[0036] (2) Data labels are used to characterize the contribution of each data in the original instruction fine-tuning dataset to the information content of the dataset. The relationship between different labels is considered, and the information content is defined to be transmitted along the edge of the label graph. This effectively depicts the distribution of information content on the label graph and improves the quality and reliability of the instruction fine-tuning dataset screening.

[0037] (3) The information gain function in the present invention is a monotonically increasing function with a decreasing rate of increase. If the derivative of the information gain function decreases too quickly, the diversity of selection is emphasized, which can effectively balance the diversity and quality of data selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of the method of the present invention;

[0039] Figure 2 Schematic diagram of the structure for modeling semantic space. DETAILED DESCRIPTION

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0041] Example 1

[0042] like Figure 1 As shown, this embodiment provides a method for screening an instruction fine-tuning dataset based on maximizing information gain, the method comprising:

[0043] S1. Obtain the original instruction fine-tuning dataset;

[0044] Among them, the i-th data in the original instruction fine-tuning dataset is expressed as:

[0045]

[0046] Where: Indicates M rounds of conversation data, is the query of the jth round of dialogue in the i-th data, r i j is the reply of the i-th data in the j-th round of dialogue, L i Indicates the label list corresponding to the i-th data; s i It represents the score of the i-th data and can be directly set to . In order to make the score proportional to the data quality, the IFD score and DEITA score can also be used. DEITA is the preferred solution.

[0047] S2. Use labels as graph nodes and the relationships between labels as edge weights to construct a label graph, such as Figure 2 As shown, the semantic space modeling is performed on the original instruction fine-tuning dataset; wherein, the contribution of each data in the original instruction fine-tuning dataset to the amount of information in the dataset comes from the label of the corresponding data.

[0048] In this embodiment, the edge weights between tags are characterized by semantic similarity. When the semantic similarity between tags is lower than a preset threshold, the edge weights between tags are reset to 0. Specifically, the edge weight ω between tag p and tag q is pq for:

[0049] ω pq =σ[ω(l p ,l q )≥T]ω(l p ,l q )

[0050] Where: l p and l q Respectively represent the pth label and the qth label in the label list; ω(l p ,l q ) is the semantic similarity between the p-th label and the q-th label. A semantic similarity function such as Word2Vec and BERT can be used. T is the semantic similarity threshold, which can be set to 0.6.

[0051] For the label graph: the information content of the entire dataset is the sum of the information content of each label; the contribution of each data point to the information content comes from its label, and the value of this information content is positively correlated with the quality of the data point. For example, if the data point's label is [storytelling, writing] and its score is 0.9, then its contribution to the "storytelling" label is 0.9, and its contribution to the "writing" label is also 0.9, for a total contribution of 1.8; to better characterize the distribution of information content on the label graph, the relationship between different labels is considered, and information content is defined as information transmission along the edges of the label graph.

[0052] For the original instruction fine-tuning dataset D P, set the upper limit N, information gain function E, the goal is to select a subset D S , so that E(D) takes the maximum value:

[0053]

[0054] Information gain function E(D) of data subset:

[0055]

[0056] E i =s i v i

[0057] Where: D is the data subset; Φ is the information gain function, which satisfies the monotonically increasing and decreasing increasing speed; A is the transfer matrix, and the element a in the transfer matrix A is pq represents the amount of information transferred between the pth label and the qth label, which is calculated based on the edge weight between the labels; E i is the information amount of the i-th data, s i is the score of the i-th data, v i is the label vector of the i-th data.

[0058] Specifically, the element a in the transfer matrix A is pq It represents the amount of information transferred between the p-th label and the q-th label, which is calculated based on the edge weight between the labels. The calculation expression is:

[0059]

[0060] Where: α is the transfer parameter, α=0 means no transfer, the larger the value, the more transfers are made. In this embodiment, the best effect is achieved when the transfer parameter is set to 1; ω p is the similarity of the p-th tag itself, which is set to a constant value of 1 in this embodiment.

[0061] Specifically, the label vector v of the i-th data i :

[0062]

[0063] Where: is the amount of information corresponding to the kth word in the i-th data, σ is the activation function, L i is the label list corresponding to the i-th data, l k is the label corresponding to the kth word in the i-th data item, and K is the number of words in the i-th data item. For example, v1 = (1, 0, 1, 0, 0, 0) means that the first data item has two labels, 1 and 3.

[0064] S3. Considering the information transfer on the label graph, calculate the information gain value of the data subset. Select the data subset with the largest information gain value as the screening target, and filter out the final instruction fine-tuning data subset from the original instruction fine-tuning data set. The filtered instruction fine-tuning data subset can be used for large-model knowledge question answering, large-model mathematical ability, and large-model logical reasoning training.

[0065] The present invention models the semantic space by constructing a label graph, while taking into account the quality and diversity of the dataset, thereby overcoming the defects of the existing technology that only defines data quality evaluation criteria or adopts heuristic rules to maintain diversity, and can evaluate the dataset more accurately as a whole.

[0066] This embodiment conducted a verification experiment on Tulu3. The amount of data in the original data pool was 939K, 50K training data were sampled, and the base model for training was Llama3.1-8B.

[0067] In Table 1 below, HE is the abbreviation of HumanEval, AE stands for AlpacaEvalv2, MT stands for MTBench, Wild stands for WildBench, and Avg is the average of the scores of the corresponding evaluation benchmarks normalized to percentages.

[0068] Table 1

[0069]

[0070]

[0071] The present invention achieved the best results on multiple test benchmarks, using only 5% of the data. On average, it achieved the same effect as training with a full data pool across multiple benchmarks. It was able to select high-quality and diverse training data for instruction fine-tuning tasks, achieving the best results compared to other data screening methods. The proposed data screening method has good effectiveness and robustness.

[0072] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0073] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0074] The processing unit performs the various methods and processes described above, such as methods S1 to S3. For example, in some embodiments, methods S1 to S3 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S3 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S3 by any other appropriate means (e.g., by means of firmware).

[0075] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), and the like.

[0076] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0077] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0078] Example 2

[0079] The difference between this embodiment and embodiment 1 is that the information gain function Φ is a monotonically increasing function with a decreasing rate of increase, so as to balance the diversity and quality of data selection. If the derivative of the information gain function decreases too quickly, the diversity of selection is emphasized.

[0080] Specifically, the information gain function Φ is expressed as follows:

[0081] Φ(x)=x a ,0<a<1

[0082] Or, the information gain function Φ, the mathematical expression is:

[0083] Φ(x)=-e -x

[0084] Where: x is the input data.

[0085] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for screening instruction fine-tuning datasets, characterized in that: include: Get the original instruction fine-tuning dataset; Using labels as graph nodes and relationships between labels as edge weights, a label graph is constructed to perform semantic space modeling on the original instruction fine-tuning dataset. The contribution of each piece of data in the original instruction fine-tuning dataset to the information content of the dataset comes from the label of the corresponding piece of data. Considering the transfer of information on the label graph, the information gain value of the data subset is calculated. The data subset with the maximum information gain value is used as the screening target, and the final instruction fine-tuning data subset is screened out from the original instruction fine-tuning data set.

2. The method for screening instruction fine-tuning datasets according to claim 1, characterized in that: The edge weights between the tags are characterized by semantic similarity. When the semantic similarity between the tags is lower than a preset threshold, the edge weights between the tags are reset to 0.

3. The method for screening instruction fine-tuning datasets according to claim 1, characterized in that: In the original instruction fine-tuning dataset, the i-th data is represented as: Where: Indicates M rounds of conversation data, is the query of the jth round of dialogue in the i-th data, is the reply of the i-th data in the j-th round of dialogue, L i Indicates the label list corresponding to the i-th data, s i Indicates the score of the i-th data.

4. The method for screening instruction fine-tuning datasets according to claim 1, characterized in that: The information gain value of the data subset is calculated as follows: E i =s i v i Where: D is the data subset; Φ is the information gain function, which satisfies the monotonically increasing and decreasing increasing speed; A is the transfer matrix, and the element a in the transfer matrix A pq represents the amount of information transferred between the pth label and the qth label, which is calculated based on the edge weight between the labels; E i is the information amount of the i-th data, s i is the score of the i-th data, v i is the label vector of the i-th data.

5. The method for screening instruction fine-tuning data sets according to claim 4, characterized in that: The information gain function Φ is a function that increases monotonically but at a decreasing rate of increase.

6. The method for screening instruction fine-tuning datasets according to claim 5, characterized in that: The information gain function Φ is expressed as follows: Φ(x)=x a ,0<a<1 Or, the information gain function Φ is expressed as follows: Φ(x)=-e -x Where: x is the input data.

7. The method for screening instruction fine-tuning data sets according to claim 4, characterized in that: The label vector v of the i-th data i , the calculation expression is: Where: is the amount of information corresponding to the kth word in the i-th data, σ is the activation function, L i is the label list corresponding to the i-th data, l k is the label corresponding to the kth word in the i-th data, and K is the number of words in the i-th data.

8. The method for screening instruction fine-tuning datasets according to claim 4, characterized in that: The element a in the transfer matrix A pq It represents the amount of information transferred between the p-th label and the q-th label, which is calculated based on the edge weight between the labels. The calculation expression is: Where: α is the transfer parameter; ω p is the similarity of the p-th label itself, which is a constant value.

9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Processing method and device for optimizing fine tuning data set of large language model

    CN118260429A

  • Analysis method and device of graph network structure and computer equipment

    CN111814006A

  • Civil aviation field image-text retrieval method based on multi-modal pre-training model

    CN116662630A

  • Multi-modal large model training method and device, storage medium and electronic equipment

    CN119150997A

  • METHOD FOR SELECTING A SET OF MICRO-INSERTS FOR CONTROLLING A DATA PROCESSING SYSTEM WITH SELECTED SYSTEM FEATURES

    DE2308859A1