Training data set construction method and device, storage medium and electronic equipment
By calling the network model to generate and filter the original data set, the target fine-tuning data set is constructed, which solves the problem of high manual labeling cost, and achieves rapid and efficient acquisition of high-quality training data sets, reducing the cost of model training.
Patent Information
- Application Number
- CN202510211813.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, it is costly to build training data sets by manual annotation, and it is difficult to quickly obtain high-quality training data sets, which affects the training efficiency and cost of large models.
By calling the first type of network model, the original data set in the target vertical field is generated and processed, the Q&A pairs that meet the preset rules are selected, and the target fine-tuning data set is constructed for fine-tuning optimization training.
Fast and efficiently construct target fine-tuning data sets, reducing the difficulty of obtaining training data and model training costs, and improving the quality and speed of training data sets.
Smart Images

Figure CN120067267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network models, and more particularly, to a method, apparatus, storage medium, and electronic device for constructing a training data set. Background Art
[0002] Large Language Models (LLMs) are a breakthrough technology in the field of artificial intelligence, especially in natural language processing (NLP). LLMs are typically based on deep learning architectures such as Transformer and are trained with a large amount of text data to capture the complex structures and semantic relationships of language. Through self-supervised learning, these models can understand and generate natural language text and perform well in various tasks such as language generation, translation, text summarization, and dialogue systems.
[0003] During the training process of large models, a large amount of training data is required. Currently, the construction of training data mainly relies on manual annotation, resulting in high labor costs. Therefore, getting rid of human limitations and constructing high-quality training data sets has become the focus of those skilled in the art. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, apparatus, storage medium, and electronic device for constructing a training data set to improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for constructing a training data set, the method including: calling a first type of network model to perform question-and-answer pair generation processing on an original data set in a target vertical domain to obtain an initial data set in the form of question-and-answer pairs; performing screening processing on the question-and-answer pairs in the initial data set to screen out the question-and-answer pairs that meet preset rules to construct a target fine-tuning data set.
[0007] By calling a first type of network model to perform question-and-answer pair generation processing on an original data set in a target vertical domain to obtain an initial data set in the form of question-and-answer pairs, and by performing screening processing on the question-and-answer pairs in the initial data set, excluding the question-and-answer pairs with too low quality assessment scores and those that are too simple and not suitable for fine-tuning optimization training, and only retaining the high-quality question-and-answer pairs that meet the preset rules, a target fine-tuning data set is constructed quickly and efficiently for fine-tuning optimization training, reducing the difficulty of obtaining training data and the model training cost.
[0008] Optionally, screening the question-and-answer pairs in the initial dataset to filter out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning dataset, including: calling a second type of network model to screen the question-and-answer pairs in the initial dataset to filter out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning dataset. Thereby ensuring the screening efficiency and effect of the question-and-answer pairs, and improving the quality and speed of training dataset construction.
[0009] Optionally, the calling of the second type of network model to screen the question-and-answer pairs in the initial dataset to filter out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning dataset includes: obtaining the preference score and / or perplexity of each question-and-answer pair in the initial dataset, where the preference score includes a question difficulty evaluation score and an answer quality evaluation score; screening out the question-and-answer pairs that meet the preset rules according to the preference score and / or perplexity of the question-and-answer pairs to construct a target fine-tuning dataset.
[0010] The purpose of screening the question-and-answer pairs in the initial dataset is to obtain answers that the large model considers to be of high quality and complex, and the preference alignment with the large model has a relatively low acquisition difficulty, only relying on partially manually labeled preference alignment data, and there is only a single loss function in the training process, which is relatively easy to converge. Therefore, it can be widely used in various vertical fields with answer preferences.
[0011] Optionally, the obtaining of the preference score of each question-and-answer pair in the initial dataset includes: evaluating the difficulty of the question in the question-and-answer pair to obtain the question difficulty evaluation score; obtaining the quality evaluation index of the answer in the question-and-answer pair, where the quality evaluation index includes any one or more of the detail index, the correctness index, the description tone style index, and the standardization index; determining the answer quality evaluation score according to the quality evaluation index.
[0012] Optionally, obtaining the perplexity of each question-and-answer pair in the initial dataset includes: determining the perplexity of the question-and-answer pair according to the probability of each token in the answer of the question-and-answer pair; where the token probability represents the probability of predicting the i-th token in the case of knowing the previous i - 1 tokens in the process of generating the answer, 1 ≤ i ≤ N, and N is the total number of tokens in the answer.
[0013] Optionally, the calling of the first type of network model to perform question-and-answer pair generation processing on the original dataset of the target vertical field to obtain an initial dataset in the form of question-and-answer pairs includes: the first type of network model extracts questions from the original dataset; obtaining the answer corresponding to each question under the range constraint condition to generate an initial dataset in the form of question-and-answer pairs; where the range constraint condition represents extracting the answer corresponding to the question from the original dataset.
[0014] Optionally, after calling the first type of network model to perform question-and-answer pair generation processing on the original data set of the target vertical domain to obtain an initial data set in the form of question-and-answer pairs, the method further includes: calling a third type of network model to perform text migration processing on the initial data set, where the text migration processing refers to adjusting the answer part in the question-and-answer pairs based on the language description style of the user group. After the text migration processing, the question-and-answer pairs in the initial data set are more biased towards the language description style of the user group in terms of language description style.
[0015] In a second aspect, an embodiment of the present invention provides a training data set construction device, where the device includes:
[0016] A first processing unit, configured to call a first type of network model to perform question-and-answer pair generation processing on the original data set of the target vertical domain to obtain an initial data set in the form of question-and-answer pairs;
[0017] A second processing unit, configured to perform screening processing on the question-and-answer pairs in the initial data set, and screen out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning data set.
[0018] In a third aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0019] In a fourth aspect, an embodiment of the present invention provides an electronic device, where the electronic device includes: a processor and a memory, and the memory is used to store one or more programs; when the one or more programs are executed by the processor, the above method is implemented.
[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention.
[0023] Figure 2 It is one of the schematic flowcharts of the training data set construction method provided by the embodiment of the present invention.
[0024] Figure 3 This is the second flowchart diagram of the training dataset construction method provided by the embodiments of the present invention.
[0025] Figure 4 This is the unit diagram of the training dataset construction device provided by the embodiments of the present invention.
[0026] In the figure: 10 - processor; 11 - memory; 12 - bus; 13 - communication interface; 501 - first processing unit; 502 - second processing unit. Detailed implementation manners
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0028] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first" and "second" are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0030] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or sequence between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0031] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0032] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "arrangement" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0033] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0034] The embodiments of the present invention provide an electronic device, which can be a mobile phone device, a computer device, and a server device. Please refer to Figure 1 , the structural schematic diagram of the electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected through the bus 12, and the processor 10 is used to execute the executable module stored in the memory 11, such as a computer program.
[0035] The processor 10 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the training dataset construction method can be completed by the integrated logic circuit in the hardware of the processor 10 or the instructions in software form. The above-mentioned processor 10 can be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it can also be a digital signal processor (Digital Signal Processor, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0036] The memory 11 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0037] The bus 12 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Figure 1 Although only one bidirectional arrow is used in the figure, it does not mean that there is only one bus 12 or only one type of bus 12 .
[0038] The memory 11 is used to store programs, such as programs corresponding to the training data set construction device. The training data set construction device includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or fixed in the operating system (OS) of the electronic device. After receiving the execution instruction, the processor 10 executes the program to implement the training data set construction method.
[0039] Possibly, the electronic device provided by the embodiment of the present invention further includes a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.
[0040] It should be understood that Figure 1 The structure shown is only a schematic diagram of a portion of the electronic device. The electronic device may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0041] A training data set construction method provided by an embodiment of the present invention can be applied to, but not limited to, Figure 1 For detailed procedures, please refer to the electronic equipment shown in Figure 2 ,The training data set construction method includes: S10 and S30, which are described in detail as follows.
[0042] S10, calling the first type of network model to perform question-answer pair generation processing on the original data set of the target vertical field to obtain an initial data set in the form of question-answer pairs.
[0043] Among them, the first type of network model can be a closed-source model, and this closed-source model can but is not limited to using GPT4. The original dataset includes any one or more of document data, text data, table data, and image data related to the target vertical domain. The target vertical domain can but is not limited to the kitchen domain, cooking domain, autonomous driving domain, tourism domain, etc. More specifically, it can be the traditional Chinese medicine diet therapy domain, etc.
[0044] S30. Perform screening processing on the question-and-answer pairs in the initial dataset, and screen out the question-and-answer pairs that meet the preset rules to construct the target fine-tuning dataset.
[0045] Among them, the target fine-tuning dataset is used to combine with the low-rank adaptation training method to fine-tune and optimize the training of the open-source model in the target vertical domain, so as to improve the inference effect of the open-source model in the target vertical domain and improve the text generation quality of the open-source model in the target vertical domain.
[0046] Low-Rank Adaptation (LoRA for short) is a technique for training large language models. It introduces parameter matrices A and B, and the parameter scales of both are much smaller than the original parameters. During training, the original parameters of the large model are frozen. The parameter matrix A reduces the rank of the input high-dimensional features to low-dimensional features, and the parameter matrix B maps the low-dimensional features back to high-dimensional features. Compared with the original parameter matrix, it greatly reduces the number of parameters participating in training without affecting the input and output high-dimensional features.
[0047] In the training dataset construction method provided by the embodiments of the present invention, a first type of network model is called to generate question-and-answer pairs for the original dataset in the target vertical domain to obtain an initial dataset in the form of question-and-answer pairs. By performing screening processing on the question-and-answer pairs in the initial dataset, the question-and-answer pairs with too low quality assessment scores and those that are too simple and not suitable for fine-tuning and optimizing training are excluded, and only the high-quality question-and-answer pairs that meet the preset rules are retained, quickly and efficiently constructing the target fine-tuning dataset for fine-tuning and optimizing training, reducing the difficulty of obtaining training data, and reducing the model training cost.
[0048] On Figure 2 this basis, regarding how to ensure the screening efficiency and effect of question-and-answer pairs and improve the quality and speed of training dataset construction, the embodiments of the present invention also provide an optional implementation manner. Please refer to the following text. S30. Perform screening processing on the question-and-answer pairs in the initial dataset, and screen out the question-and-answer pairs that meet the preset rules to construct the target fine-tuning dataset, including: S31, which is specifically described as follows.
[0049] S31. Call a second type of network model to perform screening processing on the question-and-answer pairs in the initial dataset, and screen out the question-and-answer pairs that meet the preset rules to construct the target fine-tuning dataset.
[0050] Among them, the second type of network model can be an open-source model for preference alignment, an open-source model after preference alignment.
[0051] In an alternative implementation, the direct preference optimization (DPO) method can be adopted. Using a small amount of preference datasets of the target user group designed or labeled in advance, the open-source model is trained for preference alignment to obtain a large model that aligns with the preferences of the vertical domain. By conducting preference alignment training, the output content of the large model can be controlled to a certain extent, making it basically conform to the values and preferences of the target user group, and making its output content conform to the user's description style, values, etc., which can prevent the large model from generating harmful content and being misused.
[0052] Please refer to the following text, which is an example illustration of the differences between preferred answers and non-preferred answers:
[0053] Question: What is the taste of West Lake Fish in Vinegar Sauce?
[0054] Preferred answer:
[0055] West Lake Fish in Vinegar Sauce is a classic Chinese dish, especially popular in the Hangzhou area of Zhejiang Province. Its taste characteristics mainly include:
[0056] Sour and sweet: Since the name of the dish contains "vinegar", sourness is one of its main characteristics. After the fish is cooked with vinegar, the sourness is moderate, with a hint of sweetness, forming a unique sour and sweet taste.
[0057] Fresh and fragrant: The fresh fish itself has a faint freshness, and with the seasoning of ingredients such as scallions and ginger, the taste is even more delicious.
[0058] Smooth and tender: During the cooking process, the fish is finely processed, maintaining a smooth and tender texture, making it very comfortable to eat.
[0059] Light: The seasoning of this dish is relatively light, not greasy, and is very suitable for people who like light flavors.
[0060] Generally speaking, West Lake Fish in Vinegar Sauce is a delicious food with a rich taste and unique flavor, suitable for various occasions.
[0061] Non-preferred answer:
[0062] West Lake Fish in Vinegar Sauce is a classic Hangzhou dish, with fresh grass carp as the main ingredient, paired with a special vinegar sauce. The taste is sour with a hint of sweetness, fresh and delicious. The fish is tender, and the sourness of the vinegar sauce can well enhance the freshness of the fish, making the overall taste more rich. Generally speaking, the texture of this dish is soft, and the sourness is moderate, which can stimulate the appetite without being too pungent, suitable for most people's tastes.
[0063] It should be understood that in the Direct Preference Optimization (DPO) approach, a small dataset of preferences from a pre-designed or labeled target user group is used to perform preference alignment training on an open-source model, enabling the large model to be more inclined to output answers preferred by the target user group. Moreover, the large model has the ability to evaluate the quality of questions / answers, which is directly used to filter the dataset after style alignment and exclude samples with low text quality or overly simple questions / answers.
[0064] The reason for using the second type of network model to screen the Q&A pairs in the initial dataset is that even models with extremely large numbers of parameters like GPT-4 may perform slightly worse in a specific vertical domain, especially in some niche fields (such as traditional Chinese medicine and biomolecular science), where they may produce hallucinations or give incorrect answers, and this problem cannot be avoided even after text style transfer. Using texts with overly simple answers or incorrect answers for fine-tuning will skew the output distribution of the large model after training, causing it to output incorrect characters or prematurely output the symbol, resulting in the end of the output and the user not getting a satisfactory result.
[0065] Please refer to the following. In the field of traditional Chinese medicine diet therapy, an example process of training the original open-source large model with DPO based on manually labeled preference alignment data is as follows:
[0066] Question: Why do middle-aged people like to drink wolfberry water?
[0067] Preferred answer:
[0068] The reasons why middle-aged people like to drink wolfberry water can be analyzed from multiple perspectives. The following are some possible reasons:
[0069] 1. Enhanced health awareness
[0070] Health preservation concept: As they age, middle-aged people pay more attention to health and health preservation. Wolfberry is considered to have various health care effects, such as enhancing immunity, anti-aging, and improving eyesight, so it has become one of their choices for daily health preservation.
[0071] Disease prevention: Middle-aged people start to face more health problems, such as high blood pressure and diabetes. Wolfberry is considered helpful for regulating blood sugar and blood pressure, so it has become a way for them to prevent diseases.
[0072] 2. Cultural tradition
[0073] Traditional Chinese medicine theory: Wolfberry is regarded as a good medicine for nourishing the liver and kidneys, improving eyesight, and moistening the lungs in traditional Chinese medicine theory. Middle-aged people may be influenced by traditional Chinese medicine health preservation culture and believe that wolfberry water helps to regulate the body and maintain health.
[0074] Family inheritance: Some middle-aged people may have been influenced by their families since childhood. The elders in their families have the habit of drinking wolfberry water, and this habit continues after they reach adulthood.
[0075] 3. Taste and convenience
[0076] Mild taste: Wolfberry water has a mild taste with a faint sweetness, making it suitable for daily consumption. Middle-aged people may prefer this mild beverage rather than overly stimulating drinks.
[0077] Easy to make: Wolfberry water is simple to prepare. Just soak wolfberries in water, which is suitable for busy middle-aged people.
[0078] 4. Social and psychological factors
[0079] Social recognition: In some social situations, drinking wolfberry water may be regarded as a healthy and tasteful choice. Middle-aged people may hope to show their attitude towards life and health awareness in this way.
[0080] Psychological comfort: Even though the health benefits of wolfberries have not been fully scientifically proven, middle-aged people may obtain a certain psychological comfort from drinking wolfberry water, believing that they are making efforts for their health.
[0081] 5. Economic factors
[0082] Moderate price: The price of wolfberries is relatively moderate, suitable for the consumption level of middle-aged people, making it an affordable option for their daily health care.
[0083] Summary
[0084] There are multiple reasons why middle-aged people like to drink wolfberry water, including enhanced health awareness, influence of cultural traditions, taste and convenience, social and psychological factors, and economic factors. These factors work together to make wolfberry water a common beverage in the daily life of middle-aged people.
[0085] Non-preferred answers: The reasons why middle-aged people like to drink wolfberry water may be as follows:
[0086] Increased awareness of health preservation: Middle-aged people generally pay more attention to health and health preservation. Wolfberries are regarded as a beneficial medicinal material in traditional Chinese medicine, with effects such as tonifying the kidney and essence, nourishing the liver and improving eyesight. Drinking wolfberry water is considered to be able to enhance immunity and improve health conditions.
[0087] Antioxidant effect: Wolfberries contain rich antioxidants such as vitamin C and carotenoids, which can help reduce the damage of free radicals to the body and delay aging, which is attractive to middle-aged people who pay attention to beauty and anti-aging.
[0088] Convenient and easily accessible: Chinese wolfberries are readily available in the market at a moderate price, and making wolfberry water is very simple, just by brewing with hot water, which provides convenience for busy middle-aged people.
[0089] Taste and habit: Wolfberry water has a mild taste and is easily acceptable to most people. In addition, influenced by the long-standing health preservation culture, many people have developed the habit of drinking wolfberry water.
[0090] Cultural influence: In many Asian cultures, especially in China, the habit of drinking wolfberries has taken root in people's hearts. Many people have been influenced by their families or elders since childhood and believe that drinking wolfberry water is beneficial to the body.
[0091] In short, drinking wolfberry water has become part of the daily health preservation of many middle-aged people, and they hope to maintain good health in this way.
[0092] It can be seen that the preference of the answer is more hierarchical and the description is more comprehensive and clear. The former answer is the one that the large model hopes to provide to the user. Through similar data above, the preference of the large model can be trained.
[0093] Based on the above, for the specific process of screening and processing by the second type of network model, the embodiments of the present invention also provide an optional processing method. Please refer to the following. S31. Call the second type of network model to screen and process the question-and-answer pairs in the initial data set, and screen out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning data set, including: S311 and S312, specifically as follows.
[0094] S311. Obtain the preference score and / or perplexity of each question-and-answer pair in the initial data set.
[0095] Among them, the preference score includes a question difficulty evaluation score and an answer quality evaluation score.
[0096] S312. According to the preference score and / or perplexity of the question-and-answer pair, screen out the question-and-answer pairs that meet the preset rules to construct a target fine-tuning data set.
[0097] Optionally, filter out the question-and-answer pairs that do not meet the preset rules. The question-and-answer pairs that do not meet the preset rules are those with a preference score lower than the first threshold and / or a perplexity lower than the second threshold.
[0098] The purpose of screening and processing the question-and-answer pairs in the initial data set is to obtain answers that the large model considers to be of high quality and complex, and the preference alignment with the large model has a relatively low acquisition difficulty, relying only on partially manually labeled preference alignment data, and there is only a single loss function in the training process, which is relatively easy to converge. Therefore, it can be widely used in various vertical fields with answer preferences.
[0099] Optionally, obtain the preference scores for each question-answer pair in the initial dataset, including: S311A, S311B, and S311C, as follows.
[0100] S311A, evaluate the difficulty of the question in the question-answer pair to obtain the question difficulty evaluation score.
[0101] S311B, obtain the quality evaluation index of the answer in the question-answer pair.
[0102] Among them, the quality evaluation index includes any one or more of the detail index, correctness index, description tone style index, and standardization index.
[0103] S311C, determine the answer quality evaluation score according to the quality evaluation index.
[0104] Optionally, obtain the perplexity of each question-answer pair in the initial dataset, including: S311D, as follows.
[0105] S311D, determine the perplexity of the question-answer pair according to the probability of each token in the answer of the question-answer pair.
[0106] Among them, the token probability represents the probability of predicting the i-th token given the previous i - 1 tokens during the process of generating the answer, where 1 ≤ i ≤ N and N is the total number of tokens in the answer.
[0107] Optionally, the formula for perplexity is:
[0108]
[0109] Among them, perplexity(W) represents the perplexity, P(ω i |ω 1 ,ω 2 ,...,ω i-1 ) represents the token probability corresponding to the i-th token. W represents an answer obtained after inputting the question into the second type of network model, and N is the total number of tokens in the answer, that is, the number of tokens included in this answer. The perplexity can be understood as the probability of the model generating a certain corpus, that is, the probability of the model generating this sentence. The reason for applying the (-1 / N) power is to consider the influence of the corpus length. If a sentence is longer, the probability of this sentence appearing may be lower (for example, for the two sentences "What do you want to eat" and "Do you want to eat braised crucian carp today", the former has a very high probability of appearing). The (-1 / N) power is equivalent to a penalty factor. For a number in the range of (0, 1), after applying the (-1 / N) power, the larger N is, the smaller the value after application. A question corresponding to a higher perplexity of the generated answer indicates a higher complexity of the question. Therefore, question-answer pairs with high perplexity are preferentially retained. ωi represents the token corresponding to the i-th smallest phrase unit, P(ω i |ω 1 , ω 2 ,..., ω i-1 ) represents the probability of predicting the i-th token based on the previous i - 1 tokens.
[0110] Based on Figure 2 , the embodiments of the present invention also provide an alternative implementation manner on how to quickly and accurately construct the initial dataset. Please refer to the following. S10, call the first type of network model to perform question - answer pair generation processing on the original dataset of the target vertical domain to obtain the initial dataset in the form of question - answer pairs, including: S11 and S12, specifically as follows.
[0111] S11, the first type of network model extracts questions from the original dataset.
[0112] S12, obtain the answer corresponding to each question under the range constraint conditions to generate the initial dataset in the form of question - answer pairs.
[0113] Among them, the range constraint condition means extracting the answer corresponding to the question from the original dataset.
[0114] It should be noted that when the first type of network model obtains the answer corresponding to each question under the range constraint conditions, there is also a certain possibility of exceeding the range constraint conditions.
[0115] Please refer to the following. The following is an example description of the question - answer pairs obtained by the first type of network model (which can but is not limited to the closed - source large model - GPT4).
[0116] Question:
[0117] What is the origin story of the dish Buddha Jumping Over the Wall?
[0118] Answer:
[0119] Buddha Jumping Over the Wall is a famous traditional dish originating from Fuzhou, Fujian Province and belongs to Fujian cuisine. Its name and origin story are full of cultural tales.
[0120] Origin story:
[0121] The name of Buddha Jumping Over the Wall is said to come from its extremely attractive aroma. It is said that during the Qing Dynasty, a famous chef in Fuzhou cooked a dish that combined various delicacies from mountains and seas in the imperial court or the home of a high - ranking official. Its ingredients were rich, including abalone, sea cucumber, shark fin, scallop, dried mushroom, pigeon egg, chicken, pig tendon, ham, etc., and was simmered with seasonings such as Shaoxing yellow rice wine. This dish was slowly stewed over a long time, and the aroma spread far away.
[0122] This dish emphasizes the combination of various precious ingredients and cooking techniques. It is famous for its delicious taste and rich nutrition, becoming one of the representative dishes of Fujian cuisine. There is also a praise saying, "When the altar is opened, the meaty fragrance wafts to the neighbors, and even the Buddha abandons meditation and jumps over the wall for it."
[0123] It should be noted that when using the first type of network model (which can be, but is not limited to, a closed-source large model - GPT4) to generate question-and-answer pairs from the original dataset in the target vertical domain, the distribution of the obtained initial dataset may deviate from that of the open-source large model currently expected to be trained. This will cause a certain difference between the output result of the fine-tuned large model and the expected correct result when directly using this data as training samples.
[0124] To address the above situation, an alternative implementation is also provided in the embodiments of the present invention. Please refer to Figure 3 , after using the first type of network model to generate question-and-answer pairs from the original dataset in the target vertical domain to obtain an initial dataset in the form of question-and-answer pairs, the training dataset construction method further includes: S20, which is specifically described as follows.
[0125] S20, use the third type of network model to perform text migration processing on the initial dataset.
[0126] Among them, text migration processing refers to adjusting the answer part in the question-and-answer pairs based on the language description style of the user group.
[0127] Among them, the third type of network model can be an original open-source model. As the basic model for preference alignment training, after preference alignment training, a second type of network model is obtained.
[0128] It should be understood that after text migration processing, the question-and-answer pairs in the initial dataset are more biased towards the language description style of the user group in terms of language description style.
[0129] Please refer to the following text, which is an example description of text migration processing by the third type of network model (which can be, but is not limited to, the open-source large model Qwen2.5 - 14B).
[0130] Question:
[0131] What is the origin story of the Buddha Jumping Over the Wall dish?
[0132] Answer:
[0133] Buddha Jumping Over the Wall is a traditional famous dish originating from Fuzhou, Fujian Province, belonging to Fujian cuisine. It is famous for its complex production process and diverse ingredient selection. It usually includes precious ingredients such as sea cucumbers, abalones, shark fins, pigeon eggs, tendon of pig's trotter, cuttlefish, and ham, and is slowly stewed with seasonings such as Shaoxing wine.
[0134] There are several theories about the origin of the name "Buddha Jumps Over the Wall". One of the more popular stories is that it happened during the Qing Dynasty, when a restaurant owner in Fuzhou invented this dish. Legend has it that the aroma of this dish was so tempting that even the monks in the temple could not help but jump over the wall to taste it, hence the name "Buddha Jumps Over the Wall". Another theory is that the deliciousness of this dish made the monks forget the rules and regulations, and even jumped over the wall and left the temple to taste it.
[0135] In fact, the true origin of "Buddha Jumps Over the Wall" may not be so dramatic. According to some historical records, this dish first appeared in the late 19th century and was created by an official or a chef in Fuzhou. It gradually became a high-end dish in Fuzhou and even the whole country.
[0136] As time goes by, Buddha Jumps Over the Wall has also undergone improvements and developments. While maintaining the traditional flavor, the modern version of Buddha Jumps Over the Wall also pays more attention to the concept of nutritional matching and healthy eating. In some important banquets, Buddha Jumps Over the Wall is still one of the important dishes that demonstrates class and taste.
[0137] From the above examples, we can see that the content of the two large models is not much different, but the style and description of the answers are different: GPT4 is more inclined to directly answer the user's questions, while Qwen2.5-14B will make a more detailed description based on the answer. The reason for the above results is that according to the principle of the language model, it is based on the known text token (similar to a word) to predict the next token from the token library. If the corpus text distribution used in the pre-training stage is different, the token predicted in the inference stage is likely to be different, and GPT4 is more inclined to use multiple languages as pre-training text (mainly English). Qwen2.5-14B's service objects are mostly Chinese-speaking regions, and the proportion of Chinese pre-training corpus is relatively higher, which obviously makes Qwen2.5-14B's style description closer to the expression of Chinese user needs; in addition, the fine-tuning, alignment, and training convergence of the model will also affect the answer style of the large model. Therefore, in order to make the fine-tuning stage of the open source big model as aligned as possible with the pre-training stage of the big model, it is necessary to adopt the text transfer method. For the question and answer pairs generated by the closed-source big model, write prompts and call the open source big model to be fine-tuned to rewrite the answer content corresponding to the question. At this time, the answer written by the open source big model is aligned with the pre-training stage, and the answer contains vertical domain knowledge, so the training effect is better.
[0138] The following is an example of the whole process of obtaining training data in the field of Chinese medicine diet therapy based on the training data set construction method provided in an embodiment of the present invention:
[0139] Call the first type of network model to generate question-and-answer pairs from the original dataset in the target vertical domain to obtain the initial dataset in the form of question-and-answer pairs, the initial version of the Q&A set for traditional Chinese medicine diet therapy data. Example of the generated data:
[0140] Q1: What benefits can pears have in traditional Chinese medicine?
[0141] A1: In traditional Chinese medicine theory, pears are considered to have multiple health benefits, mainly including the following aspects:
[0142] Moistening the lungs and relieving cough: Pears have the effects of moistening dryness, clearing heat, and promoting the production of body fluids, and are often used to relieve coughs caused by dryness and heat. Especially in autumn and winter when the air is dry, pears can help moisten the lungs and relieve discomfort.
[0143] Clearing heat and resolving phlegm: Pears are rich in juice and can help clear heat and resolve phlegm, and have an adjuvant effect on heat coughs and sore throats.
[0144] Promoting the production of body fluids and quenching thirst: Pears contain a large amount of water and natural sugars, which can help supplement body fluids, relieve thirst, and increase saliva secretion.
[0145] Promoting digestion: Pears are rich in dietary fiber, which helps promote intestinal peristalsis, aids digestion, and prevents constipation.
[0146] Protecting the liver and promoting diuresis: In traditional Chinese medicine, pears have a certain protective effect on the liver and can also promote urine excretion, helping to detoxify.
[0147] In traditional Chinese medicine applications, pears can be eaten alone or paired with other ingredients such as tremella and rock sugar, boiled into soup or made into stewed dishes to enhance their effect of moistening the lungs and relieving dryness. However, it should be noted that pears are cold in nature, and those with weak spleen and stomach or diarrhea should not eat too much. When applying specifically, it should be eaten in moderation according to individual physique and health status.
[0148] For the above Q&A, call the third type of network model to perform text migration processing on the initial dataset. Example of the data after text migration:
[0149] The following content is an instruction and the corresponding original answer. Please modify the original answer based on your own style and answer the question:
[0150] Instruction: What benefits can pears have in traditional Chinese medicine?
[0151] Original answer: In traditional Chinese medicine theory, pears are considered to have multiple health benefits, mainly including the following aspects:
[0152] Moistening the lungs and relieving cough: Pears have the effects of moistening dryness, clearing heat, and promoting the production of body fluids, and are often used to relieve coughs caused by dryness and heat. Especially in autumn and winter when the air is dry, pears can help moisten the lungs and relieve discomfort.
[0153] Clearing heat and resolving phlegm: Pears are rich in juice, which can help clear heat and resolve phlegm, and has an adjuvant effect on febrile cough and sore throat.
[0154] Promoting fluid production to quench thirst: Pears contain a large amount of water and natural sugars, which can help supplement body fluids, relieve thirst, and increase saliva secretion.
[0155] Promoting digestion: Pears are rich in dietary fiber, which helps promote intestinal peristalsis, aids digestion, and prevents constipation.
[0156] Protecting the liver and promoting diuresis: Traditional Chinese medicine believes that pears have a certain protective effect on the liver, and can also promote urine excretion, helping to detoxify the body.
[0157] In traditional Chinese medicine applications, pears can be eaten alone or paired with other ingredients such as tremella and rock sugar, cooked into soup or made into stews to enhance their effect of moistening the lungs and relieving dryness. However, it should be noted that pears are cool in nature, and those with weak spleen and stomach or diarrhea should not eat too much. When applying specifically, it should be consumed in moderation according to individual constitution and health status.
[0158] Answer, the answer obtained by the third type of network model (Qwen2.5 - 14B), that is, the text migration result is as follows:
[0159] In the wisdom of traditional Chinese medicine, pears are not only a delicious fruit, but also a versatile "health assistant". Its benefits are extensive and profound, mainly reflected in the following aspects:
[0160] Moistening the lungs and relieving cough: Pears are known as natural "lung moistening agents". Their properties of moistening dryness, clearing heat, and promoting fluid production make them an ideal choice for relieving coughs caused by dryness. Especially in autumn and winter when the air becomes dry, pears can gently moisten your lungs and keep you away from cough troubles.
[0161] Clearing heat and resolving phlegm: The rich juice of pears not only quenches thirst, but also helps clear heat and resolve phlegm. For those troubled by febrile cough or sore throat, pears are undoubtedly a good medicine.
[0162] Promoting fluid production to quench thirst: Pears are rich in water and natural sugars, which can quickly supplement body fluids, relieve thirst, keep your mouth moist, and naturally increase saliva secretion.
[0163] Promoting digestion: The dietary fiber in pears is like a little helper in the intestines, which can promote intestinal peristalsis, aid digestion, prevent constipation, and keep your digestive system smooth.
[0164] Protecting the liver and promoting diuresis: In traditional Chinese medicine, pears have a certain protective effect on the liver, and at the same time can promote urine excretion, helping the body detoxify, allowing your liver and kidneys to work easily.
[0165] In the dietary therapy of traditional Chinese medicine, pears can be enjoyed alone or paired with other ingredients such as tremella and rock sugar to make delicious soups or stews, further enhancing their effect of moistening the lungs and relieving dryness. However, pears are cold in nature, so friends with weak spleen and stomach or diarrhea should consume them in moderation to avoid overeating. When enjoying the deliciousness and health benefits of pears, remember to make a wise choice according to your own constitution and health condition.
[0166] It can be seen that the content of the answer after text style alignment is basically the same, but the description is richer, the expression is more diverse, and it potentially aligns with the pre-trained text of the open-source large model.
[0167] After text migration, the second type of network model is called to screen the question-and-answer pairs in the initial dataset, and the question-and-answer pairs that meet the preset rules are selected to construct an example of the target fine-tuning dataset as follows:
[0168] The following content is an instruction and the corresponding answer. Please evaluate the answer corresponding to the instruction and score it according to the preference score and / or perplexity of the question. The score value is an integer from 0 to 10, and the higher the score, the more difficult the question is to answer and the higher the quality of the answer:
[0169] Instruction:
[0170] What benefits can pears have in traditional Chinese medicine?
[0171] Answer:
[0172] In the wisdom of traditional Chinese medicine, pears are not only a delicious fruit but also a versatile "health assistant". Its benefits are extensive and profound, mainly reflected in the following aspects:
[0173] Moistening the lungs and relieving cough: Pears are known as natural "lung moistening agents". Their properties of moistening dryness, clearing heat, and promoting fluid production make them an ideal choice for relieving coughs caused by dryness and heat. Especially in autumn and winter when the air becomes dry, pears can gently moisten your lungs and keep you away from cough troubles.
[0174] Clearing heat and resolving phlegm: The rich juice of pears not only quenches thirst but also helps clear heat and resolve phlegm. For those troubled by heat-induced coughs or sore throats, pears are undoubtedly a good remedy.
[0175] Promoting fluid production and quenching thirst: Pears are rich in water and natural sugars, which can quickly replenish body fluids, relieve thirst, keep your mouth moist, and naturally increase saliva secretion.
[0176] Promoting digestion: The dietary fiber in pears is like a little helper in the intestines, which can promote intestinal peristalsis, aid digestion, prevent constipation, and keep your digestive system smooth.
[0177] Liver protection and diuresis: In traditional Chinese medicine, pears have a certain protective effect on the liver. At the same time, they can promote urine excretion, help the body detoxify, and enable your liver and kidneys to work easily.
[0178] In the dietary therapy of traditional Chinese medicine, pears can be enjoyed alone or paired with other ingredients such as tremella and rock sugar to make delicious soups or stews, further enhancing their effect of moistening the lungs and relieving dryness. However, pears are cold in nature, so friends with weak spleen and stomach or diarrhea should consume them in moderation to avoid overeating. When enjoying the deliciousness and health benefits of pears, remember to make a wise choice according to your own constitution and health status.
[0179] Evaluation results:
[0180] An example of the evaluation results of the preference alignment large model is as follows:
[0181] Perplexity: 6
[0182] Preference score: 9
[0183] Explanation: Perplexity (6). This question involves traditional Chinese medicine theory and the health benefits of pears, requiring certain professional knowledge and understanding of dietary therapy in traditional Chinese medicine. Although not very complex, it is not a simple question either, so the difficulty is moderate.
[0184] Preference score (9). The answer details and comprehensively explains the various benefits of pears in traditional Chinese medicine, including moistening the lungs and relieving cough, clearing heat and resolving phlegm, promoting fluid production and quenching thirst, promoting digestion, and liver protection and diuresis, etc. The answer content is rich, the logic is clear, the language expression is fluent, and practical eating suggestions and precautions are provided. The overall quality is very high. The only point deducted is that the nutritional components of pears are not mentioned, but this does not affect the excellence of the overall answer.
[0185] It is considered that the difficulty of this question is moderate and the answer quality is high, so it can be used as a fine-tuning data set for the large model.
[0186] Perplexity calculation
[0187] Since this answer is relatively long, only partial conditional token probabilities of the large model are shown, that is, P(Y|X), where Y is the token to be predicted and X is the known token:
[0188] P(more|In the wisdom of traditional Chinese medicine, pears are not only a delicious fruit,) = 0.78;
[0189] P(cough|In the wisdom of traditional Chinese medicine, pears are not only a delicious fruit, but also a versatile "health assistant". Its benefits are extensive and profound, mainly reflected in the following aspects: Moistening the lungs and stopping) = 0.95;
[0190] Finally, according to the formula, the calculated result is 157.78, which belongs to the normal perplexity range (50 - 1000), so this Q&A pair can be retained.
[0191] Please refer to Figure 4 , Figure 4 a training dataset construction device provided by an embodiment of the present invention. Optionally, this training dataset construction device is applied to the electronic device described above.
[0192] The training dataset construction device includes: a first processing unit 501 and a second processing unit 502.
[0193] The first processing unit 501 is used to call a first type of network model to perform Q&A pair generation processing on the original dataset of the target vertical domain, so as to obtain an initial dataset in the form of Q&A pairs.
[0194] The second processing unit 502 is used to screen the Q&A pairs in the initial dataset, and screen out the Q&A pairs that meet the preset rules, so as to construct a target fine-tuning dataset.
[0195] It should be noted that the training dataset construction device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiment.
[0196] An embodiment of the present invention also provides a storage medium, which stores computer instructions and programs. When the computer instructions and programs are read and run, they execute the training dataset construction method of the above embodiment. The storage medium may include memory, flash memory, registers, or a combination thereof, etc.
[0197] The following provides an electronic device, which can be a mobile phone device, a computer device, and a server device. As Figure 1 shown, this electronic device can implement the above training dataset construction method; specifically, this electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs. When the one or more programs are executed by the processor 10, the training dataset construction method of the above embodiment is executed.
[0198] In summary, a method, apparatus, storage medium, and electronic device for constructing a training data set provided by an embodiment of the present invention generate question-and-answer pairs from an original data set in a target vertical domain by invoking a first type of network model to obtain an initial data set in the form of question-and-answer pairs. By screening the question-and-answer pairs in the initial data set, question-and-answer pairs with too low quality assessment scores and those that are too simple and not suitable for fine-tuning optimization training are excluded, and only high-quality question-and-answer pairs that meet the preset rules are retained, so as to quickly and efficiently construct a target fine-tuning data set for fine-tuning optimization training, reducing the difficulty of obtaining training data and the cost of model training.
[0199] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0200] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference numerals in the claims should not be regarded as limiting the corresponding claims.
Claims
1. A method for constructing a training data set, characterized in that: The method comprises: Call the first type of network model to generate question-answer pairs for the original data set in the target vertical field to obtain an initial data set in the form of question-answer pairs; The question-answer pairs in the initial data set are screened to select the question-answer pairs that meet the preset rules to construct a target fine-tuning data set.
2. The method for constructing a training data set according to claim 1, wherein: The screening process of the question-answer pairs in the initial data set to select the question-answer pairs that meet the preset rules to construct the target fine-tuning data set includes: The second type of network model is called to screen the question-answer pairs in the initial data set, and the question-answer pairs that meet the preset rules are screened out to construct a target fine-tuning data set.
3. The method for constructing a training data set according to claim 2, wherein: The calling of the second type of network model to screen the question-answer pairs in the initial data set, and screening out the question-answer pairs that meet the preset rules to construct a target fine-tuning data set, includes: Obtaining a preference score and / or perplexity of each question-answer pair in the initial data set, wherein the preference score includes a question difficulty assessment score and an answer quality assessment score; According to the preference scores and / or perplexities of the question-answer pairs, the question-answer pairs that meet the preset rules are screened out to construct a target fine-tuning dataset.
4. The method for constructing a training data set according to claim 3, wherein: The obtaining of a preference score for each question-answer pair in the initial data set includes: Performing a difficulty assessment on the question in the question-answer pair to obtain a difficulty assessment score for the question; Obtaining a quality assessment index of the answer in the question-answer pair, wherein the quality assessment index includes any one or more of a detail index, a correctness index, a description tone and style index, and a standardization index; The answer quality assessment score is determined according to the quality assessment indicator.
5. The method for constructing a training data set according to claim 3, wherein: Obtain the perplexity of each question-answer pair in the initial data set, including: Determining the perplexity of the question-answer pair according to the probability of each word in the answer to the question-answer pair; The word-gram probability represents the probability of predicting the i-th word-gram when the first i-1 words are known in the process of generating the answer, 1≤i≤N, N is the total number of words in the answer.
6. The method for constructing a training data set according to claim 1, wherein: The calling of the first type of network model to perform question-answer pair generation processing on the original data set in the target vertical field to obtain an initial data set in the form of question-answer pairs includes: The first type of network model extracts questions from the original data set; Obtain the answer to each question under the scope constraint to generate an initial dataset in the form of question-answer pairs; The range constraint condition represents extracting the answer corresponding to the question from the original data set.
7. The method for constructing a training data set according to any one of claims 1 to 6, characterized in that: After calling the first type of network model to perform question-answer pair generation processing on the original data set in the target vertical field to obtain an initial data set in the form of question-answer pairs, the method further includes: The third type of network model is called to perform text migration processing on the initial data set, wherein the text migration processing refers to adjusting the answer part in the question-answer pair based on the language description style of the user group.
8. A training data set construction device, characterized in that: The device comprises: A first processing unit is used to call the first type of network model to perform question-answer pair generation processing on the original data set in the target vertical field to obtain an initial data set in the form of question-answer pairs; The second processing unit is used to screen the question-answer pairs in the initial data set to select the question-answer pairs that meet the preset rules to construct a target fine-tuning data set.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store one or more programs; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Knowledge updating method, device and equipment of model based on industrial internet, medium and program product
CN120973892A
Data screening method and device, equipment and medium
CN121661443A