Method, device, electronic device and storage medium for constructing complex instruction training data for model training

By extracting and presetting content, answers and questions, combining seed instructions and generalizing them, the problems of low efficiency in building complex instruction training data and poor answer quality in the existing technology are solved, and the generation of high-quality training data and intelligent improvement of the model are achieved.

CN118228839BActive Publication Date: 2025-05-06BEIJING FACE WALL INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202410494963.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-05-06
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency and low answer quality when building training data of complex instruction types, making it difficult to effectively deal with extremely complex instruction data structures.

Method used

By obtaining the initial training data of the large language model, the initial content, the initial answers and the initial questions are extracted, the preset content, the preset answers and the preset questions are determined based on these contents, and the seed instructions are combined to form, and high-quality training data is obtained through generalization technology.

Benefits of technology

It realizes the expansion of the initial training data of large language models, generates high-quality and diverse training data, and improves the richness of the training data and the understanding of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118228839B_ABST
    Figure CN118228839B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a method, device, electronic device and storage medium for constructing complex instruction training data for model training, which relates to the field of artificial intelligence and can generate high-quality training data. The method includes: obtaining initial training data of a large language model, the initial training data including initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; determining preset content, preset answers and preset questions based on the initial content, initial answers and initial questions; combining preset tasks and the preset content, preset answers and preset questions to obtain seed instructions; generalizing the seed instructions to obtain training data. The present invention is suitable for scenarios where complex instruction training data is constructed for large language models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and specifically to a method, device, electronic device and storage medium for constructing complex instruction training data for model training. Background Art

[0002] In Large Language Model (LLM) training, the ability to understand and follow complex instructions is very important, and building training data for complex instruction types requires a lot of manpower and time. Existing training data construction solutions have the following shortcomings: there are limitations when processing specific types of data, and it is difficult to effectively process extremely complex instruction data structures; because the questions and answers of the constructed training data are completely dependent on the generation of existing language models, the answer quality is low.

[0003] Therefore, for training data of complex instruction types, a solution needs to be designed to efficiently produce a large amount of high-quality training data. Summary of the invention

[0004] In view of this, the present application provides a method, device, electronic device and storage medium for constructing complex instruction training data for model training to generate high-quality training data.

[0005] In a first aspect, an embodiment of the present invention provides a method for constructing complex instruction training data for model training, comprising: obtaining initial training data of a large language model, the initial training data comprising initial content, an initial answer extracted from the initial content, and an initial question corresponding to the initial answer; determining preset content, preset answers, and preset questions based on the initial content, the initial answer, and the initial question; combining a preset task with the preset content, the preset answer, and the preset question to obtain a seed instruction; and generalizing the seed instruction to obtain training data.

[0006] In a specific implementation manner, determining preset content, preset answers and preset questions based on the initial content, initial answers and initial questions includes: extracting the initial content, initial answers and initial questions from the initial training data; determining preset content based on the initial content, determining preset answers based on the initial answers, and determining preset questions based on the initial questions in accordance with a target format.

[0007] In a specific implementation scheme, the method for determining the preset content includes: obtaining the initial content; using the id of the initial content as the first field content, and configuring the corresponding first field to represent the first field content; using the initial content as the second field content, and configuring the second field to represent the second field content; combining the first field and the first field content, the second field and the second field content to obtain the preset content.

[0008] In a specific implementation scheme, the method for determining the preset question includes: obtaining the initial question; using the ID of the initial question as the content of the third field, and configuring the corresponding third field to characterize the content of the third field; using the initial question as the content of the fourth field, and configuring the corresponding fourth field to characterize the content of the fourth field; combining the third field and the third field content, and the fourth field and the fourth field content to obtain the preset question.

[0009] In a specific implementation scheme, the method for determining the preset answer includes: obtaining the initial answer; taking the ID of the initial question corresponding to the initial answer as the content of the fifth field, and configuring the corresponding fifth field to represent the content of the fifth field; taking the initial answer as the content of the sixth field, and configuring the corresponding sixth field to represent the content of the sixth field; taking the starting position of the initial answer in the initial content as the content of the seventh field, and configuring the corresponding seventh field to represent the content of the seventh field; combining the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain the preset answer.

[0010] In a specific implementation, generalizing the seed instruction includes: determining a generalization example; and generalizing the sentences in the seed instruction using the large language model based on contextual learning of the generalization example by the large language model.

[0011] In a specific implementation manner, after generalizing the seed instructions, the method further includes: scoring the generalized seed instructions, and selecting the seed instructions whose scoring values ​​exceed a preset threshold as training data.

[0012] In a specific implementation manner, the target format is any one of JSON format, YAML format, and XML format.

[0013] In a specific implementation scheme, the preset task includes a task description and a task guide, the task description includes a task example and a description of the target format, and the task guide sequentially guides the preset content, preset questions, and preset answers.

[0014] In a second aspect, an embodiment of the present invention further provides a device for constructing complex instruction training data for model training, the device for constructing complex instruction training data for model training comprising: an acquisition unit for acquiring initial training data of a large language model, the initial training data comprising initial content, an initial answer extracted from the initial content, and an initial question corresponding to the initial answer; a determination unit for determining preset content, preset answers and preset questions based on the initial content, initial answers and initial questions; a combination unit for combining a preset task and the preset content, preset answers and preset questions to obtain a seed instruction; and a generalization unit for generalizing the seed instruction to obtain training data.

[0015] In a specific implementation scheme, the determination unit includes: an extraction module for extracting the initial content, initial answers and initial questions from the initial training data; a format determination module for determining preset content according to the initial content, determining preset answers according to the initial answers, and determining preset questions according to the initial questions in accordance with the target format.

[0016] In a specific implementation scheme, the format determination module includes a preset content determination sub-block, which is used to: obtain the initial content; use the id of the initial content as the first field content, and configure the corresponding first field to characterize the first field content; use the initial content as the second field content, and configure the second field to characterize the second field content; combine the first field and the first field content, the second field and the second field content to obtain the preset content.

[0017] In a specific implementation scheme, the format determination module also includes a preset question determination sub-block, which is used to: obtain the initial question; use the id of the initial question as the content of the third field, and configure the corresponding third field to characterize the content of the third field; use the initial question as the content of the fourth field, and configure the corresponding fourth field to characterize the content of the fourth field; combine the third field and the third field content, and the fourth field and the fourth field content to obtain the preset question.

[0018] In a specific implementation scheme, the format determination module also includes a preset answer determination sub-block, which is used to: obtain the initial answer; use the id of the initial question corresponding to the initial answer as the content of the fifth field, and configure the corresponding fifth field to represent the content of the fifth field; use the initial answer as the content of the sixth field, and configure the corresponding sixth field to represent the content of the sixth field; use the starting position of the initial answer in the initial content as the content of the seventh field, and configure the corresponding seventh field to represent the content of the seventh field; combine the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain a preset answer.

[0019] In a specific implementation, the generalization unit includes: a generalization example module, used to determine a generalization example; a sentence generalization module, used to generalize the sentence in the seed instruction based on the context learning of the generalization example by the large language model.

[0020] In a specific implementation manner, the generalization unit further includes: a scoring module, which is used to score the generalized seed instructions and select the seed instructions whose scoring values ​​exceed a preset threshold as training data.

[0021] In a specific implementation manner, the target format is any one of JSON format, YAML format, and XML format.

[0022] In a specific implementation scheme, the preset task includes a task description and a task guide, the task description includes a task example and a description of the target format, and the task guide sequentially guides the preset content, preset questions, and preset answers.

[0023] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising: a housing, a processor, a memory, a circuit board and a power supply circuit, wherein the circuit board is placed inside the space enclosed by the housing, and the processor and the memory are arranged on the circuit board; a power supply circuit for supplying power to various circuits or devices of the above-mentioned electronic device; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute any one of the methods for constructing complex instruction training data for model training provided in the embodiments of the present invention.

[0024] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement any method for constructing complex instruction training data for model training provided in an embodiment of the present invention.

[0025] The embodiments of the present invention provide a method, device, electronic device and storage medium for constructing complex instruction training data for model training, which obtains initial training data of a large language model, wherein the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; then based on the initial content, initial answers and initial questions, preset content, preset answers and preset questions are determined; preset tasks and preset content, preset answers and preset questions are combined to obtain seed instructions; and the seed instructions are generalized to obtain training data. This method can expand the initial training data of a large language model, generate high-quality training data, and improve the diversity and richness of training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0027] Figure 1 A flow chart of a method for constructing complex instruction training data for model training provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of a structure of a device for constructing complex instruction training data for model training provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0031] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] Existing training data construction solutions are difficult to effectively handle extremely complex instruction data structures when processing specific types of data. In addition, the questions and answers of the constructed training data are completely dependent on the generation of existing language models, and the answer quality is low. Therefore, for training data of complex instruction types, the first aspect is to Figure 1 As shown, an embodiment of the present invention provides a method for constructing complex instruction training data for model training, which may include:

[0033] S11. Obtaining initial training data of the large language model, where the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers.

[0034] For different natural language tasks, the large language model can perform deep learning based on the corresponding initial training data to more accurately understand the complexity of the language and provide users with more intelligent and personalized services. In this embodiment, the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers. For example, the initial training data can be selected based on CMRC (Chinese Machine Reading Comprehension, Chinese machine reading comprehension data set), where the CMRC data set contains question and answer pairs, as well as related article paragraphs. By allowing the large language model to read and understand article paragraphs to answer questions, the ability to understand natural language can be improved. The initial content, initial questions, and initial answers in the initial training data can be the corresponding article paragraphs, questions, and answer pairs of a certain training data in the CMRC data set.

[0035] S12. Based on the initial content, initial answers and initial questions, determine the preset content, preset answers and preset questions.

[0036] After obtaining the initial training data, preset content, preset answers and preset questions can be constructed based on the initial content, initial answers and initial questions in the initial training data according to the training requirements of the large language model. For example, if the initial content is only a background article and there is no additional information about the background article such as index position, number or label, then additional information about the background article can be added to the initial content. In this way, the preset content includes not only the initial content but also relevant additional information, thereby improving the degree of dataization of the relevant content of the training material and facilitating further data processing to generate training data.

[0037] S13, combining the preset task and the preset content, the preset answer, and the preset question to obtain a seed instruction.

[0038] After determining the preset content, preset answers and preset questions according to the training requirements of the large language model, the preset tasks and preset content, preset answers and preset questions can be further combined to obtain seed instructions. When combining, the corresponding combination method can be designed based on the task type of the preset task and the training purpose of the large language model. For example, in order to facilitate data management of training data, the preset tasks, preset content, preset answers, preset questions and other combination items can be designed in a modular manner to facilitate data maintenance or expansion of each combination item.

[0039] S14. Generalize the seed instruction to obtain training data.

[0040] After obtaining the seed instruction, the preset content, preset answers and preset questions in the seed instruction may be relatively concise or more serious language descriptions. When the large language model trained by the seed instruction interacts with a specific user object, such as a child, its language style may not match the user object, thereby affecting the human-computer understanding and interaction. Therefore, the preset content, preset answers or preset questions in the seed instruction can be generalized to obtain interesting and easy-to-understand preset content, preset answers and preset questions in multiple language expression styles, so that the seed instruction can be expanded and richer training data can be constructed. The large language model trained based on the training data can better interact with the user object and improve the intelligence and personalization of the large language model.

[0041] The method for constructing complex instruction training data for model training provided by an embodiment of the present invention obtains initial training data of a large language model, wherein the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; then based on the initial content, initial answers, and initial questions, preset content, preset answers, and preset questions are determined; preset tasks and preset content, preset answers, and preset questions are combined to obtain seed instructions; and the seed instructions are generalized to obtain training data. This method can expand the initial training data of a large language model, generate high-quality training data, and improve the diversity and richness of training data.

[0042] There may be multiple existing training sets for the large language model of the same task type, and the data structures and content expressions of the multiple existing training sets are often different. In order to facilitate the organization and use of initial training data in different data sets, optionally, in one embodiment of the present invention, step S12 determines preset content, preset answers and preset questions based on the initial content, initial answers and initial questions, including: extracting the initial content, initial answers and initial questions from the initial training data; determining the preset content according to the initial content, determining the preset answer according to the initial answer, and determining the preset question according to the initial question in accordance with the target format.

[0043] For a certain task type of a large language model, different developers provide a first training set and a second training set. According to the present method, no matter it is the first training set represented in the first format or the second training set represented in the second format, it is only necessary to extract the corresponding initial content, initial answer and initial question from a certain initial training data, and convert the initial content into preset content in the corresponding format, convert the initial answer into preset answer in the corresponding format, and convert the initial question into preset question in the corresponding format according to the target format of the present embodiment. Through the present method, when organizing and using the existing training set, it is no longer limited by the format restrictions that may exist in the existing training set, and at the same time improves the modularity of the data and the efficiency of data organization.

[0044] Optionally, in one embodiment of the present invention, a method for determining preset content includes: obtaining initial content; using the id of the initial content as the first field content, and configuring the corresponding first field to characterize the first field content; using the initial content as the second field content, and configuring the second field to characterize the second field content; combining the first field and the first field content, the second field and the second field content to obtain preset content.

[0045] To facilitate efficient data organization, this embodiment uses the id (identifier) ​​of the acquired initial content as the first field content and configures the corresponding first field for representation, uses the initial content as the second field content and configures the second field for representation. The id of the initial content can be the identification information of the entire document of the initial training data in which it is located, or it can be the identification information reconfigured after the initial content is acquired, so as to facilitate data management of the acquired initial content. For example, if the initial content is a background article "Samurai Warriors 3 is the third legitimate sequel to the Samurai Warriors series developed by Koei and ω-force. This work is based on three main stories, namely, "Kanto Three Kingdoms" with Takeda Shingen and others as the main character, "Three Heroes of the Warring States" with Oda Nobunaga and others as the main character, and "Young Warrior of Sekigahara" with Ishida Mitsunari and others as the main character, enriching the plot in the game." The id "DEV_0" of the initial content is used as the first field content, and the first field is configured for representation. For example, the first field can be configured as "id", the above background article is used as the second field content, and the second field is configured for representation. For example, the second field can be configured as "context", and then the first field and the first field content, the second field and the second field content are combined to obtain the preset content as shown in the following table:

[0046] Through the above method, the initial content in the initial training data can be further digitized and modularized, thereby improving the data processing efficiency and data integration utilization of the initial training data.

[0047] Optionally, in one embodiment of the present invention, the method for determining a preset question includes: obtaining an initial question; using the ID of the initial question as the content of the third field, and configuring a corresponding third field to characterize the content of the third field; using the initial question as the content of the fourth field, and configuring a corresponding fourth field to characterize the content of the fourth field; combining the third field and the third field content, and the fourth field and the fourth field content to obtain the preset question.

[0048] When determining the preset question based on the initial question, the data structure of the initial question can also be modularized. For example, in the above example, the initial question obtained from the initial training data is "Which two companies jointly developed "Samurai Warriors 3"?", and the id "DEV_0_QUERY_0" of the initial question can be used as the third field content, and the corresponding third field "id" can be configured for representation. The id of the initial question can be its identification information in the initial training data, or it can be the identification information reconfigured after the initial question is obtained; the initial question "Which two companies jointly developed "Samurai Warriors 3"?" is used as the fourth field content, and the corresponding fourth field "question" is configured for representation, so as to realize data and modularization of the initial question in the initial training data, and then the preset question can be obtained by combining the third field and the third field content, and the fourth field and the fourth field content, as shown in the following table:

[0049] Among them, according to the preset content, there can be multiple different preset questions, and a preset question set can be formed based on multiple preset questions.

[0050] Optionally, in one embodiment of the present invention, a method for determining a preset answer includes: obtaining an initial answer; using the ID of the initial question corresponding to the initial answer as the content of the fifth field, and configuring the corresponding fifth field to represent the content of the fifth field; using the initial answer as the content of the sixth field, and configuring the corresponding sixth field to represent the content of the sixth field; using the starting position of the initial answer in the initial content as the content of the seventh field, and configuring the corresponding seventh field to represent the content of the seventh field; combining the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain a preset answer.

[0051] For example, for the initial answer, the id "DEV_0_QUERY_0" of the corresponding initial question can be used as the fifth field content, and the corresponding fifth field "id" can be configured for representation to build a corresponding relationship between the initial answer and the initial question, and the initial answer "glory and ω-force" can be used as the sixth field content, and the corresponding sixth field "text" can be configured for representation. Furthermore, the starting position of the initial answer in the initial content, such as text sorting 10, can be used as the seventh field content, and the corresponding seventh field "answer_start" can be configured for representation. In this way, the data relationship between the initial answer and the initial content can be further explored, and then the preset answer can be obtained by combining the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content, as shown in the following table:

[0052] Among them, multiple preset answers can be constructed according to different situations, such as preset answers for brief description mode or preset answers for detailed description mode, etc. A preset answer set can be formed based on multiple preset answers to improve the accuracy of the model when the generated training data is applied to a large language model.

[0053] In order to generalize the expression of seed instructions and improve the ability of the trained large language model to understand complex instructions, optionally, in one embodiment of the present invention, the seed instructions are generalized, including: determining generalization examples; based on the context learning of the generalization examples by the large language model, generalizing the sentences in the seed instructions using the large language model.

[0054] In this embodiment, the sentence expressions in the seed instruction can be generalized through the ICL (In-Context Learning) of the large language model. The large language model can understand the input task according to the given task description or several examples and give the result. For example, when generalizing the preset question of the seed instruction in this embodiment, the following generalization example can be determined first:

[0055] Please generalize the following statements:

[0056] Which two companies jointly developed "Samurai Warriors 3"?

[0057] Which company developed Dynasty Warriors 3?

[0058] In the above generalization examples, "Please generalize the following statements:" is the generalization task description, "Which two companies jointly developed "Samurai Warriors 3"?" is the generalization task, and "Which company developed Warriors 3?" is the generalization result. The language expression of the generalization result in the generalization example used to train the large language model can be more concise and more colloquial. Multiple generalization tasks and generalization results can be designed for the generalization example, and different generalization language styles can be designed for different user groups to improve the generalization ability of the large language model after learning and training. In this way, the generalization results of the sentences in the seed instructions using the trained large language model are richer and more diverse, thereby generalizing and expanding the seed instructions to obtain a large number of high-quality complex instructions as training data, further improving the understanding ability of the large language model.

[0059] After generalizing the seed instructions, in order to further screen out high-quality and efficient generalization results as training data, optionally, in one embodiment of the present invention, after generalizing the seed instructions, the process further includes: scoring the generalized seed instructions, and screening out seed instructions with scoring values ​​exceeding a preset threshold as training data. For example, the generalized seed instructions can be scored using a large language model, data with low scores can be discarded, and data with scores exceeding a preset threshold can be screened out as training data, thereby improving the quality of the training data finally generated.

[0060] When determining the preset content according to the initial content, determining the preset answer according to the initial answer, and determining the preset question according to the initial question, the existing data structure can also be used. Optionally, in one embodiment of the present invention, the target format is any one of the JSON format, the YAML format, and the XML format. For example, NLP (Natural Language Processing) data sets of different task types are sorted in JSON format, the preset content, the preset answer, and the preset question are determined, and the seed instructions are obtained by combining them with the preset tasks, and then converted into the required data structure format through the conversion script between JSON, YAML, and XML formats, thereby further improving the structuring and generalization of the seed instruction constituent data.

[0061] Since the preset task needs to clearly indicate the task type and task operation of the training data, when constructing the preset task, optionally, in one embodiment of the present invention, the preset task includes a task description and a task guide, the task description includes a task example and a description of the target format, and the task guide guides the preset content, preset questions, and preset answers in sequence. In this way, the task description can display the task type and task operation method, and the preset content, preset questions, and preset answers can be combined through the guidance prompts. For example, the following templated preset tasks can be designed:

[0062]

[0063]

[0064]

[0065] In the above table, the first cell is the task description, and the second and third cells are task guidance. Among them, the first cell contains a task example and a description of the target format, which can explain and display the task type and task operation method of the preset task in detail. The second cell can be used to guide the display of preset content and preset questions, such as replacing the input in the second cell with preset content and preset questions to achieve a combination of preset content, preset questions and preset tasks. The third cell can be used to guide the display of preset answers, such as replacing the output in the third cell with a preset answer to achieve a combination of preset answers and preset tasks. It can be seen that this templated configuration of preset tasks further improves the modularization, structuring and formatting of data. The seed instructions constructed in this way facilitate further data processing and integration to quickly and efficiently generate a large amount of high-quality training data.

[0066] In a second aspect, an embodiment of the present invention also provides a device for constructing complex instruction training data for model training, which is capable of generating high-quality training data.

[0067] like Figure 2 As shown, the construction device for complex instruction training data for model training provided in an embodiment of the present application may include: an acquisition unit 11, used to acquire initial training data of a large language model, the initial training data including initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; a determination unit 12, used to determine preset content, preset answers and preset questions based on the initial content, the initial answers and the initial questions; a combination unit 13, used to combine preset tasks and preset content, preset answers and preset questions to obtain seed instructions; a generalization unit 14, used to generalize the seed instructions to obtain training data.

[0068] The embodiment of the present invention provides a device for constructing complex instruction training data for model training, which obtains initial training data of a large language model, wherein the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; then based on the initial content, initial answers, and initial questions, preset content, preset answers, and preset questions are determined; the preset tasks and preset content, preset answers, and preset questions are combined to obtain seed instructions; and the seed instructions are generalized to obtain training data. This method can expand the initial training data of a large language model, generate high-quality training data, and improve the diversity and richness of the training data.

[0069] Optionally, in one embodiment of the present invention, the determination unit includes: an extraction module for extracting initial content, initial answers and initial questions from initial training data; a format determination module for determining preset content according to the initial content, determining preset answers according to the initial answers, and determining preset questions according to the initial questions in accordance with a target format.

[0070] Optionally, in one embodiment of the present invention, the format determination module includes a preset content determination sub-block, which is used to: obtain initial content; use the id of the initial content as the first field content, and configure the corresponding first field to characterize the first field content; use the initial content as the second field content, and configure the second field to characterize the second field content; combine the first field and the first field content, the second field and the second field content to obtain preset content.

[0071] Optionally, in one embodiment of the present invention, the format determination module also includes a preset question determination sub-block, which is used to: obtain an initial question; use the id of the initial question as the content of the third field, and configure the corresponding third field to characterize the content of the third field; use the initial question as the content of the fourth field, and configure the corresponding fourth field to characterize the content of the fourth field; combine the third field and the third field content, and the fourth field and the fourth field content to obtain the preset question.

[0072] Optionally, in one embodiment of the present invention, the format determination module also includes a preset answer determination sub-block, which is used to: obtain an initial answer; use the ID of the initial question corresponding to the initial answer as the content of the fifth field, and configure the corresponding fifth field to represent the content of the fifth field; use the initial answer as the content of the sixth field, and configure the corresponding sixth field to represent the content of the sixth field; use the starting position of the initial answer in the initial content as the content of the seventh field, and configure the corresponding seventh field to represent the content of the seventh field; combine the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain a preset answer.

[0073] Optionally, in one embodiment of the present invention, the generalization unit includes: a generalization example module, used to determine the generalization example; a sentence generalization module, used to learn the context of the generalization example based on the large language model, and generalize the sentences in the seed instruction using the large language model.

[0074] Optionally, in one embodiment of the present invention, the generalization unit further includes: a scoring module, configured to score the generalized seed instructions, and select seed instructions with score values ​​exceeding a preset threshold as training data.

[0075] Optionally, in one embodiment of the present invention, the target format is any one of JSON format, YAML format, and XML format.

[0076] Optionally, in one embodiment of the present invention, the preset task includes a task description and a task guide, the task description includes a task example and a description of the target format, and the task guide sequentially guides the preset content, preset questions, and preset answers.

[0077] In a third aspect, an embodiment of the present invention further provides an electronic device.

[0078] like Figure 3 As shown, the electronic device provided in the embodiment of the present application includes: a shell 51, a processor 52, a memory 53, a circuit board 54 and a power supply circuit 55, wherein the circuit board 54 is arranged inside the space enclosed by the shell 51, and the processor 52 and the memory 53 are arranged on the circuit board 54; the power supply circuit 55 is used to supply power to various circuits or devices of the above-mentioned electronic device; the memory 53 is used to store executable program codes; the processor 52 runs the program corresponding to the executable program code by reading the executable program code stored in the memory 53, so as to execute any one of the construction methods of complex instruction training data for model training provided in the embodiments of the present invention.

[0079] The specific execution process of the above steps by the processor 52 and the steps further executed by the processor 52 by running the executable program code can be found in the description of the previous embodiment, which will not be repeated here.

[0080] The above electronic devices exist in many forms, including but not limited to:

[0081] (1) Mobile communication devices: These devices are characterized by their mobile communication functions and their main purpose is to provide voice and data communications. These terminals include: smart phones (such as iPhone), multimedia phones, feature phones, and low-end phones.

[0082] (2) Ultra-mobile personal computer devices: These devices fall into the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, such as iPad.

[0083] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players (such as iPods), handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0084] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general-purpose computer architecture, but because it needs to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0085] (5) Other electronic devices with data interaction functions.

[0086] In the fourth aspect, the embodiments of the present application also provide a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement any method for constructing complex instruction training data for model training provided in the embodiments of the present invention, thereby also being able to achieve the corresponding technical effects, which have been described in detail above and will not be repeated here.

[0087] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0088] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0089] In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0090] For the convenience of description, the above device is described by dividing the functions into various units / modules. Of course, when implementing the present invention, the functions of each unit / module can be implemented in the same or multiple software and / or hardware.

[0091] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0092] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for constructing complex instruction training data for model training, characterized in that: include: Acquire initial training data of a large language model, wherein the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; Determining preset content, preset answers and preset questions based on the initial content, initial answers and initial questions; Combining the preset task and the preset content, preset answer, and preset question to obtain a seed instruction; Generalizing the seed instruction to obtain training data; The determining of the preset content, the preset answer and the preset question based on the initial content, the initial answer and the initial question comprises: extracting the initial content, the initial answer and the initial question from the initial training data; determining the preset content according to the initial content, determining the preset answer according to the initial answer, and determining the preset question according to the initial question in accordance with the target format; The generalizing the seed instruction includes: determining a generalization example; learning the context of the generalization example based on the large language model, and generalizing the sentence in the seed instruction using the large language model; The target format is any one of JSON format, YAML format, and XML format; The method for determining the preset content includes: obtaining the initial content; using the ID of the initial content as the first field content, and configuring the corresponding first field to represent the first field content; using the initial content as the second field content, and configuring the second field to represent the second field content; combining the first field and the first field content, the second field and the second field content, to obtain the preset content; The method for determining the preset question includes: obtaining the initial question; using the ID of the initial question as the content of the third field, and configuring the corresponding third field to represent the content of the third field; using the initial question as the content of the fourth field, and configuring the corresponding fourth field to represent the content of the fourth field; combining the third field and the content of the third field, and the fourth field and the content of the fourth field to obtain the preset question; The method for determining the preset answer includes: obtaining the initial answer; taking the ID of the initial question corresponding to the initial answer as the content of the fifth field, and configuring the corresponding fifth field to represent the content of the fifth field; taking the initial answer as the content of the sixth field, and configuring the corresponding sixth field to represent the content of the sixth field; taking the starting position of the initial answer in the initial content as the content of the seventh field, and configuring the corresponding seventh field to represent the content of the seventh field; combining the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain the preset answer.

2. The method for constructing complex instruction training data for model training according to claim 1, characterized in that: After generalizing the seed instruction, the method further includes: The generalized seed instructions are scored, and the seed instructions with score values ​​exceeding a preset threshold are selected as training data.

3. The method for constructing complex instruction training data for model training according to claim 1, characterized in that: The preset task includes a task description and a task guide. The task description includes a task example and a description of the target format. The task guide sequentially guides the preset content, preset questions, and preset answers.

4. A device for constructing complex instruction training data for model training, characterized in that: include: An acquisition unit, configured to acquire initial training data of a large language model, wherein the initial training data includes initial content, initial answers extracted from the initial content, and initial questions corresponding to the initial answers; A determination unit, configured to determine preset content, preset answers and preset questions based on the initial content, initial answers and initial questions; A combining unit, used for combining the preset task and the preset content, preset answer, and preset question to obtain a seed instruction; A generalization unit, used for generalizing the seed instruction to obtain training data; The determination unit includes: an extraction module, used to extract the initial content, initial answer and initial question from the initial training data; a format determination module, used to determine the preset content according to the initial content, determine the preset answer according to the initial answer, and determine the preset question according to the initial question in accordance with the target format; The generalization unit includes: a generalization example module, used to determine a generalization example; a sentence generalization module, used to learn the context of the generalization example based on the large language model, and generalize the sentence in the seed instruction using the large language model; The target format is any one of JSON format, YAML format, and XML format; The format determination module includes a preset content determination sub-block, which is used to: obtain the initial content; use the ID of the initial content as the first field content, and configure the corresponding first field to represent the first field content; use the initial content as the second field content, and configure the second field to represent the second field content; combine the first field and the first field content, the second field and the second field content to obtain the preset content; The format determination module further includes a preset question determination sub-block, which is used to: obtain the initial question; use the ID of the initial question as the third field content, and configure the corresponding third field to characterize the third field content; use the initial question as the fourth field content, and configure the corresponding fourth field to characterize the fourth field content; combine the third field and the third field content, and the fourth field and the fourth field content to obtain the preset question; The format determination module also includes a preset answer determination sub-block, which is used to: obtain the initial answer; use the id of the initial question corresponding to the initial answer as the content of the fifth field, and configure the corresponding fifth field to represent the content of the fifth field; use the initial answer as the content of the sixth field, and configure the corresponding sixth field to represent the content of the sixth field; use the starting position of the initial answer in the initial content as the content of the seventh field, and configure the corresponding seventh field to represent the content of the seventh field; combine the fifth field and the fifth field content, the sixth field and the sixth field content, and the seventh field and the seventh field content to obtain a preset answer.

5. The device for constructing complex instruction training data for model training according to claim 4, characterized in that: The generalization unit also includes: The scoring module is used to score the generalized seed instructions and select the seed instructions whose scoring values ​​exceed a preset threshold as training data.

6. The device for constructing complex instruction training data for model training according to claim 4, characterized in that: The preset task includes a task description and a task guide. The task description includes a task example and a description of the target format. The task guide sequentially guides the preset content, preset questions, and preset answers.

7. An electronic device, characterized in that: The electronic device comprises: a shell, a processor, a memory, a circuit board and a power supply circuit, wherein the circuit board is placed inside the space enclosed by the shell, and the processor and the memory are arranged on the circuit board; the power supply circuit is used to supply power to various circuits or devices of the above-mentioned electronic device; the memory is used to store executable program code; the processor runs the program corresponding to the executable program code by reading the executable program code stored in the memory, and is used to execute the method for constructing complex instruction training data for model training described in any one of the preceding claims 1-3.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method for constructing complex instruction training data for model training as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Man-machine conversation and pre-training language model training method and system and electronic equipment

    CN115587175A

  • Method, apparatus and device for quality control and storage medium

    US20210326524A1

Cited By

  • Large model training-oriented multi-dimensional generalizable data generation method and device and software system

    CN121189320A