Supervision fine tuning data construction method and device, electronic equipment and storage medium

By using preset construct conditions to modify the source data in the AI ​​big model, the illusion problem of AI big model generation and supervising fine-tuning data is solved, and the data generation quality and controllability are improved.

CN120087493APending Publication Date: 2025-06-03DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411966031.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, when AI models generate supervised fine-tuning data, due to hallucination problems, the generated data is low and the data quality cannot be effectively guaranteed.

Method used

By obtaining the group data composed of problem data and instruction data, the preset big model is used to modify the format of the source data according to the preset construction conditions, and the target supervision fine-tuning data that meets the construction conditions is generated.

Benefits of technology

The source data format is rewritten according to preset construction conditions through the large model, which avoids the illusion of directly generating data from the large model, and improves the generation quality and controllability of supervised fine-tuning data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087493A_ABST
    Figure CN120087493A_ABST
Patent Text Reader

Abstract

The invention provides a supervision fine tuning data construction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining source data which is composed of problem data and instruction data; a problem, instruction gt; group data, the source data being derived from a plurality of data sets; inputting the source data and a preset construction condition into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction condition to obtain a target lt meeting the preset construction condition; a problem, instruction gt; and the group serves as target supervision fine tuning data. The source data format is rewritten by using the large model according to the preset construction condition to obtain the target supervision fine tuning data, and the data is not directly generated by using the large model, so that wrong data generation possibly caused by the illusion problem of the large model is avoided, and the generation quality of the supervision fine tuning data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model technology, and in particular to a supervised fine-tuning data construction method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of AI technology, AI big models have been widely used. In order to enable AI big models to deal with a specific problem, such as generating images based on text, generating answers based on questions, etc., it is usually necessary to supervise and fine-tune the pre-trained AI big model, that is, to train the model with labeled data and adjust the parameters in the big model by minimizing the difference between the model prediction results and the data labels. The above-mentioned pre-trained big model usually refers to a big model trained with unsupervised text data, where unsupervised means that it is not manually labeled and is directly trained with the supervisory signals in the text, that is, the order of words in the sentence.

[0003] In order to ensure the fine-tuning effect, it is usually necessary to use a large amount of supervised fine-tuning data to supervise the fine-tuning of the large model. In related technologies, supervised fine-tuning data is usually generated by related large models. However, large models may have hallucination problems, that is, the training data used in the training process of the large model may have related documents appearing too many times or being close in position, etc., which may cause the large model to output wrong data, resulting in the generated supervised fine-tuning data having low accuracy. Summary of the invention

[0004] In view of this, embodiments of the present invention provide a supervised fine-tuning data construction method, device, electronic device and storage medium to improve the generation quality of supervised fine-tuning data.

[0005] According to one aspect of the present invention, a supervised fine-tuning data construction method is provided, the method comprising:

[0006] Acquire source data, where the source data is a <question, instruction> group of data consisting of question data and instruction data, and the source data comes from multiple data sets;

[0007] The source data and preset construction conditions are input into the preset large model, so that the preset large model modifies the format of the source data according to the preset construction conditions, and obtains a target <question, instruction> group that meets the preset construction conditions as target supervision fine-tuning data.

[0008] In a possible embodiment, the question data includes a question and data storage information; the instruction data includes a query instruction obtained based on the data storage information and the question;

[0009] The preset construction conditions include: target problem data, target query instruction format, annotations, and content that cannot be included.

[0010] In one possible embodiment, the method further includes:

[0011] In the case where the format of the target supervised fine-tuning data output by the preset large model does not conform to the preset construction conditions, input the source data and the supervised fine-tuning data that conforms to the preset construction conditions into the preset large model, so that the preset large model learns the supervised fine-tuning data that conforms to the preset construction conditions, and outputs target supervised fine-tuning data for the source data according to the preset construction conditions.

[0012] In one possible embodiment, the method further includes:

[0013] Determine the type of the source data, where the type of the source data includes code data, image data, and text data;

[0014] Determine the preset construction conditions corresponding to the type of the source data.

[0015] According to another aspect of the present invention, there is provided a device for constructing supervised fine-tuning data, characterized in that the device includes:

[0016] An acquisition module, configured to acquire source data, where the source data is <problem, instruction> group data composed of problem data and instruction data, and the source data is from multiple data sets;

[0017] A construction module, configured to input the source data and the preset construction conditions into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction conditions to obtain a target <problem, instruction> group that conforms to the preset construction conditions as target supervised fine-tuning data.

[0018] In one possible embodiment, the problem data includes a problem and data storage information; the instruction data includes a query instruction obtained based on the data storage information and the problem;

[0019] The preset construction conditions include: target problem data, target query instruction format, annotations, and content that cannot be included.

[0020] In one possible embodiment, the construction module is configured to, in the case where the format of the target supervised fine-tuning data output by the preset large model does not conform to the preset construction conditions, input the source data and the supervised fine-tuning data that conforms to the preset construction conditions into the preset large model, so that the preset large model learns the supervised fine-tuning data that conforms to the preset construction conditions, and outputs target supervised fine-tuning data for the source data according to the preset construction conditions.

[0021] In one possible embodiment, the device further includes:

[0022] A determination module, configured to determine the source data type, where the source data type includes code data, image data, and text data; and determine the preset construction conditions corresponding to the source data type.

[0023] According to another aspect of the present invention, there is provided an electronic device, including:

[0024] A processor; and

[0025] A memory storing a program,

[0026] wherein the program includes instructions that, when executed by the processor, cause the processor to execute any one of the above-mentioned supervised fine-tuning data construction methods.

[0027] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the above-mentioned supervised fine-tuning data construction methods.

[0028] In one or more technical solutions provided in the embodiments of the present invention, source data is obtained, where the source data is a <question, instruction> group data composed of question data and instruction data, and the source data is derived from multiple data sets; the source data and the preset construction conditions are input into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction conditions to obtain a target <question, instruction> group that conforms to the preset construction conditions, as target supervised fine-tuning data. By using the large model to rewrite the source data format according to the preset construction conditions to obtain the target supervised fine-tuning data, rather than directly generating data using the large model, the problem of generating incorrect data caused by the large model hallucination problem is avoided, and the generation quality of the supervised fine-tuning data is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In the following description of the exemplary embodiments in conjunction with the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:

[0030] Figure 1 is a schematic flowchart of a supervised fine-tuning data construction method provided by an embodiment of the present invention;

[0031] Figure 2 is a schematic structural diagram of a supervised fine-tuning data construction device provided by an embodiment of the present invention

[0032] Figure 3 shows a structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed Implementation Modes

[0033] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0034] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0035] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.

[0036] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0037] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0038] Currently, the methods for constructing supervised fine-tuning (SFT) data can be divided into two types. One is manual annotation, and the other is to generate with the help of a more powerful large model. Since the cost of manual annotation is huge, the latter method has become a general method for obtaining large-scale instruction data. In related technologies, the supervised fine-tuning data is usually constructed by the following methods:

[0039] 1. Self-instruct: This method is not restricted by domain and can be used in general domains and other domains (such as code). The specific method is to manually write some seed data, including instructions and answers, and let the large model continue writing by referring to the seed data to generate new data. The data constructed by this method has a relatively low difficulty, and the instructions are basically one sentence.

[0040] 2. Evol-instruct: This method is also not restricted by domain and can be used in general domains and other domains (such as code). The specific method is to transform the original instructions with the help of the large model on the basis of Self-instruct data to increase the instruction difficulty (such as increasing the time complexity requirement) or make the original instructions more specific (such as specifying the input data format), etc. Then call the large model again to use the large model's own ability to regenerate answers for the transformed instructions. This method can effectively improve the difficulty and detail of Self-instruct data. However, since the Self-instruct data itself is generated by the large model and the answers are also generated by the large model, the data quality still cannot be guaranteed.

[0041] 3. Oss-instruct: This method is proposed in the code domain and can also be extended to other domains. The main idea is to randomly extract code snippets from github code files and let the large model generate instructions and answers inspired by these code snippets. This method provides reference information for the large model, and the reference information comes from real user code on github. Therefore, the generated instructions are more in line with the user questions in real scenarios, and the coverage of code languages, task scenarios, and instruction difficulties is also wider. This method has greatly improved the authenticity and specificity of the instructions compared with the previous two methods. However, generating instruction and answer data inspired by a certain piece of information ultimately still depends on the large model's ability, and the generated data quality is still uncontrollable.

[0042] Specifically, the Self-instruct method directly generates new data by the large model referring to the manually written seed data. The data constructed is often too simple, and the instructions are usually in the form of one sentence. The Evol-instruct method transforms on the basis of Self-instruct data to make the instructions more complex, but the data still comes entirely from the large model generation, and the quality still cannot be guaranteed. The Oss-instruct method constructs data by referring to the code snippets extracted from github, and the generated instructions are more real and specific. However, just providing inspiration to the large model still has a great dependence on the large model's own creative ability and cannot ensure the controllability of the generated data quality.

[0043] It can be seen that the supervised fine-tuning instruction data constructed by the prior art relies too much on large models. Due to the hallucination problem existing in large models, it is very difficult to guarantee the quality of the generated data, and there may be problems such as meaningless instructions, answers that cannot answer the instructions, overly simple instructions, and insufficient instruction diversity.

[0044] Based on this, the embodiments of the present invention provide a method, device, electronic device, and storage medium for constructing supervised fine-tuning data. The method for constructing supervised fine-tuning data provided by the embodiments of the present invention can be applied to any electronic device with a data construction function, and the electronic device can be a server, a computer, a mobile terminal, etc. The solution of the present invention will be described below with reference to the accompanying drawings:

[0045] Figure 1 FIG. is a schematic flowchart of a method for constructing supervised fine-tuning data provided by an embodiment of the present invention, which may include the following steps:

[0046] S101. Obtain source data, where the source data is <question, instruction> group data composed of question data and instruction data, and the source data is derived from multiple data sets;

[0047] S102. Input the source data and a preset construction condition into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction condition to obtain a target <question, instruction> group that meets the preset construction condition, as target supervised fine-tuning data.

[0048] In the embodiments of the present invention, source data is obtained, where the source data is <question, instruction> group data composed of question data and instruction data, and the source data is derived from multiple data sets; the source data and a preset construction condition are input into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction condition to obtain a target <question, instruction> group that meets the preset construction condition, as target supervised fine-tuning data. By using the large model to rewrite the format of the source data according to the preset construction condition to obtain the target supervised fine-tuning data, rather than directly generating data using the large model, the problem of generating incorrect data caused by the hallucination problem of the large model is avoided, and the generation quality of the supervised fine-tuning data is improved.

[0049] The above S101-S102 will be exemplarily described below:

[0050] The above source data can be any type of data, such as code data, image data, or text data, etc. The above source data can come from different data sets. In a possible embodiment, if the above source data is code data, the above data sets can include Code_contests, Bird, etc. Among them, the Code_contests data set comes from a programming website, with the input being code programming questions and the output being the code solutions corresponding to the questions; the Bird data set is an artificially constructed nl2sql (Natural Language to SQL) data set, with the input being database information and user questions and the output being the corresponding sql statements, which can be executed in the current database to query the information asked by the user.

[0051] If the above source data is image data, its source data sets can include the Multi-Turn Multi-Image Dialog Understanding dataset MMDU, the MS COCO (Microsoft Common Objects in Context, a large-scale computer vision dataset) dataset, etc. The above source data sets and source data can be flexibly selected according to the actual business scenario.

[0052] The above source data is a <question, instruction> group data composed of question data and instruction data. Among them, the question data can include a question and data storage information. The above data storage information refers to the storage database information of the data, which can include the database name, address, and the identifier of the data, etc. The above question is a question that can obtain an answer based on this data storage information, such as how to find a certain data, where is the storage location of a certain data, etc. The above instruction data can be the SQL form expression of the above question.

[0053] Exemplarily, the above question data can be: Given the database schema information and the query question, please help me complete the sql statement.

[0054] Database table information:

[0055]

[0056]

[0057] Query question:

[0058] How many players who were drafted by the Toronto Maple Leafs have played over 300 games in their first 7 years of the NHL career?

[0059] The corresponding instruction data may be SELECT COUNT(ELITEID) FROM PlayerInfo WHERE overallby='Toronto Maple Leafs' AND sum_7yr_GP>300.

[0060] In a possible embodiment, if the source data is image data, the question data may include image description information, and the instruction data may be image data. If the source data is text data, the question data may include user questions that may appear in actual business scenarios and reference materials for the questions, and the instructions may include answers to the questions generated based on the materials.

[0061] In an embodiment of the present invention, the above-mentioned <question, instruction> group and preset construction conditions can be input into a preset large model. The preset construction conditions can be pre-set according to different types of source data. For example, different construction conditions can be set for the above-mentioned code data, image data, and text data. The construction conditions can be set according to actual business scenarios, and the present invention does not make specific limitations on this.

[0062] In a possible embodiment, the preset construction conditions may include the target question data and the target query instruction format, annotations, and non-included content, wherein the target question data and the target query instruction format may be standardized code indentation and Markdown format, Markdown is a lightweight markup language that allows users to write documents in an easy-to-read and easy-to-write plain text format, and then convert them into valid XHTML (or HTML) documents. The non-included content refers to content that is not desired to appear in the target question and the target instruction, and may include logos, text, etc.

[0063] Exemplarily, based on the above example, the preset construction condition may be:

[0064] 1. Make sure that the [question] contains instructions, a detailed description of the database structure, and the query. The [instructions] contain the SQL statement corresponding to the query and a detailed text description;

[0065] 2. Rewrite the title in natural and fluent Chinese while keeping the meaning unchanged, but making it more in line with the style of how humans ask questions to ChatGPT;

[0066] 3. Describe the database structure in another way, but make sure to include the necessary information for writing SQL statements, such as **English database table names and column names**, so that correct SQL statements can be written and executed. Other information such as types and descriptions can be randomly retained;

[0067] 4. Describe the database structure as diversely as possible, such as using natural language description or the way of "table name (column name 1, column name 2)", covering as many possible description methods by humans as possible;

[0068] 5. Ensure that the description of the database structure is given in the

Question

[0069] 6. Give a more comprehensive, fluent and logically clear answer, adding an analysis of the question, necessary code comments and detailed Chinese explanations;

[0070] 7. The answer needs to accurately answer the question. If the answer is not the correct answer or the best answer, directly give a better answer;

[0071] 8. The words irrelevant to the question and answer content such as "model identifier, original question, original answer, modified question, detailed Chinese explanation, analysis" shall not appear in the output

Question

Instruction

[0072] The above preset construction conditions can be pre-stored in a preset large model, which can be any type of large model with question-and-answer function, such as ChatGPT. In a possible embodiment, since there are multiple types of source data, the above method may further include determining the type of the source data, where the type of the source data includes code data, image data, and text data; determining the preset construction conditions corresponding to the type of the source data.

[0073] As a possible implementation manner, the above type of source data can be determined based on the data type input by the user, or can be determined according to the question data and instruction data in the source data. The present invention does not make specific limitations on this.

[0074] The preset large model can generate a target <question, instruction> group according to the question data and instruction data according to the preset construction conditions, and the target <question, instruction> group includes target question data and target instruction data.

[0075] Exemplarily, based on the above example, the target <question, instruction> group that can be obtained through the preset large model is as follows:

[0076]

Question

[0077] Please write an SQL query statement based on the following database table structure. We have a database about hockey players, which contains the basic information of the players and their height and weight data. Specifically, there is a table called `PlayerInfo`, which records information such as the name, date of birth, nationality, and the number of games played in the first seven years in the NHL (National Hockey League) of the players. Among them, the `overallby` column represents the team that selected them during the draft, and the `sum_7yr_GP` column represents the total number of games played by the players in the first seven years in the NHL. I want to know how many players selected by the Toronto Maple Leafs have played more than 300 games in the first seven years of their NHL careers?

[0078]

Instruction

[0079] ```sql

[0080] SELECT COUNT(*) AS NumberOfPlayers

[0081] FROM PlayerInfo

[0082] WHERE overallby = 'Toronto Maple Leafs' AND sum_7yr_GP > 300;

[0083] The purpose of this SQL query statement is to count the number of players selected by the Toronto Maple Leafs who have played more than 300 games in the first seven years of their NHL careers.

[0084] The `COUNT(*)` function is used to calculate the number of records that meet the conditions, and the `WHERE` clause is used to filter out the players who were selected by the Toronto Maple Leafs (`overallby = 'Toronto Maple Leafs'`) and have played more than 300 games in the first seven years in the NHL (`sum_7yr_GP > 300`). The result will be displayed with the column name `NumberOfPlayers` showing the number of these players.

[0085] It can be seen that the quality of the transformed data is comparable to that of the artificially constructed data. The questions are natural and fluent, fully conforming to the style of human questions. The instructions have been further optimized based on the standard answers and include regular markdown code format and detailed text explanations.

[0086] In a possible embodiment, the embodiment of the present invention may further include cleaning the target supervised fine-tuning data, for example, cleaning the target supervised fine-tuning data based on the above-mentioned preset construction conditions. Specifically, each construction condition can be matched with the target supervised fine-tuning data. If the match is successful, it means that the target supervised fine-tuning data meets the requirements and can be retained.

[0087] In a possible embodiment, if the preset large model does not output a target <question, instruction> group that meets the preset construction conditions, the source data and the supervised fine-tuning data that meet the preset construction conditions can be input into the preset large model, so that the preset large model learns the supervised fine-tuning data that meets the preset construction conditions and outputs target supervised fine-tuning data for the source data according to the preset construction conditions.

[0088] Applying the embodiment of the present invention, by providing detailed input-output reference information to the preset large model, and only using the large model to rewrite the input and polish the output, high-quality supervised fine-tuning data can be obtained. There is no need to use the large model to generate instructions out of thin air, nor does the large model need to rely on its own knowledge to generate correct answers, which greatly reduces the dependence on the large model's capabilities when constructing data and improves the controllability of the generated data quality. Thus, the authenticity, effectiveness, diversity, and balance of difficulty of the supervised fine-tuning data constructed by the large model are improved by using the open-source code benchmark dataset, effectively improving the quality of the code supervised fine-tuning dataset, and thus improving the performance of the code supervised fine-tuning model.

[0089] Based on the same inventive concept, the embodiment of the present invention also provides a device for constructing supervised fine-tuning data, such as Figure 2 shown. The device 200 may include:

[0090] An acquisition module 201, configured to acquire source data, where the source data is <question, instruction> group data composed of question data and instruction data, and the source data is from multiple datasets;

[0091] A construction module 202, configured to input the source data and preset construction conditions into a preset large model, so that the preset large model modifies the format of the source data according to the preset construction conditions to obtain a target <question, instruction> group that meets the preset construction conditions as target supervised fine-tuning data.

[0092] In a possible embodiment, the question data includes a question and data storage information; the instruction data includes a query instruction obtained based on the data storage information and the question;

[0093] The preset construction conditions include: target question data and target query instruction format, annotations, and content that cannot be included.

[0094] In a possible embodiment, the construction module is configured to input the source data and the supervised fine-tuning data that meets the preset construction conditions into the preset large model when the format of the target supervised fine-tuning data output by the preset large model does not meet the preset construction conditions, so that the preset large model learns the supervised fine-tuning data that meets the preset construction conditions and outputs the target supervised fine-tuning data for the source data according to the preset construction conditions.

[0095] In a possible embodiment, the apparatus further includes:

[0096] A determination module, configured to determine the type of the source data, where the type of the source data includes code data, image data, and text data; and determine the preset construction conditions corresponding to the type of the source data.

[0097] Wherein, in the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0098] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.

[0099] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.

[0100] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.

[0101] Reference Figure 3, a block diagram of an electronic device 300 that can be a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0102] As Figure 3 shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0103] Multiple components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, an output unit 307, a storage unit 308, and a communication unit 309. The input unit 306 can be any type of device that can input information into the electronic device 300. The input unit 306 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 307 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 308 can include, but is not limited to, magnetic disks, optical disks. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0104] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above. For example, in some embodiments, the above-described supervised fine-tuning data construction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. In some embodiments, the computing unit 301 can be configured to execute the above-described supervised fine-tuning data construction method in any other suitable way (e.g., by means of firmware).

[0105] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0106] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0107] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0108] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0109] The systems and techniques described here can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described here), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0110] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A supervised fine-tuning data construction method, characterized in that: The method comprises: Acquire source data, where the source data is a <question, instruction> group of data consisting of question data and instruction data, and the source data comes from multiple data sets; The source data and preset construction conditions are input into the preset large model, so that the preset large model modifies the format of the source data according to the preset construction conditions, and obtains a target <question, instruction> group that meets the preset construction conditions as target supervision fine-tuning data.

2. The method according to claim 1, characterized in that The question data includes the question and data storage information; the instruction data includes a query instruction obtained based on the data storage information and the question; The preset construction conditions include: target question data and target query instruction format, annotations and non-included content.

3. The method according to claim 1, characterized in that: The method further comprises: When the format of the target supervised fine-tuning data output by the preset large model does not meet the preset construction conditions, the source data and the supervised fine-tuning data that meets the preset construction conditions are input into the preset large model so that the preset large model learns the supervised fine-tuning data that meets the preset construction conditions and outputs the target supervised fine-tuning data for the source data according to the preset construction conditions.

4. The method according to claim 1, characterized in that The method further comprises: Determine the source data type, the source data type includes code data, image data and text data; Determine a preset construction condition corresponding to the source data type.

5. A supervised fine-tuning data construction device, characterized in that: The device comprises: An acquisition module, used for acquiring source data, wherein the source data is a <question, instruction> group data consisting of question data and instruction data, and the source data comes from multiple data sets; A construction module is used to input the source data and preset construction conditions into a preset large model so that the preset large model modifies the format of the source data according to the preset construction conditions, and obtains a target <problem, instruction> group that meets the preset construction conditions as target supervision fine-tuning data.

6. The device according to claim 5, characterized in that The question data includes the question and data storage information; the instruction data includes a query instruction obtained based on the data storage information and the question; The preset construction conditions include: target question data and target query instruction format, annotations and non-included content.

7. The device according to claim 5, characterized in that The construction module is used to input the source data and the supervised fine-tuning data that meets the preset construction conditions into the preset large model when the format of the target supervised fine-tuning data output by the preset large model does not meet the preset construction conditions, so that the preset large model learns the supervised fine-tuning data that meets the preset construction conditions and outputs the target supervised fine-tuning data for the source data according to the preset construction conditions.

8. The device according to claim 5, characterized in that The device also includes: The determination module is used to determine the source data type, which includes code data, image data and text data; and determine the preset construction conditions corresponding to the source data type.

9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-4.