Sample Data Generation Method, Apparatus, Electronic Device, and Storage Medium
Through the large language model in multiple rounds of dialogue, initial problem instructions, target answer text and target problem instructions are generated, which solves the problem of insufficient quality and diversity of sample data in the prior art, and achieves efficient and diversified sample data generation.
Patent Information
- Application Number
- CN202311810096.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-12-26
AI Technical Summary
The prior art cannot efficiently obtain high-quality and diverse sample data for training large language models.
By calling multiple large language models for multiple rounds of conversations, initial question instructions, target answer text and target question instructions are generated, and sample data is constructed.
Efficiently obtain high-quality and diverse sample data, improving the complexity and diversity of problem instructions.
Smart Images

Figure CN118861218B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for generating sample data. Background Art
[0002] In the field of artificial intelligence, a large language model is short for a large language model (LLM), which refers to a deep learning model trained with a large amount of text data and can generate natural language text or understand the meaning of language text. A large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc.
[0003] Currently, sample data for training large language models can be obtained through various acquisition methods. For example, high-quality sample data can be obtained through manual construction, but the manual construction method cannot efficiently obtain diverse sample data. Another example is that a large amount of sample data can be collected through online platforms, but the quality of the sample data is generally low and it takes a lot of time for data cleaning. It can be seen that traditional sample data acquisition methods cannot efficiently obtain high-quality and diverse sample data, and there is an urgent need for a method that can efficiently obtain high-quality and diverse sample data. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this application. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of the present application provide a method, device, electronic device and storage medium for generating sample data, which can efficiently obtain high-quality and diverse sample data.
[0006] On the one hand, an embodiment of the present application provides a method for generating sample data, including:
[0007] Obtaining a first prompt text for prompting a first large language model to generate a question instruction, and calling the first large language model to perform question instruction prediction based on the first prompt text to generate an initial question instruction;
[0008] Using the initial question instruction as the first input to the question instruction of a second large language model, and calling the second large language model and a third large language model to have multiple rounds of conversations, where the second large language model is used to generate a target answer text according to the input question instruction, and the third large language model is used to generate a target question instruction input to the second large language model according to the target answer text;
[0009] Constructing sample data based on the initial question instruction, the target question instruction in multiple rounds of conversations, and the target answer text in multiple rounds of conversations.
[0010] On the other hand, an embodiment of the present application further provides a sample data generation device, including:
[0011] A first generation module, configured to obtain a first prompt text for prompting a first large language model to generate a question instruction, and call the first large language model to perform question instruction prediction based on the first prompt text to generate an initial question instruction;
[0012] A second generation module, configured to use the initial question instruction as the first input to the question instruction of a second large language model, and call the second large language model to have multiple rounds of conversations with a third large language model, where the second large language model is used to generate a target answer text according to the input question instruction, and the third large language model is used to generate a target question instruction input to the second large language model according to the target answer text;
[0013] A sample construction module, configured to construct sample data based on the initial question instruction, the target question instruction in the multiple rounds of conversations, and the target answer text in the multiple rounds of conversations.
[0014] Further, the number of knowledge granularity types of the candidate text is at least one, and the second generation module is further configured to:
[0015] Obtain a first role definition text and a second role definition text, where the first role definition text is used to prompt the second large language model to be the answering party for the target task in the multiple rounds of conversations, and the second role definition text is used to prompt the third large language model to be the questioning party in the multiple rounds of conversations, and the target task is a downstream task trained based on the sample data;
[0016] Input the first role definition text into the second large language model, and input the second role definition text into the third large language model.
[0017] Further, the second generation module is specifically configured to:
[0018] Input the initial question instruction into the second large language model for answer prediction to generate a target answer text for the first conversation round;
[0019] Construct a second prompt text for prompting the third large language model to generate a question instruction according to the target answer text whenever the target answer text is received;
[0020] Input the second prompt text and the target answer text of the first conversation round into the third large language model for question instruction prediction to generate a target question instruction for the next conversation round;
[0021] Input the target question instruction of the next conversation turn into the second large language model for answer prediction again until the number of conversation turns is equal to the preset number threshold.
[0022] Further, the above-mentioned second generation module is specifically used for:
[0023] Add the initial question instruction to the second prompt text to obtain the target prompt text;
[0024] Input the target prompt text and the target answer text of the first conversation turn into the third large language model for question instruction prediction, and generate the target question instruction of the next conversation turn, where the target prompt text is used to prompt the third large language model to generate question instructions according to each question instruction input to the second large language model and the target answer text of each conversation turn.
[0025] Further, the above-mentioned first generation module is specifically used for:
[0026] Obtain the task information of the target task, where the target task is a downstream task trained based on the sample data;
[0027] Based on the task information, construct a first prompt text for prompting the first large language model to generate question instructions according to the task information.
[0028] Further, the above-mentioned first generation module is specifically used for:
[0029] According to the task information, determine a target node in the preset task tree that matches the task information, where the task tree includes multiple layers of nodes, and each node carries its own candidate keywords;
[0030] Use the candidate keywords carried by the target node as the target keywords, and according to the target keywords, construct a first prompt text for prompting the first large language model to generate question instructions according to the task information.
[0031] Further, the above-mentioned first generation module is specifically used for:
[0032] When the level where the target node is located is the target level, construct a first prompt text for prompting the first large language model to generate question instructions according to the task information, where the target level is the level after the level where the root node of the task tree is located;
[0033] Alternatively, when the level where the target node is located is a level after the target level, determine an associated node associated with the target node in the task tree, use the candidate keyword carried by the associated node as an associated keyword, and construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information based on the associated keyword and the target keyword, where the level where the associated node is located is before the level where the target node is located.
[0034] Further, the level after the level where the root node in the task tree is located is the target level, and the nodes located at the target level carry third role definition texts. The first generation module further is configured to:
[0035] Use the node associated with the target node and located at the target level as the target role definition node, and use the third role definition text carried by the target role definition node as the target role definition text, where the target role definition text is used to prompt the first large language model to be the questioner for the target task;
[0036] Input the target role definition text into the first large language model.
[0037] Further, the number of the initial question instructions is multiple, and the sample data generation device further includes a third generation module. The third generation module is specifically configured to:
[0038] Randomly sample a first question instruction from the multiple initial question instructions;
[0039] Construct a third prompt text for prompting the first large language model to perform instruction expansion according to a preset expansion strategy based on the preset expansion strategy;
[0040] Input the third prompt text and the first question instruction into the first large language model, perform instruction expansion on the first question instruction, generate a first expanded question instruction, and use the first expanded question instruction as the initial question instruction.
[0041] Further, the number of the expansion strategies is multiple, and the third generation module is specifically configured to:
[0042] Perform instruction expansion on the first question instruction multiple times to obtain a first expanded question instruction generated by each instruction expansion;
[0043] Wherein, whenever performing instruction expansion on the first question instruction, perform instruction expansion on the first question instruction according to at least one of the multiple expansion strategies.
[0044] Further, the number of the initial question instructions is multiple, and the first generation module further is configured to:
[0045] Randomly sample a second problem instruction from multiple said initial problem instructions;
[0046] Construct a fourth prompt text for prompting the first large language model to generate a problem instruction with reference to the second problem instruction, input the fourth prompt text and the second problem instruction into the first large language model for problem instruction prediction, generate a second extended problem instruction, and use the second extended problem instruction as the initial problem instruction.
[0047] Furthermore, the above first generation module is further configured to:
[0048] Obtain a target sampling temperature value;
[0049] Configure the temperature parameter of the first large language model as the target sampling temperature value, where the target sampling temperature value is used to indicate the randomness of the output result of the first large language model.
[0050] Furthermore, the above first generation module specifically is configured to:
[0051] Call the first large language model to perform problem instruction prediction based on the first prompt text, and sequentially generate multiple initial problem instructions;
[0052] Wherein, whenever an initial problem instruction is generated, determine the instruction similarity between the currently generated initial problem instruction and the historically generated initial problem instruction, and when the instruction similarity is greater than or equal to a preset similarity threshold, eliminate the currently generated initial problem instruction.
[0053] On the other hand, an embodiment of the present application further provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above sample data generation method is implemented.
[0054] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the above sample data generation method is implemented.
[0055] On the other hand, an embodiment of the present application further provides a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes to implement the above sample data generation method.
[0056] The embodiments of the present application at least include the following beneficial effects: By generating an initial question instruction through a first large language model, and then using the initial question instruction as the first input to the question instruction of the second large language model, which is equivalent to using the initial question instruction as the starting point of a multi-round conversation. Through multi-round conversations between the second large language model and the third large language model, target answer texts and target question instructions for multiple conversation rounds can be obtained. Then, sample data is constructed through the initial question instruction, the target question instruction, and the target answer text. Since the first large language model can generate high-quality and diverse initial question instructions, the second large language model can generate high-quality and diverse target answer texts, and the third large language model can generate high-quality and diverse target question instructions, it is possible to improve the complexity and diversity of the question instructions by combining multiple large language models, thereby efficiently obtaining high-quality and diverse sample data.
[0057] Other features and advantages of the present application will be described in the subsequent description, and, in part, will become apparent from the description or will be understood by implementing the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the description. They are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0059] Figure 1 It is a schematic diagram of an optional implementation environment provided for the embodiments of the present application;
[0060] Figure 2 It is an optional flowchart of the sample data generation method provided for the embodiments of the present application;
[0061] Figure 3 It is a first optional structural diagram of the first large language model provided for the embodiments of the present application;
[0062] Figure 4 It is an optional structural diagram of the multi-round conversation provided for the embodiments of the present application;
[0063] Figure 5 It is an optional structural diagram of the task tree provided for the embodiments of the present application;
[0064] Figure 6 It is an optional interface diagram of the task configuration interface provided for the embodiments of the present application;
[0065] Figure 7 It is another optional interface diagram of the task configuration interface provided for the embodiments of the present application;
[0066] Figure 8The second optional structural schematic diagram of the first large language model provided by the embodiments of the present application;
[0067] Figure 9 The third optional structural schematic diagram of the first large language model provided by the embodiments of the present application;
[0068] Figure 10 An optional architecture schematic diagram of the sample data generation method provided by the embodiments of the present application;
[0069] Figure 11 An optional structural schematic diagram of the sample data generation device provided by the embodiments of the present application;
[0070] Figure 12 A partial structural block diagram of the terminal provided by the embodiments of the present application;
[0071] Figure 13 A partial structural block diagram of the server provided by the embodiments of the present application. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0073] It should be noted that in each specific implementation manner of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as the attribute information or set of attribute information of the target object, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. Among them, the target object can be a user. In addition, when the embodiments of the present application need to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present application will be obtained.
[0074] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0075] To facilitate the understanding of the technical solutions provided by the embodiments of this application, some key terms used in the embodiments of this application are explained here first:
[0076] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system background, which can only be achieved through cloud computing.
[0077] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0078] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0079] Currently, sample data for training large language models can be obtained through various acquisition methods. For example, high-quality sample data can be obtained through manual construction, but the manual construction method cannot efficiently obtain diverse sample data. Another example is that a large amount of sample data can be collected through online platforms, but the quality of the sample data is generally low and it takes a lot of time for data cleaning. It can be seen that traditional sample data acquisition methods cannot efficiently obtain high-quality and diverse sample data, and there is an urgent need for a method that can efficiently obtain high-quality and diverse sample data.
[0080] Based on this, the embodiments of the present application provide a method, device, electronic device, and storage medium for generating sample data, which can efficiently obtain high-quality and diverse sample data.
[0081] Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of an optional implementation environment provided by the embodiments of the present application. This implementation environment includes a terminal 101 and a server 102, where the terminal 101 and the server 102 are connected through a communication network.
[0082] Exemplarily, the server 102 can obtain a first prompt text sent by the terminal 101 for prompting the first large language model to generate a problem instruction, call the first large language model to predict a problem instruction based on the first prompt text, and generate an initial problem instruction; use the initial problem instruction as the first input problem instruction to the second large language model, and call the second large language model and the third large language model to conduct multiple rounds of conversations. Among them, the second large language model is used to generate a target answer text according to the input problem instruction, and the third large language model is used to generate a target problem instruction input to the second large language model according to the target answer text; based on the initial problem instruction, the target problem instructions in multiple rounds of conversations, and the target answer texts in multiple rounds of conversations, construct sample data and send it to the terminal 101.
[0083] Server 102 generates an initial question instruction through the first large language model, and then uses the initial question instruction as the first input to the question instruction of the second large language model. That is, the initial question instruction is used as the starting point of the multi-round conversation. Through multi-round conversations between the second large language model and the third large language model, target response texts and target question instructions for multiple conversation rounds can be obtained. Then, sample data is constructed through the initial question instruction, target question instruction, and target response text. Since the first large language model can generate high-quality and diverse initial question instructions, the second large language model can generate high-quality and diverse target response texts, and the third large language model can generate high-quality and diverse target question instructions, it is possible to improve the complexity and diversity of question instructions by combining multiple large language models, thereby efficiently obtaining high-quality and diverse sample data.
[0084] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, Server 102 can also be a node server in a blockchain network.
[0085] Terminal 101 can be a mobile phone, computer, intelligent voice interaction device, smart home appliance, vehicle terminal, etc., but is not limited thereto. Terminal 101 and Server 102 can be directly or indirectly connected through wired or wireless communication methods, and this application embodiment does not make any restrictions here.
[0086] The method provided in the embodiments of this application can be applied to various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0087] Refer to Figure 2 , Figure 2 FIG. is an optional flowchart of the sample data generation method provided in the embodiments of this application. This sample data generation method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. This text type determination method includes but is not limited to the following steps 201 to 203.
[0088] Step 201: Obtain a first prompt text for prompting the first large language model to generate a question instruction, and call the first large language model to perform question instruction prediction based on the first prompt text to generate an initial question instruction.
[0089] Among them, the first prompt text is used to prompt the first large language model to generate a question instruction. That is to say, the first prompt text can guide the first large language model to generate specific outputs. By providing the first prompt text, the text content generated by the first large language model can be affected.
[0090] Among them, the first large language model is a pre-trained large language model. A large language model is a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Generally, a large language model adopts a recurrent neural network (RNN) or its variants, such as long short-term memory network (LSTM) and gated recurrent unit (GRU), to capture the context information in the text sequence, so as to realize tasks such as natural language text generation, language model evaluation, text classification, and sentiment analysis. In the field of natural language processing, large language models have been widely used, such as speech recognition, machine translation, automatic summarization, dialogue systems, intelligent question answering, etc.
[0091] Specifically, refer to Figure 3 , Figure 3 which is the first optional structural schematic diagram of the first large language model provided by the embodiment of the present application.
[0092] Among them, inputting the first prompt text into the first large language model to generate an initial question instruction through the first large language model can improve the accuracy of the initial question instruction, have a high generation quality, and the initial question instruction generated by the first large language model has randomness, which is equivalent to the first large language model being able to generate high-quality and diverse target question instructions.
[0093] Step 202: Use the initial question instruction as the first input to the question instruction of the second large language model, and call the second large language model and the third large language model to conduct multiple rounds of conversations;
[0094] Among them, the second large language model is used to generate a target answer text according to the input question instruction, and the third large language model is used to generate a target question instruction input to the second large language model according to the target answer text. The second large language model is a pre-trained large language model. Therefore, by generating a target answer text through the second large language model, the accuracy of the target answer text can be improved, and it has a high generation quality. The third large language model is also a pre-trained large language model. Therefore, by generating a target question instruction through the third large language model, the accuracy of the target question instruction can be improved, and it also has a high generation quality.
[0095] Based on this, taking the initial question instruction as the first input to the question instruction of the second large language model is equivalent to taking the initial question instruction as the starting point of a multi-round conversation. The second large language model generates the target answer text in the answering stage, and the target answer text generated by the second large language model is random. The third large language model generates the target question instruction in the questioning stage, and the target question instruction generated by the third large language model is random. Therefore, under the interaction of the second large language model and the third large language model, high-quality and diverse target answer texts and target question instructions can be generated.
[0096] Specifically, one of the second large language model and the third large language model can be the same large language model as the first large language model.
[0097] In a possible implementation manner, before taking the initial question instruction as the first input to the question instruction of the second large language model, the sample data generation method further includes: obtaining a first role definition text and a second role definition text, where the first role definition text is used to prompt the second large language model to act as the answering party for the target task in a multi-round conversation, and the second role definition text is used to prompt the third large language model to act as the questioning party in a multi-round conversation. The target task is a downstream task based on the sample data for training; inputting the first role definition text into the second large language model and inputting the second role definition text into the third large language model.
[0098] Among them, the target task is a downstream task in a specific application field. For example, the target task can be the generation of JavaScript code in front-end development, and the target task can be the generation of a travel planning plan; the content of the first role definition text is used to display and define the role of the second large language model, and the content of the second role definition text is used to display and define the role of the third large language model.
[0099] Based on this, before taking the initial question instruction as the first input to the question instruction of the second large language model, inputting the first role definition text into the second large language model to prompt the second large language model to act as the answering party for the target task in a multi-round conversation. The second large language model can better understand the answer generation task in the multi-round conversation, can improve the generation quality of the target answer text, and the second large language model can generate the target answer text related to the target task. Subsequently, sample data related to the target task can be obtained, and the large language model is trained using the sample data related to the target task to fine-tune the large language model. The fine-tuned large language model can better adapt to the target task, which helps to improve the performance of the large language model in a specific application field;
[0100] Meanwhile, input the second role definition text into the third large language model to prompt the third large language model to act as the questioner in multi-round conversations. The third large language model can better understand the problem instructions in multi-round conversations to generate tasks, which can improve the generation quality of the target problem instructions.
[0101] Exemplarily, when the target task is JavaScript code generation in front-end development, the second role definition text can be "System role: You are a front-end development engineer with rich experience, especially good at JavaScript programming. I will ask you some questions related to JavaScript code generation in front-end development, and you need to give detailed answers"; the first role definition text can be "System role: You are chatting with a front-end code programming assistant, and it will give answers. You need to output problem instructions."
[0102] It can be seen that by inputting the first role definition text into the second large language model, it is equivalent to endowing the second large language model with the role of the answerer, and by inputting the second role definition text into the third large language model, it is equivalent to endowing the third large language model with the role of the questioner. The third large language model interacts with the second large language model by simulating the question operations of relevant personnel, realizing the multi-round conversation by calling the second large language model and the third large language model. In a real scenario, a reliable conversation assistant often needs to have the ability to have a multi-round conversation with the object and be able to understand and solve the object's problems based on the multi-round conversation. Therefore, through the multi-round conversation between the second large language model and the third large language model to construct multi-round conversation data, it can better imitate the real conversation scenario, and higher-quality sample data can be obtained subsequently.
[0103] In a possible implementation, the second role definition text can be used to prompt the third large language model to act as the questioner for the target task in multi-round conversations. The third large language model can generate target problem instructions related to the target task, and relevant sample data related to the target task can be obtained subsequently. Use the sample data related to the target task to train the large language model to achieve fine-tuning of the large language model. The fine-tuned large language model can better adapt to the target task, which helps to improve the performance of the large language model in a specific application field.
[0104] In a possible implementation, the initial question instruction is used as the question instruction for the first input to the second largest language model. The second largest language model and the third largest language model are called for multiple rounds of conversation. Specifically, the initial question instruction can be input into the second largest language model for answer prediction to generate the target answer text for the first conversation round. A second prompt text is constructed to prompt the third largest language model to generate a question instruction according to the target answer text whenever the target answer text is received. The second prompt text and the target answer text of the first conversation round are input into the third largest language model for question instruction prediction to generate the target question instruction for the next conversation round. The target question instruction for the next conversation round is input into the second largest language model again for answer prediction until the number of conversation rounds is equal to the preset number threshold.
[0105] Among them, the number threshold is the preset number of conversation rounds. For example, the number threshold can be set to 5, or set to other values. Inputting the initial question instruction into the second largest language model for answer prediction can specifically be taking the set of multiple initial question instructions as the seed instruction set. The initial question instructions in the seed instruction set can be regarded as seeds. The initial question instruction input into the second largest language model can be selected sequentially from the seed instruction set, or randomly selected from the seed instruction set.
[0106] Based on this, the initial question instruction is used as the question instruction for the first conversation round. After the initial question instruction is input into the second largest language model, the second largest language model answers the initial question instruction to generate the target answer text. Assuming that the initial question instruction is a question raised for a specific downstream task, the target answer text is the answer made for that downstream task. Then, the second prompt text and the target answer text of the first conversation round are input into the third largest language model to generate the target question instruction for the next conversation round. Since the second prompt text can prompt the third largest language model to generate a question instruction according to the target answer text, the target question instruction is also a question raised for that downstream task. In subsequent conversation rounds, the target answer text generated by the second largest language model and the target question instruction generated by the third largest language model can be considered to be for the same downstream task. Subsequently, sample data related to the downstream task can be obtained, and the large language model can be trained using the sample data related to the downstream task to fine-tune the large language model. The fine-tuned large language model can better adapt to the downstream task, which helps to improve the performance of the large language model in a specific application field.
[0107] Specifically, while inputting the target response text of the first conversation turn into the third large language model, the second prompt text is input into the third large language model, enabling the third large language model to understand the question instruction generation task. When inputting the target response text of subsequent conversation turns into the third large language model, the second prompt text can be input again. Since the third large language model has strong context processing capabilities, the third large language model will consider the previously input second prompt text when generating a new target question instruction, so the second prompt text does not necessarily need to be input again.
[0108] Exemplarily, the second prompt text can be "You need to give question instructions in combination with the other party's response". When the target task is JavaScript code generation in front-end development, the second prompt text can also be "You need to give question instructions in combination with the other party's response, for example, any question instructions related to JavaScript code generation in front-end development".
[0109] In a possible implementation, the second prompt text and the target response text of the first conversation turn are input into the third large language model for question instruction prediction to generate the target question instruction for the next conversation turn. Specifically, the initial question instruction can be added to the second prompt text to obtain the target prompt text; the target prompt text and the target response text of the first conversation turn are input into the third large language model for question instruction prediction to generate the target question instruction for the next conversation turn;
[0110] Among them, the target prompt text is used to prompt the third large language model to generate question instructions based on each question instruction input into the second large language model and the target response text of each conversation turn; adding the initial question instruction to the second prompt text is equivalent to updating the second prompt text and can obtain the target prompt text; inputting the target prompt text into the third large language model enables the third large language model to refer to the target response text and question instructions of each historical conversation turn when generating the target question instruction.
[0111] Specifically, referring to Figure 4 , Figure 4 is an optional structural schematic diagram of the multi-round conversation provided by the embodiments of the present application.
[0112] Among them, since the initial question instruction is equivalent to the question instruction of the first conversation turn, when the third large language model generates the target question instruction of the second conversation turn, the initial question instruction belongs to the question instruction of the historical conversation turn, and the target prompt text needs to include the initial question instruction. Then, the target prompt text and the target response text of the first conversation turn are input into the third large language model, and the third large language model can generate the target question instruction of the second conversation turn based on the initial question instruction and the target response text of the first conversation turn;
[0113] It can be seen that assuming the initial question instruction is a question proposed for a specific downstream task, the target answer text corresponding to the initial question instruction is an answer made for the downstream task. The third large language model generates a target question instruction based on the question instructions and target answer texts of each conversation turn, so that the newly generated target question instruction is also a question proposed for the downstream task. The newly generated target question instruction can be closer to the real question instruction of the relevant personnel, and sample data related to the downstream task can be obtained subsequently. The large language model is trained using the sample data related to the downstream task to fine-tune the large language model. The fine-tuned large language model can better adapt to the downstream task, which helps to improve the performance of the large language model in a specific application field.
[0114] Exemplarily, when the target task is JavaScript code generation in front-end development, the initial question instruction can be "How to implement an image carousel function using JavaScript?"; adding the initial question instruction to the second prompt text can obtain the target prompt text;
[0115] The target prompt text can be "You have an initial question: How to implement an image carousel function using JavaScript? You need to give a question instruction based on the other party's reply; you need to give a question instruction based on the other party's reply, such as: 1. Questions related to the previous question instruction; 2. Further questions regarding the previous answer of the other party; 3. Any question instruction related to JavaScript code generation in front-end development".
[0116] In addition, the format of the newly generated target question instruction can also be restricted in the target prompt text. For example, a requirement can be added to the target prompt text: "Your output format is: Instruction: XXXX".
[0117] Step 203: Construct sample data based on the initial question instruction, the target question instructions in the multi-round conversation, and the target answer texts in the multi-round conversation.
[0118] Based on this, sample data is constructed through the initial question instruction, the target question instruction, and the target answer text. Since the first large language model can generate high-quality and diverse initial question instructions, and the second large language model can generate high-quality and diverse target answer texts, and the third large language model can generate high-quality and diverse target question instructions, it is possible to enhance the complexity and diversity of the question instructions by combining multiple large language models, thereby efficiently obtaining high-quality and diverse sample data;
[0119] Therefore, subsequent use of sample data to train a large language model and fine-tune the large language model can help improve the performance of the large language model in specific application fields, and stimulate a large language model that can correctly understand various instructions and provide high-quality responses.
[0120] In a possible implementation, obtain a first prompt text for prompting the first large language model to generate problem instructions. Specifically, it can be to obtain the task information of the target task, where the target task is a downstream task trained based on sample data; based on the task information, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information.
[0121] Based on this, the target task is a downstream task in a specific application field. Construct a first prompt text based on the task information of the target task. Subsequently, sample data related to the target task can be obtained, and the large language model can be trained using the sample data related to the target task to fine-tune the large language model. The fine-tuned large language model can better adapt to the target task, which helps improve the performance of the large language model in specific application fields.
[0122] Specifically, the task content included in the task information can be filled into a preset first prompt template to obtain the first prompt text. For example, assume that the task information of the target task is "generate JavaScript code used in front-end development". The task content included in the task information includes code generation, front-end development, and JavaScript. The first prompt template can be "Please provide some <task content> problem instructions". Based on the preset prompt construction strategy, the task content included in the task information is constructed accordingly, and the construction result is filled into <task content> in the first prompt template to obtain the first prompt text as "Please provide some problem instructions related to JavaScript code generation in front-end development".
[0123] In a possible implementation, based on the task information, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information. Specifically, according to the task information, determine a target node in the preset task tree that matches the task information, where the task tree includes multiple layers of nodes, and each node carries its own candidate keywords; use the candidate keywords carried by the target node as the target keywords, and based on the target keywords, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information.
[0124] Among them, the task tree includes multiple layers of nodes. The candidate keywords carried by each node are used to indicate the specific task content of the target task. Two nodes with a parent-child relationship in the task tree are connected to each other. Based on the connection relationship between each node, each node can correspond to a specific downstream task. If there is a parent-child relationship between two nodes in the task tree, it means that there is a hierarchical relationship between the downstream tasks corresponding to these two nodes. The task granularity ranges corresponding to nodes at different levels are usually different, and the task granularity ranges corresponding to nodes at the same level are usually the same.
[0125] Specifically, referring to Figure 5 , Figure 5 is an optional structural schematic diagram of the task tree provided by the embodiment of the present application.
[0126] Among them, assume that the candidate keyword carried by the node node11 as the parent node is code generation. Assume that the child nodes corresponding to the node node11 include the node node22, and the candidate keyword carried by the node node22 can be mathematical reasoning. It is equivalent that the downstream task task11 corresponding to the node node11 is a code generation task, and the downstream task task22 corresponding to the node node22 is a code generation task for mathematical reasoning. It can be seen that the downstream task task22 can be regarded as a subtask of the downstream task task11. There is a hierarchical relationship between the downstream task task22 and the downstream task task11. The downstream task task22 is a more specific task than the downstream task task11, that is, the task granularity range of the downstream task task22 is smaller than the task granularity range of the downstream task task11. Therefore, the task granularity ranges corresponding to nodes at different levels are different;
[0127] Furthermore, assume that the child nodes corresponding to the node node22 include the node node32 and the node node33. The candidate keyword carried by the node node32 can be junior high school math problems, and the candidate keyword carried by the node node33 can be high school math problems. It is equivalent that the downstream task task32 corresponding to the node node32 is a code generation task for mathematical reasoning about junior high school math problems, and the downstream task task33 corresponding to the node node33 is a code generation task for mathematical reasoning about high school math problems. It can be seen that both the downstream task task32 and the downstream task task33 can be regarded as subtasks of the downstream task task22. There is a hierarchical relationship between the downstream task task32 and the downstream task task22, and between the downstream task task33 and the downstream task task22. The task granularity ranges of the downstream task task32 and the downstream task task33 are the same. Therefore, the task granularity ranges corresponding to nodes at the same level are the same.
[0128] There are various ways to determine the target node, and two of them will be described in detail below.
[0129] Method 1: Refer to Figure 6 , Figure 6 which is an optional interface schematic diagram of the task configuration interface provided by the embodiments of the present application.
[0130] Among them, the task configuration interface may be configured with multiple first checkbox controls 610, multiple second checkbox controls 620, multiple third checkbox controls 630, and a first configuration determination control 640. The task configuration interface may also be configured with more hierarchical checkbox controls, which are not limited in the embodiments of the present application;
[0131] Taking the checkbox controls with three levels as an example, each checkbox control has a corresponding task content, and the text of the corresponding task content will be displayed on one side of each checkbox control. The first checkbox controls 610, the second checkbox controls 620, and the third checkbox controls 630 respectively correspond to the nodes at different levels of the task tree. Each first checkbox control 610 has an association relationship with the corresponding node at the second level, each second checkbox control 620 has an association relationship with the corresponding node at the third level, and each third checkbox control 630 has an association relationship with the corresponding node at the fourth level. Therefore, there is also a hierarchical relationship among the first checkbox controls 610, the second checkbox controls 620, and the third checkbox controls 630.
[0132] In the task configuration interface of the terminal, multiple first checkbox controls 610 are usually displayed. Relevant personnel can click on any one of the displayed first checkbox controls 610 in the task configuration interface. At one moment, at most one first checkbox control 610 is allowed to be in the checked state. When one of the first checkbox controls 610 is in the checked state, each second checkbox control 620 at the next level of the first checkbox control 610 will be displayed in the task configuration interface;
[0133] Relevant personnel can click on any one of the displayed second checkbox controls 620 in the task configuration interface. At one moment, at most one second checkbox control 620 is allowed to be in the checked state. When one of the second checkbox controls 620 is in the checked state, each third checkbox control 630 at the next level of the second checkbox control 620 will be displayed in the task configuration interface;
[0134] Relevant personnel can click on any one of the displayed third checkbox controls 630 in the task configuration interface. At one moment, at most one third checkbox control 630 is allowed to be in the checked state;
[0135] After the relevant personnel complete the recheck, the first configuration determination control 640 can be triggered. The relevant personnel can trigger the first configuration determination control 640 after selecting the first checkbox control 610, or can also trigger the first configuration determination control 640 after selecting the first checkbox control 610 and the second checkbox control 620. By triggering the first configuration determination control 640, the terminal can respond to the operation on the first configuration determination control 640, and determine the target node in the task tree according to the check states of each checkbox control, and use the node corresponding to the last checkbox control in the checked state as the target node. In Figure 6 Among them, the first checkbox control 610 with the task content of code generation is in the checked state, the second checkbox control 620 with the task content of front-end development is in the checked state, and the third checkbox control 630 with the task content of front-end JavaScript is in the checked state. Therefore, the target node of JavaScript can be determined in the task tree.
[0136] Method 2: Refer to Figure 7 , Figure 7 is another optional interface schematic diagram of the task configuration interface provided by the embodiments of the present application.
[0137] Among them, the task configuration interface can be configured with a text input control 710 and a second configuration determination control 720;
[0138] In the task configuration interface of the terminal, the relevant personnel can input the task information of the target task through the text input control 710, and then the relevant personnel can trigger the second configuration determination control 720. By triggering the second configuration determination control 720, the terminal can respond to the operation on the second configuration determination control 720, and determine the target node in the task tree according to the task information in the text input control 710.
[0139] Specifically, assuming that the task information is "generate relevant code of JavaScript in front-end development", it can be referred to Figure 5, first determine the nodes that match the task information among the nodes at the second level of the task tree. Assume that the node that matches at the second level is node node11. For example, the candidate keyword carried by node node11 is code generation. Then determine the nodes that match the task information among the child nodes of node node11. Assume that the node that matches at the third level is node node21. For example, the candidate keyword carried by node node21 is front-end development. Then determine the nodes that match the task information among the child nodes of node node21. Assume that the node that matches at the fourth level is node node31. For example, the candidate keyword carried by node node31 is JavaScript. Then determine the nodes that match the task information among the child nodes of node node31. Assume that there is no matching node at the fourth level, and take node node31 as the target node.
[0140] It can be seen that among the nodes that match the task information, the task granularity range corresponding to the target node is the smallest. Take the candidate keyword carried by the target node as the target keyword, and the task content indicated by the target keyword is the most specific. Since the first prompt text is constructed based on the target keyword, the first prompt text can more specifically prompt the first large language model to generate an initial question instruction for the target task. For example, if the target keyword carried by the target node is JavaScript, the constructed first prompt text can be "Please provide some question instructions related to JavaScript".
[0141] In a possible implementation manner, according to the target keyword, construct the first prompt text for prompting the first large language model to generate a question instruction according to the task information. Specifically, it can be:
[0142] When the level where the target node is located is the target level, according to the target keyword, construct the first prompt text for prompting the first large language model to generate a question instruction according to the task information, where the target level is the level immediately after the level where the root node of the task tree is located;
[0143] Among them, after determining the target node that matches the task information in the task tree, judge whether the level where the target node is located belongs to the target level. When the level where the target node is located is the target level, the task information only matches the target node in the task tree, and the task information can be more accurately characterized by the target keyword. Therefore, construct the first prompt text according to the target keyword.
[0144] Alternatively, when the level where the target node is located is a level after the target level, determine an associated node associated with the target node in the task tree, use the candidate keywords carried by the associated node as associated keywords, and construct a first prompt text for prompting the first large language model to generate a question instruction based on the task information according to the associated keywords and the target keywords, where the level where the associated node is located is before the level where the target node is located.
[0145] Among them, in the case where the level where the target node is located is a level after the target level, the task information will match multiple nodes in the task tree. It is necessary to determine an associated node associated with the target node, and then use the candidate keywords of the associated node as associated keywords. The task information can be more accurately characterized by the combination of the associated keywords and the target keywords. Therefore, construct the first prompt text according to the associated keywords and the target keywords.
[0146] Exemplarily, the target level is the level after the level where the root node in the task tree is located, that is, the target level is the second level. Assuming that the level where the target node is located is the fourth level, that is, the level where the target node is located is a level after the target level, it is possible to determine that the node connected to the target node in the third level is the first associated node, and it is possible to determine that the node connected to the first associated node in the target level is the second associated node. Then, construct the first prompt text according to the associated keywords carried by the two associated nodes and the target keywords carried by the target node. The first prompt text can more specifically prompt the first large language model to generate an initial question instruction for the target task. For example, the target keyword carried by the target node is JavaScript, the associated keyword carried by the first associated node is front-end development, and the associated keyword carried by the second associated node is code generation. The constructed first prompt text can be "Please provide some question instructions related to JavaScript code generation in front-end development."
[0147] In a possible implementation manner, the level after the level where the root node in the task tree is located is the target level, and the node located at the target level carries a third role definition text. Before using the candidate keywords carried by the target node as target keywords, the sample data generation method further includes: using the node associated with the target node and located at the target level as the target role definition node, and using the third role definition text carried by the target role definition node as the target role definition text, where the target role definition text is used to prompt the first large language model to be the questioner for the target task; inputting the target role definition text into the first large language model.
[0148] Based on this, the target level of the task tree can include one or more nodes. The nodes at the target level in the task tree are used as role definition nodes. Since the candidate keywords carried by different role definition nodes are different, which is equivalent to the specific task contents corresponding to different role definition nodes being different, it is necessary to configure the corresponding third role definition text for each role definition node. The content of the third role definition text is used to display the role of defining the first large language model. Moreover, since the task granularity range corresponding to the role definition node in the task tree is the largest, and the task granularity ranges corresponding to the other nodes associated with the role definition node belong to the sub-range of the task granularity range corresponding to the role definition node, the target role definition text can be adapted to the target role definition node and the other nodes associated with the target role definition node. By inputting the target role definition text into the first large language model, the first large language model can better understand the target task and improve the generation quality of the initial problem instruction.
[0149] Refer again to Figure 5 , assuming that the candidate keyword carried by the role definition node node11 is code generation, the third role definition text configured for the role definition node node11 can be "System role: You are an assistant good at code programming and problem-related", and for another example, the candidate keyword carried by the role definition node node12 is travel planning, and the third role definition text configured for the role definition node node12 can be "System role: You are a virtual tour guide good at travel planning";
[0150] Assume that the role definition node node11 is the target role definition node, and the target role definition text is "System role: You are an assistant good at code programming and problem-related". By inputting the target role definition text into the first large language model, the first large language model can better understand the target task of code generation and improve the generation quality of the initial problem instruction.
[0151] In a possible implementation manner, the number of initial problem instructions is multiple. Before using the initial problem instructions as the first problem instructions input to the second large language model, the sample data generation method further includes: randomly sampling a first problem instruction from the multiple initial problem instructions; constructing a third prompt text for prompting the first large language model to perform instruction expansion according to a preset expansion strategy; inputting the third prompt text and the first problem instruction into the first large language model to perform instruction expansion on the first problem instruction and generate a first expanded problem instruction, and using the first expanded problem instruction as the initial problem instruction.
[0152] Specifically, refer to Figure 8 , Figure 8 This is the second optional structural schematic diagram of the first large language model provided by the embodiments of this application.
[0153] Among them, the first problem instruction is randomly sampled from multiple initial problem instructions. Specifically, a set of multiple initial problem instructions generated by the first large language model can be used as a seed instruction set, and the initial problem instructions in the seed instruction set can be regarded as seeds. The first problem instruction is randomly sampled from the seed instruction set; the expansion strategy can be a strategy text for complicating problem instructions. The expansion strategy is used to prompt the first large language model on how to complicate problem instructions. The expansion strategy can be used to make requirements for the first expanded problem instruction generated by the first large language model. For example, the expansion strategy can make requirements for the number of words, word usage, number of reasoning steps, or complexity of the first expanded problem instruction, etc. The expansion strategies required for different downstream tasks are usually different. Therefore, multiple expansion strategies need to be preset to cope with different downstream tasks.
[0154] Based on this, a third prompt text is constructed based on the expansion strategy. The third prompt text can prompt the first large language model to handle the task of complicating problem instructions. The third prompt text and the first problem instruction are input into the first large language model, so that the first large language model can complicate the first problem instruction. The first large language model can generate the first expanded problem instruction based on the model input. The number of the first expanded problem instructions can be one or more, which is not limited in the embodiments of the present application; the first expanded problem instruction can be regarded as the result of complicating the first problem instruction. Since relevant personnel may input relatively complex problem instructions, compared with the first problem instruction with lower complexity, the first expanded problem instruction with higher complexity is usually closer to the relatively complex problem instructions input by relevant personnel in the real scenario. When using the first expanded problem instruction to train the large language model in subsequent downstream tasks, the training effect of the large language model can be effectively improved.
[0155] Specifically, before taking the initial problem instruction as the first problem instruction input to the second large language model, multiple first expanded problem instructions can be generated through multiple complicating rounds. In each complicating round, the first problem instruction needs to be randomly sampled from multiple initial problem instructions, so that some initial problem instructions can be complicated. Using the first expanded problem instruction obtained by complicating the problem instruction as the initial problem instruction can further enhance the diversity of the initial problem instructions and improve the training effect of the large language model subsequently.
[0156] Taking the downstream task of JavaScript code generation in front-end development as an example, the first expanded problem instruction can be obtained in multiple ways. Here, taking one of the ways as an example, the process of obtaining the first expanded problem instruction will be described in detail.
[0157] Method 1: Assume that the first question instruction randomly sampled is "How to implement an image carousel function using JavaScript?", and assume that the extension strategy is "Add new constraints and requirements to the original instruction to increase the content length of the instruction". The third prompt text constructed based on this extension strategy can be "Please increase the complexity of the given question instruction. The following methods can be used to increase the complexity of the question instruction: Method 1: Add new constraints and requirements to the original instruction to increase the content length".
[0158] Fill the third prompt text and the first question instruction into the second prompt template. The first input text obtained is "Please increase the complexity of the given question instruction. The following methods can be used to increase the complexity of the instruction: Method 1: Add new constraints and requirements to the original instruction to increase the content length. Original instruction: How to implement an image carousel function using JavaScript?".
[0159] Then input the first input text into the first large language model, and the first large language model generates the first extended question instruction.
[0160] Among them, the content of the fourth role definition text is used to display the role of defining the first large language model. Before inputting the first input text into the first large language model, input the fourth role definition text into the first large language model to prompt the first large language model to act as the questioner for the target task. For example, in the downstream task of JavaScript code generation in front-end development, the fourth role definition text can be "System role: You are an assistant good at code programming and problem-related". By inputting the fourth role definition text into the first large language model, the first large language model can better understand the task of complicating the question instruction and improve the generation quality of the second extended question instruction.
[0161] In a possible implementation, the number of extension strategies is multiple. The first question instruction is extended to generate an extended question instruction. Specifically, it can be: The first question instruction is extended multiple times to obtain the first extended question instruction generated by each instruction extension; among them, whenever the first question instruction is extended, at least one of the multiple extension strategies is used to extend the first question instruction.
[0162] Based on this, after inputting the third prompt text and the first question instruction into the first large language model, the first large language model will perform multiple instruction expansions on the first question instruction. Therefore, by complicating the question instruction multiple times for the same first question instruction, multiple first expanded question instructions can be obtained. Moreover, each time an instruction expansion is performed, one or more of the expansion strategies are randomly selected to expand the first question instruction, which can improve the randomness and diversity of the first expanded question instructions. Taking each first expanded question instruction as the initial question instruction can further enhance the diversity of the initial question instructions.
[0163] Specifically, assume that there are four preset expansion strategies. Expansion strategy a is "adding new constraints and requirements to the original instruction to increase the content length of the instruction", expansion strategy b is "replacing common concepts in the original instruction with more specific but less common concepts", expansion strategy c is "increasing the reasoning steps of the original instruction", and expansion strategy d is "increasing the time complexity or space complexity of the original instruction". The third prompt text constructed based on this expansion strategy can be "Please increase the complexity of the given question instruction while ensuring the integrity of the question instruction. The methods for increasing the complexity of the question instruction include but are not limited to: Method 1: Adding new constraints and requirements to the original instruction to increase the content length of the instruction; Method 2: Replacing common concepts in the original instruction with more specific but less common concepts; Method 3: Increasing the reasoning steps of the original instruction; Method 4: Increasing the time complexity or space complexity of the original instruction".
[0164] In a possible implementation, the number of initial question instructions is multiple. Before taking the initial question instructions as the first input to the question instruction of the second large language model, the sample data generation method further includes: randomly sampling a second question instruction from multiple initial question instructions; constructing a fourth prompt text for prompting the first large language model to generate a question instruction with reference to the second question instruction, inputting the fourth prompt text and the second question instruction into the first large language model for question instruction prediction, generating a second expanded question instruction, and taking the second expanded question instruction as the initial question instruction.
[0165] Specifically, referring to Figure 9 , Figure 9 This is the third optional structural schematic diagram of the first large language model provided by the embodiments of the present application.
[0166] Among them, the second problem instruction is randomly sampled from multiple initial problem instructions. Specifically, a set of multiple initial problem instructions generated by the first large language model can be used as a seed instruction set. The initial problem instructions in the seed instruction set can be regarded as seeds. A preset number of problem instructions are randomly sampled from the seed instruction set to obtain the second problem instruction. Therefore, constructing the fourth prompt text for prompting the first large language model to generate problem instructions with reference to the second problem instruction is equivalent to regarding the second problem instruction as a reference example for the problem instruction generation task. Inputting the fourth prompt text and the second problem instruction into the first large language model, the first large language model can generate a second extended problem instruction based on the model input. Then, the second extended problem instruction is used as a new initial problem instruction, and the new initial problem instruction is added to the seed instruction set. The number of second extended problem instructions can be one or more, which is not limited in the embodiments of the present application.
[0167] Based on this, under the action of the fourth prompt text, regarding the second problem instruction as a reference example, the first large language model learns the task through several reference examples organized in a demonstration form, and temporarily inserts the knowledge contained in the second problem instruction into the first large language model, so that the first large language model can better understand the current problem instruction generation task, and thus generate a second extended problem instruction with better quality. Since the first large language model generates the second extended problem instruction based on the knowledge contained in the second problem instruction, the second extended problem instruction is similar to the knowledge contained in the second problem instruction, which is equivalent to that the second extended problem instruction and the second problem instruction can be used to process the same downstream task. Therefore, the second extended problem instruction can be used as a new initial problem instruction. Moreover, the second extended problem instruction generated by the first large language model is usually different from the second problem instruction, that is, the first large language model can output generalized instructions, effectively expand the initial problem instructions, and increase the diversity of the initial problem instructions.
[0168] Specifically, before using the initial problem instruction as the first problem instruction input to the second large language model, the second extended problem instruction can be generated multiple times through multiple generation rounds. In each generation round, the second problem instruction needs to be randomly sampled from multiple initial problem instructions. Since the second extended problem instruction will be used as a new initial problem instruction, the second problem instruction sampled in the current generation round may be the second extended problem instruction generated in the previous generation round.
[0169] Taking the downstream task of JavaScript code generation in front-end development as an example, the second extended problem instruction can be obtained in multiple ways. Here, taking two of them as examples, the process of obtaining the second extended problem instruction will be described in detail.
[0170] Method 1: Assume that two second question instructions are randomly sampled. The second question instruction a is "How to implement an image carousel function using JavaScript?", and the second question instruction b is "How to create a timer using JavaScript?". The constructed fourth hint text can be "The following are some question instructions related to JavaScript code generation in front-end development.", where the specific form of the fourth hint text in the embodiments of the present application is not limited;
[0171] Fill the fourth hint text and the second question instructions into the third hint template. The obtained second input text is "The following are some question instructions related to JavaScript code generation in front-end development. Example 1: Instruction: How to implement an image carousel function using JavaScript? Example 2: Instruction: How to create a timer using JavaScript? Example 3: ";
[0172] Then input the second input text into the first large language model. The first large language model generates a second extended question instruction. Assume that the output of the first large language model is "Instruction: How to implement click event detection using JavaScript?", then the second extended question instruction is "How to implement click event detection using JavaScript?".
[0173] Method 2: Assume that two second question instructions are randomly sampled. The second question instruction a is "How to implement an image carousel function using JavaScript?", and the second question instruction b is "How to create a timer using JavaScript?". The constructed fourth hint text can be "Please refer to the following question instructions related to JavaScript code generation in front-end development to generate new question instructions.", where the specific form of the fourth hint text in the embodiments of the present application is not limited;
[0174] Concatenate the fourth hint text, the second question instruction a, and the second question instruction b in sequence. The obtained first concatenated text is "Please refer to the following question instructions related to JavaScript code generation in front-end development to generate new question instructions. How to implement an image carousel function using JavaScript? How to create a timer using JavaScript?";
[0175] Then input the first concatenated text into the first large language model. The first large language model generates a second extended question instruction. Assume that the output of the first large language model is "How to implement click event detection using JavaScript?", then the second extended question instruction is "How to implement click event detection using JavaScript?".
[0176] In a possible implementation, before calling the first large language model to predict problem instructions based on the first prompt text, the sample data generation method further includes: obtaining a target sampling temperature value; configuring the temperature parameter of the first large language model as the target sampling temperature value, where the target sampling temperature value is used to indicate the randomness degree of the output result of the first large language model.
[0177] Among them, the temperature parameter of the first large language model is a hyperparameter. A hyperparameter is a parameter configured for the first large language model before the first large language model starts learning. The hyperparameter is not obtained through model training. Usually, the hyperparameter is assigned based on existing experience to configure the hyperparameter for the first large language model. Since the initial problem instructions generated by the first large language model are usually obtained by sampling the vocabulary based on a probability distribution, the temperature parameter is used to adjust the probability distribution. The target sampling temperature value can be manually input or pre-set.
[0178] Based on this, when the target sampling temperature value is higher, the probability distribution is smoother, which is equivalent to smoothing the initial probabilities of each token in the vocabulary, increasing the generation probability of tokens with lower initial probabilities, increasing the randomness of the initial problem instructions, making the initial problem instructions more diverse, but the initial problem instructions are more likely to have quality problems, such as inaccurate grammar or inaccurate content, etc.; conversely, when the target sampling temperature value is lower, the probability distribution is sharper, reducing the generation probability of tokens with lower initial probabilities, so that the first large language model usually generates tokens with higher initial probabilities, which can improve the quality of the initial problem instructions, but reduce the randomness of the initial problem instructions; usually, on the premise of ensuring quality, a higher target sampling temperature value needs to be selected to improve the diversity of the initial problem instructions generated by the first large language model, and a diverse and high-quality seed instruction set formed by the initial problem instructions can be obtained.
[0179] Specifically, before calling the first large language model to predict problem instructions based on the first prompt text, configuring the temperature parameter of the first large language model as the target sampling temperature value is equivalent to the temperature parameter corresponding to the first large language model when generating initial problem instructions being the target sampling temperature value. Suppose the temperature parameter corresponding to the first large language model when generating the first extended problem instruction is the first sampling temperature value, and the temperature parameter corresponding to the first large language model when generating the second extended problem instruction is the second sampling temperature value. The target sampling temperature value, the first sampling temperature value, and the second sampling temperature value can be the same, partially the same, or completely different. This application embodiment does not make a limitation here. Usually, a higher first sampling temperature value and a higher second sampling temperature value also need to be selected to improve the diversity of the generation results of the first large language model, so as to obtain a diverse and high-quality seed instruction set.
[0180] In a possible implementation, the first large language model is called to predict problem instructions based on the first prompt text to generate initial problem instructions. Specifically, it can be: calling the first large language model to predict problem instructions based on the first prompt text to sequentially generate multiple initial problem instructions; among them, whenever an initial problem instruction is generated, the instruction similarity between the currently generated initial problem instruction and the historically generated initial problem instructions is determined, and when the instruction similarity is greater than or equal to a preset similarity threshold, the currently generated initial problem instruction is excluded.
[0181] Among them, calling the first large language model to predict problem instructions based on the first prompt text to sequentially generate multiple initial problem instructions, the implementation methods include but are not limited to: Method 1: After inputting the first prompt text into the first large language model, the first large language model can continuously generate new initial problem instructions until the stop condition is met. It is equivalent to inputting the first prompt text into the first large language model once, and the first large language model can sequentially generate multiple initial problem instructions based on this first prompt text. At this time, the first prompt text can be "Please provide some problem instructions related to JavaScript code generation in front-end development"; Method 2: Inputting the first prompt text into the first large language model multiple times. Whenever the first prompt text is input into the first large language model, the first large language model can generate initial problem instructions. At this time, the first prompt text can be "Please provide problem instructions related to JavaScript code generation in front-end development".
[0182] Specifically, the stop condition can be that the total number of initial problem instructions is equal to a preset first threshold, or the stop condition can be that the relevant personnel manually stop the generation process of the first large language model. For example, the relevant personnel click the stop button on the display interface or input a stop instruction. The embodiments of the present application do not limit the specific form of the stop condition.
[0183] Generally, in the same target task, the first prompt text input to the first large language model is the same. Since the initial problem instructions generated by the first large language model are usually obtained by sampling the vocabulary based on a probability distribution, the output of the first large language model is random. Even if the input text of the first large language model is the same, the first large language model can generate different initial problem instructions and can obtain multiple different initial problem instructions.
[0184] Based on this, since the output of the first large language model is random, the initial question instructions generated by the first large language model are usually different. However, the first large language model may generate similar initial question instructions. To improve the diversity of the initial question instructions, whenever an initial question instruction is generated, the instruction similarity between the currently generated initial question instruction and the historically generated initial question instructions is determined. Then, based on the size relationship between the instruction similarity and the similarity threshold, it is determined whether to retain or eliminate the currently generated initial question instruction. When the instruction similarity is less than the similarity threshold, it means that the content difference between the currently generated initial question instruction and the historically generated initial question instructions is relatively large, and the currently generated initial question instruction is retained. On the contrary, when the instruction similarity is greater than or equal to the similarity threshold, it means that the content difference between the currently generated initial question instruction and the historically generated initial question instructions is relatively small, and the currently generated initial question instruction is eliminated; since the initial question instructions with high content repetition have been eliminated, the retained initial question instructions have high diversity. Using the set of retained initial question instructions as the seed instruction set can obtain a seed instruction set with high diversity.
[0185] Among them, if the number of historically generated initial question instructions is zero, the instruction similarity between the currently generated initial question instruction and the historically generated initial question instructions is zero, and the similarity threshold is usually greater than zero, so the currently generated initial question instruction will be retained; if the number of historically generated initial question instructions is multiple, the instruction similarity between the currently generated initial question instruction and each of the historically generated initial question instructions can be calculated respectively, and then the maximum instruction similarity is selected. Based on the size relationship between this instruction similarity and the similarity threshold, it is determined whether to retain or eliminate the currently generated initial question instruction.
[0186] Among them, there are various ways to determine the instruction similarity. For example, the cosine similarity between the currently generated initial question instruction and the historically generated initial question instructions can be used as the instruction similarity. Another example is to determine the number of common tokens and the total number of tokens between the currently generated initial question instruction and the historically generated initial question instructions, and use the ratio of the number of common tokens to the total number of tokens as the instruction similarity; it is also possible to determine the instruction similarity through other means, which is not limited in this embodiment of the present application.
[0187] Among them, the similarity threshold can be a real number between 0 and 1. It can be seen that when the similarity threshold is larger, the currently generated initial problem instructions that are more similar to the historically generated initial problem instructions can be retained. This is equivalent to being able to retain the initial problem instructions with relatively small differences in the instruction content generated historically. Usually, the initial problem instructions that are less frequently retained may have quality problems, which can improve the overall quality of the initial problem instructions generated by the first large language model, but will reduce the diversity of the initial problem instructions generated by the first large language model. On the contrary, when the similarity threshold is smaller, the currently generated initial problem instructions that are less similar to the historically generated initial problem instructions can be eliminated, which can improve the diversity of the initial problem instructions generated by the first large language model, but only the initial problem instructions with relatively large differences in the instruction content generated historically will be retained. It is possible that many of the initially retained problem instructions have quality problems, such as inaccurate grammar or inaccurate content, etc., which reduces the overall quality of the initial problem instructions generated by the first large language model. Therefore, it is necessary to select an appropriate similarity threshold based on the existing experience obtained from multiple experiments or based on other strategies, so that the initial problem instructions generated by the first large language model can achieve a balance between diversity and quality, and a diverse and high-quality seed instruction set formed by the initial problem instructions can be obtained.
[0188] The complete process of the sample data generation method will be described in detail below.
[0189] Refer to Figure 10 , Figure 10 which is an optional schematic structural diagram of the sample data generation method provided by the embodiments of the present application.
[0190] Among them, the overall architecture may include an instruction generation module, an instruction complication module, and a multi-round dialogue module;
[0191] The processing process of the instruction generation module will be described in detail below.
[0192] First, the instruction generation module obtains a first prompt text for prompting the first large language model to generate problem instructions;
[0193] Then, the instruction generation module obtains a target sampling temperature value; configures the temperature parameter of the first large language model to the target sampling temperature value, where the target sampling temperature value is used to indicate the randomness of the output result of the first large language model
[0194] Then, the instruction generation module obtains the task information of the target task, where the target task is a downstream task trained based on the sample data;
[0195] Then, the instruction generation module determines a target node matching the task information in a preset task tree according to the task information, where the task tree includes multiple layers of nodes, each node carries its own candidate keywords, and the task tree is pre-constructed;
[0196] Then, the instruction generation module uses the nodes associated with the target node and located at the target level as target role definition nodes, and uses the third role definition text carried by the target role definition nodes as the target role definition text. Among them, the target role definition text is used to prompt the first large language model to be the questioner for the target task; the target role definition text is input into the first large language model;
[0197] Then, the instruction generation module uses the candidate keywords carried by the target node as the target keywords;
[0198] Then, when the level where the target node is located is the target level, the instruction generation module constructs a first prompt text for prompting the first large language model to generate a question instruction according to the task information based on the target keywords. Among them, the target level is the level immediately after the level where the root node of the task tree is located;
[0199] Or, when the level where the target node is located is a level after the target level, the instruction generation module determines an associated node associated with the target node in the task tree, uses the candidate keywords carried by the associated node as the associated keywords, and constructs a first prompt text for prompting the first large language model to generate a question instruction according to the task information based on the associated keywords and the target keywords. Among them, the level where the associated node is located is before the level where the target node is located;
[0200] Then, the instruction generation module calls the first large language model to predict question instructions based on the first prompt text, and sequentially generates multiple initial question instructions; among them, whenever an initial question instruction is generated, the instruction similarity between the currently generated initial question instruction and the historically generated initial question instructions is determined. When the instruction similarity is greater than or equal to the preset similarity threshold, the currently generated initial question instruction is eliminated. Among them, the initial question instruction is a single-round question instruction, and the set of each initial question instruction can be used as the seed instruction set, and the initial question instructions in the seed instruction set are all seed instructions;
[0201] Then, the instruction generation module randomly samples a first question instruction from multiple initial question instructions;
[0202] Then, the instruction generation module constructs a third prompt text for prompting the first large language model to perform instruction expansion according to the preset expansion strategy;
[0203] Then, the instruction generation module inputs the third prompt text and the first question instruction into the first large language model, performs multiple rounds of instruction expansion on the first question instruction, and obtains the first expanded question instruction generated by each round of instruction expansion. Among them, whenever the first question instruction is expanded, at least one of multiple expansion strategies is used to expand the first question instruction. The newly generated first expanded question instruction is also a single-round question instruction, and the first expanded question instruction can be used as the initial question instruction, and the newly generated initial question instruction is added to the seed instruction set to diversify the expansion of the seed instruction set.
[0204] The processing process of the instruction complication module is described in detail below.
[0205] First, the instruction complication module randomly samples a second question instruction from multiple initial question instructions.
[0206] Then, the instruction complication module constructs a fourth prompt text for prompting the first large language model to generate a question instruction with reference to the second question instruction, inputs the fourth prompt text and the second question instruction into the first large language model for question instruction prediction, and generates a second expanded question instruction. The newly generated second expanded question instruction is also a single-round question instruction, and the second expanded question instruction can be used as the initial question instruction.
[0207] The processing process of the multi-round dialogue module is described in detail below.
[0208] First, the multi-round dialogue module obtains the first role definition text and the second role definition text. The first role definition text is used to prompt the second large language model to act as the answerer for the target task in the multi-round dialogue, and the second role definition text is used to prompt the third large language model to act as the questioner in the multi-round dialogue. The target task is a downstream task trained based on sample data.
[0209] Then, the multi-round dialogue module inputs the first role definition text into the second large language model and inputs the second role definition text into the third large language model.
[0210] Then, the multi-round dialogue module inputs the initial question instruction into the second large language model for answer prediction to generate the target answer text for the first dialogue round.
[0211] Then, the multi-round dialogue module constructs a second prompt text for prompting the third large language model to generate a question instruction according to the target answer text whenever it receives the target answer text.
[0212] Then, the multi-round dialogue module adds the initial question instruction to the second prompt text to obtain the target prompt text.
[0213] Then, the multi-turn dialogue module inputs the target prompt text and the target response text of the first dialogue turn into the third large language model for question instruction prediction, and generates the target question instruction for the next dialogue turn. Among them, the target prompt text is used to prompt the third large language model to generate question instructions based on each question instruction input to the second large language model and the target response text of each dialogue turn;
[0214] Then, the multi-turn dialogue module inputs the target question instruction for the next dialogue turn into the second large language model for response prediction again until the number of dialogue turns is equal to the preset number threshold;
[0215] Finally, based on the initial question instruction, the target question instructions in the multi-turn dialogue, and the target response texts in the multi-turn dialogue, sample data is constructed;
[0216] Among them, sample data can also be constructed through the single-turn question instructions generated by the instruction generation module and the single-turn question instructions generated by the instruction complication module. The embodiments of the present application do not limit this here.
[0217] Based on this, the initial question instruction is generated by the first large language model, and then the initial question instruction is used as the first question instruction input to the second large language model. Equivalent to using the initial question instruction as the starting point of the multi-turn dialogue, through multi-turn dialogue between the second large language model and the third large language model, multiple target response texts and target question instructions for dialogue turns can be obtained. Then, sample data is constructed through the initial question instruction, the target question instruction, and the target response text. Since the first large language model can generate high-quality and diverse initial question instructions, and the second large language model can generate high-quality and diverse target response texts, and the third large language model can generate high-quality and diverse target question instructions, therefore, the complexity and diversity of question instructions can be improved by combining multiple large language models, so as to efficiently obtain high-quality and diverse sample data.
[0218] The sample data generation method provided by the embodiments of the present application can be applied to various scenarios.
[0219] For example, in a scenario where a target large language model is applied to an image text retrieval task, first, task information of the image text retrieval task is obtained; then, based on the task information, a first prompt text is constructed for prompting the first large language model to generate a question instruction according to the task information; then, the first prompt text for prompting the first large language model to generate a question instruction is obtained, and the first large language model is called to perform question instruction prediction based on the first prompt text to generate an initial question instruction; then, the initial question instruction is used as the first input question instruction to the second large language model, and the second large language model and the third large language model are called to have multiple rounds of conversations, where the second large language model is used to generate a target answer text according to the input question instruction, and the third large language model is used to generate a target question instruction input to the second large language model according to the target answer text; then, based on the initial question instruction, the target question instruction in the multiple rounds of conversations, and the target answer text in the multiple rounds of conversations, sample data is constructed; then, based on the sample data, the target large language model is trained so that the target large language model can better adapt to the image text retrieval task.
[0220] Specifically, the training effects of the target large language model are shown in Table 1 below:
[0221] Table 1
[0222] Model Target large language model a Target large language model b Baseline 24.39 23.17 After training based on sample data 53.66 51.83
[0223] Among them, sample data is constructed based on the sample data generation method provided in the embodiments of the present application. For example, 100,000 pieces of sample data are constructed, and then the target large language model a is trained based on the sample data, and the target large language model b is trained based on the sample data. It can be seen that both the target large language model a and the target large language model b have significant improvement effects.
[0224] It can be understood that although the steps in the above respective flowcharts are sequentially shown according to the indication of the arrows, these steps do not necessarily need to be sequentially executed according to the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times, and the execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0225] Referring to Figure 11 , Figure 11 , which is an optional structural schematic diagram of the sample data generation device provided in the embodiments of the present application. The sample data generation device 1100 includes:
[0226] The first generation module 1101 is configured to obtain a first prompt text for prompting a first large language model to generate a question instruction, call the first large language model to perform question instruction prediction based on the first prompt text, and generate an initial question instruction.
[0227] The second generation module 1102 is configured to use the initial question instruction as the first input to the question instruction of the second large language model, and call the second large language model and the third large language model to conduct multiple rounds of conversations. Among them, the second large language model is configured to generate a target answer text according to the input question instruction, and the third large language model is configured to generate a target question instruction input to the second large language model according to the target answer text.
[0228] The sample construction module 1103 is configured to construct sample data based on the initial question instruction, the target question instruction in the multiple rounds of conversations, and the target answer text in the multiple rounds of conversations.
[0229] Furthermore, the number of knowledge granularity types of the candidate text is at least one, and the second generation module 1102 is further configured to:
[0230] Obtain a first role definition text and a second role definition text. Among them, the first role definition text is used to prompt the second large language model to act as the answering party for the target task in the multiple rounds of conversations, and the second role definition text is used to prompt the third large language model to act as the questioning party in the multiple rounds of conversations. The target task is a downstream task trained based on the sample data.
[0231] Input the first role definition text into the second large language model, and input the second role definition text into the third large language model.
[0232] Furthermore, the second generation module 1102 is specifically configured to:
[0233] Input the initial question instruction into the second large language model for answer prediction to generate the target answer text for the first conversation round;
[0234] Construct a second prompt text for prompting the third large language model to generate a question instruction according to the target answer text whenever the target answer text is received;
[0235] Input the second prompt text and the target answer text of the first conversation round into the third large language model for question instruction prediction to generate the target question instruction for the next conversation round;
[0236] Input the target question instruction of the next conversation round into the second large language model to perform answer prediction again until the number of conversation rounds is equal to the preset number threshold.
[0237] Further, the above-mentioned second generation module 1102 is specifically configured to:
[0238] Add the initial problem instruction to the second prompt text to obtain the target prompt text;
[0239] Input the target prompt text and the target answer text of the first conversation turn into the third large language model for problem instruction prediction, and generate the target problem instruction of the next conversation turn. Among them, the target prompt text is used to prompt the third large language model to generate problem instructions according to each problem instruction input to the second large language model and the target answer text of each conversation turn.
[0240] Further, the above-mentioned first generation module 1101 is specifically configured to:
[0241] Obtain the task information of the target task, where the target task is a downstream task trained based on sample data;
[0242] Based on the task information, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information.
[0243] Further, the above-mentioned first generation module 1101 is specifically configured to:
[0244] According to the task information, determine the target node matching the task information in the preset task tree. The task tree includes multiple layers of nodes, and each node carries its own candidate keywords;
[0245] Use the candidate keywords carried by the target node as the target keywords, and according to the target keywords, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information.
[0246] Further, the above-mentioned first generation module 1101 is specifically configured to:
[0247] When the level where the target node is located is the target level, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information according to the target keywords;
[0248] Or, when the level where the target node is located is a level after the target level, determine the associated node associated with the target node in the task tree, use the candidate keywords carried by the associated node as the associated keywords, and according to the associated keywords and the target keywords, construct a first prompt text for prompting the first large language model to generate problem instructions according to the task information, where the level where the associated node is located is before the level where the target node is located.
[0249] Further, the node located at the target level carries the third role definition text. The above-mentioned first generation module 1101 is specifically configured to:
[0250] Nodes associated with the target node and located at the target level are used as target role definition nodes, and the third role definition text carried by the target role definition nodes is used as the target role definition text, where the target role definition text is used to prompt the first large language model to be the questioner for the target task;
[0251] According to the target keyword, construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information, and add the target role definition text to the first prompt text.
[0252] Furthermore, the number of initial question instructions is multiple, and the sample data generation device further includes a third generation module (not shown in the figure), and the third generation module is specifically used for:
[0253] Randomly sample a first question instruction from multiple initial question instructions;
[0254] Construct a third prompt text for prompting the first large language model to perform instruction expansion according to the expansion strategy based on the preset expansion strategy;
[0255] Input the third prompt text and the first question instruction into the first large language model, perform instruction expansion on the first question instruction, generate a first expanded question instruction, and use the first expanded question instruction as the initial question instruction.
[0256] Furthermore, the number of expansion strategies is multiple, and the above-mentioned third generation module is specifically used for:
[0257] Perform instruction expansion on the first question instruction multiple times to obtain the first expanded question instruction generated by each instruction expansion;
[0258] Among them, whenever instruction expansion is performed on the first question instruction, the first question instruction is expanded according to at least one of multiple expansion strategies.
[0259] Furthermore, the number of initial question instructions is multiple, and the above-mentioned first generation module 1101 is further used for:
[0260] Randomly sample a second question instruction from multiple initial question instructions;
[0261] Construct a fourth prompt text for prompting the first large language model to generate a question instruction with reference to the second question instruction, input the fourth prompt text and the second question instruction into the first large language model for question instruction prediction, generate a second expanded question instruction, and use the second expanded question instruction as the initial question instruction.
[0262] Furthermore, the above-mentioned first generation module 1101 is further used for:
[0263] Obtain the target sampling temperature value;
[0264] Configure the temperature parameter of the first large language model to a target sampling temperature value, where the target sampling temperature value is used to indicate the randomness of the output result of the first large language model.
[0265] Further, the above-mentioned first generation module 1101 is specifically configured to:
[0266] Call the first large language model to predict problem instructions based on the first prompt text, and sequentially generate multiple initial problem instructions;
[0267] Wherein, whenever an initial problem instruction is generated, determine the instruction similarity between the currently generated initial problem instruction and the historically generated initial problem instruction. When the instruction similarity is greater than or equal to a preset similarity threshold, eliminate the currently generated initial problem instruction.
[0268] The above sample data generation device 1100 and the sample data generation method applied to the central node are based on the same inventive concept. By generating initial problem instructions through the first large language model, and then using the initial problem instructions as the first input to the problem instructions of the second large language model, which is equivalent to using the initial problem instructions as the starting point of multiple rounds of conversations. Through multiple rounds of conversations between the second large language model and the third large language model, target answer texts and target problem instructions for multiple conversation rounds can be obtained. Then, sample data is constructed through the initial problem instructions, target problem instructions, and target answer texts. Since the first large language model can generate high-quality and diverse initial problem instructions, the second large language model can generate high-quality and diverse target answer texts, and the third large language model can generate high-quality and diverse target problem instructions, therefore, the complexity and diversity of problem instructions can be improved by combining multiple large language models, so as to efficiently obtain high-quality and diverse sample data.
[0269] The electronic device provided in the embodiments of the present application for executing the above sample data generation method may be a terminal. Refer to Figure 12 , Figure 12 which is a partial structural block diagram of the terminal provided in the embodiments of the present application. The terminal includes: a camera component 1210, a memory 1220, an input unit 1230, a display unit 1240, a sensor 1250, an audio circuit 1260, a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290, etc. Those skilled in the art can understand that Figure 12 the terminal structure shown in
[0270] The camera component 1210 can be used to collect images or videos. Optionally, the camera component 1210 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the background blurring function by fusing the main camera and the depth camera, the panoramic shooting and VR (Virtual Reality) shooting functions or other fusion shooting functions by fusing the main camera and the wide-angle camera.
[0271] The memory 1220 can be used to store software programs and modules. The processor 1280 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 1220.
[0272] The input unit 1230 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1230 can include a touch panel 1231 and other input devices 1232.
[0273] The display unit 1240 can be used to display the input information or the provided information and various menus of the terminal. The display unit 1240 can include a display panel 1241.
[0274] The audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface.
[0275] The power supply 1290 can be alternating current, direct current, a primary battery, or a rechargeable battery.
[0276] The number of the sensors 1250 can be one or more. The one or more sensors 1250 include, but are not limited to: an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, etc. Among them:
[0277] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1280 can control the display unit 1240 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for game or user motion data collection.
[0278] The gyroscope sensor can detect the body direction and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the 3D actions of the user on the terminal. Based on the data collected by the gyroscope sensor, the processor 1280 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0279] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1240. When the pressure sensor is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1280 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1240, the processor 1280 can control the operable controls on the UI interface according to the pressure operation of the user on the display unit 1240. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0280] The optical sensor is used to collect the ambient light intensity. In one embodiment, the processor 1280 can control the display brightness of the display unit 1240 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1240 is increased; when the ambient light intensity is low, the display brightness of the display unit 1240 is decreased. In another embodiment, the processor 1280 can also dynamically adjust the shooting parameters of the camera module 1210 according to the ambient light intensity collected by the optical sensor.
[0281] In this embodiment, the processor 1280 included in the terminal can execute the sample data generation method of the previous embodiment.
[0282] The electronic device for executing the above sample data generation method provided by the embodiments of the present application can also be a server. Refer to Figure 13 , Figure 13This is a partial structural block diagram of the server provided by the embodiments of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1322 (for example, one or more processors) and a memory 1332, and one or more storage media 1330 (for example, one or more mass storage devices) for storing application programs 1342 or data 1344. Among them, the memory 1332 and the storage media 1330 may be transient storage or persistent storage. The program stored in the storage media 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1300. Further, the central processing unit 1322 may be configured to communicate with the storage media 1330 and execute a series of instruction operations in the storage media 1330 on the server 1300.
[0283] The server 1300 may further include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1358, and / or one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0284] The processor in the server 1300 may be used to execute the sample data generation method.
[0285] The embodiments of the present application further provide a computer-readable storage medium for storing program codes for executing the sample data generation method of the foregoing various embodiments.
[0286] The embodiments of the present application further provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the sample data generation method described above.
[0287] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0288] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c may be single or plural.
[0289] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0290] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0291] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0292] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0293] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0294] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0295] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A method for generating sample data, characterized in that, Including: Obtain a first prompt text for prompting a first large language model to generate a question instruction, call the first large language model to perform question instruction prediction based on the first prompt text, and generate an initial question instruction; Use the initial question instruction as the first input to the question instruction of the second large language model, and call the second large language model and the third large language model to have multiple rounds of conversations. Among them, the second large language model is used to generate a target answer text according to the input question instruction, and the third large language model is used to generate a target question instruction input to the second large language model according to a second prompt text. The second prompt text is used to prompt the third large language model to generate the target question instruction according to the target answer text whenever the target answer text is received; Construct sample data based on the initial question instruction, the target question instruction in the multiple rounds of conversations, and the target answer text in the multiple rounds of conversations.
2. The sample data generation method according to claim 1, wherein Before using the initial question instruction as the first input to the question instruction of the second large language model, the sample data generation method further includes: Obtain a first role definition text and a second role definition text. Among them, the first role definition text is used to prompt the second large language model to act as an answerer for the target task in the multiple rounds of conversations, and the second role definition text is used to prompt the third large language model to act as a questioner in the multiple rounds of conversations. The target task is a downstream task trained based on the sample data; Input the first role definition text into the second large language model, and input the second role definition text into the third large language model.
3. The sample data generation method according to claim 1, characterized in that, Using the initial question instruction as the first input to the question instruction of the second large language model and calling the second large language model and the third large language model to have multiple rounds of conversations includes: Input the initial question instruction into the second large language model for answer prediction to generate a target answer text for the first conversation round; Construct the second prompt text, and input the second prompt text and the target answer text of the first conversation round into the third large language model for question instruction prediction to generate a target question instruction for the next conversation round; Input the target question instruction of the next conversation round into the second large language model for answer prediction again until the number of conversation rounds is equal to a preset number threshold.
4. The method for generating sample data according to claim 3, wherein, Inputting the second prompt text and the target answer text of the first conversation round into the third large language model for question instruction prediction to generate a target question instruction for the next conversation round includes: Add the initial question instruction to the second prompt text to obtain a target prompt text; Input the target prompt text and the target answer text of the first conversation round into the third large language model for question instruction prediction to generate a target question instruction for the next conversation round. Among them, the target prompt text is used to prompt the third large language model to generate a question instruction according to each question instruction input to the second large language model and the target answer text of each conversation round.
5. The sample data generation method according to claim 1, wherein The obtaining of the first prompt text for prompting the first large language model to generate a question instruction includes: Obtain the task information of the target task, where the target task is a downstream task trained based on the sample data; Based on the task information, construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information.
6. The sample data generation method according to claim 5, characterized in that The constructing of the first prompt text for prompting the first large language model to generate a question instruction according to the task information includes: According to the task information, determine a target node in a preset task tree that matches the task information, where the task tree includes multiple layers of nodes, and each of the nodes carries its respective candidate keywords; Use the candidate keywords carried by the target node as target keywords, and according to the target keywords, construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information.
7. The sample data generation method according to claim 6, wherein The constructing of the first prompt text for prompting the first large language model to generate a question instruction according to the target keywords includes: When the level where the target node is located is the target level, construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information, where the target level is the level after the level where the root node of the task tree is located; Or, when the level where the target node is located is a level after the target level, determine an associated node associated with the target node in the task tree, use the candidate keywords carried by the associated node as associated keywords, and according to the associated keywords and the target keywords, construct a first prompt text for prompting the first large language model to generate a question instruction according to the task information, where the level where the associated node is located is before the level where the target node is located.
8. The sample data generation method according to claim 6, characterized in that, The level after the level where the root node of the task tree is located is the target level, and the nodes located at the target level carry third role definition texts. Before using the candidate keywords carried by the target node as target keywords, the sample data generation method further includes: Use the node associated with the target node and located at the target level as the target role definition node, and use the third role definition text carried by the target role definition node as the target role definition text, where the target role definition text is used to prompt the first large language model to be the questioner for the target task; Input the target role definition text into the first large language model.
9. The method for generating sample data according to claim 1, wherein The number of the initial question instructions is multiple. Before using the initial question instruction as the first question instruction input to the second large language model, the sample data generation method further includes: Randomly sample a first question instruction from the multiple initial question instructions; Based on a preset expansion strategy, construct a third prompt text for prompting the first large language model to perform instruction expansion according to the expansion strategy. Input the third prompt text and the first question instruction into the first large language model, expand the instruction of the first question instruction to generate a first expanded question instruction, and use the first expanded question instruction as the initial question instruction.
10. The sample data generation method according to claim 9, wherein The number of the expansion strategies is multiple. The expanding the instruction of the first question instruction to generate an expanded question instruction includes: Performing multiple expansions on the first question instruction to obtain a first expanded question instruction generated by each instruction expansion; Wherein, whenever the instruction of the first question instruction is expanded, the first question instruction is expanded according to at least one of the multiple expansion strategies.
11. The method for generating sample data according to claim 1, wherein The number of the initial question instructions is multiple. Before using the initial question instruction as the first question instruction input to the second large language model, the sample data generation method further includes: Randomly sampling a second question instruction from the multiple initial question instructions; Constructing a fourth prompt text for prompting the first large language model to generate a question instruction with reference to the second question instruction, inputting the fourth prompt text and the second question instruction into the first large language model for question instruction prediction, generating a second expanded question instruction, and using the second expanded question instruction as the initial question instruction.
12. The sample data generation method according to claim 1, wherein Before calling the first large language model to perform question instruction prediction based on the first prompt text, the sample data generation method further includes: Obtaining a target sampling temperature value; Configuring the temperature parameter of the first large language model as the target sampling temperature value, where the target sampling temperature value is used to indicate the randomness degree of the output result of the first large language model.
13. The sample data generation method according to claim 1, wherein The calling the first large language model to perform question instruction prediction based on the first prompt text to generate an initial question instruction includes: Calling the first large language model to perform question instruction prediction based on the first prompt text to sequentially generate multiple initial question instructions; Wherein, whenever an initial question instruction is generated, determining the instruction similarity between the currently generated initial question instruction and the historically generated initial question instruction, and when the instruction similarity is greater than or equal to a preset similarity threshold, removing the currently generated initial question instruction.
14. A sample data generation device, characterized in that, including: A first generation module, configured to obtain a first prompt text for prompting a first large language model to generate a question instruction, and call the first large language model to perform question instruction prediction based on the first prompt text to generate an initial question instruction; A second generation module, configured to use the initial question instruction as the first question instruction input to the second large language model, and call the second large language model and the third large language model to have multiple rounds of conversations, where the second large language model is configured to generate a target answer text according to the input question instruction, the third large language model is configured to generate a target question instruction input to the second large language model according to a second prompt text, and the second prompt text is used to prompt the third large language model to generate the target question instruction according to the target answer text whenever the target answer text is received; A sample construction module, configured to construct sample data based on the initial problem instruction, the target problem instruction in the multi-round conversation, and the target answer text in the multi-round conversation.
15. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the sample data generation method according to any one of claims 1 to 13.
16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the sample data generation method according to any one of claims 1 to 13.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the sample data generation method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Large language model training method and device and text processing method and device
CN117149989A