Data generation method and apparatus, and requirement processing method and apparatus
By generating dialogue streams and response data that meet specific needs, and training language models, the problem of the capability boundary of language models when users book flights and hotels is solved, realizing the improvement of language models' capabilities in specific demand directions and the controllability and fluency of data generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2026-03-19
AI Technical Summary
Existing language models have limitations in fulfilling users' needs such as booking flights and hotels, and it is difficult to improve their expected capabilities through continued training.
By generating dialogue streams and response data that meet specific needs, the language model is trained, including dialogue responses and tool call statements, ensuring data consistency and controllability to expand the capabilities of the language model.
It enhances the language model's capabilities in specific needs, improves the controllability and fluency of generated data, ensures the correctness and rationality of training data, and avoids invalid generation.
Smart Images

Figure CN2025114686_19032026_PF_FP_ABST
Abstract
Description
A method for generating data, a method and device for demand processing
[0001] The present application claims priority to the Chinese patent application No. 202411266536.8, filed on September 10, 2024, entitled "A method for generating data, a method and device for demand processing", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of artificial intelligence, and more particularly, to a method for generating data, a method and device for model training. BACKGROUND
[0003] Language models are excellent in various text understanding and generation tasks through text. However, there is an obvious ability boundary for a naked model of a language model, such as being unable to complete a user's demand for booking a ticket or a hotel. In order to enable the language model to call external tools and expand the ability boundary of the language model, the language model can be trained based on data to learn how to call external tools.
[0004] However, the data for generating the continued training language model at present is difficult to improve the language model in the direction of the expected ability. SUMMARY
[0005] The present application provides a method for generating data, a method and device for demand processing, which generates data that can be used for language model continued training, and the ability of the language model trained based on the data can be improved in the direction of the expected ability.
[0006] In a first aspect, a method for generating data is provided, comprising: generating a dialogue flow according to a first demand, the dialogue flow comprising a dialogue process conforming to the first demand; generating first demand data according to the dialogue flow; generating first response data according to the first demand data, the first demand data and the first response data being used for training a language model.
[0007] Based on the scheme provided in the embodiments of the present application, by generating the first requirement data according to the dialogue flow, and generating the first response data responding to the first requirement data, since the dialogue flow generated according to the first requirement has a limiting effect on generating the first requirement data, on the one hand, training data (i.e. the first requirement data and the first response data) capable of improving the ability of the trained language model in the direction of the first requirement can be obtained, compared with the process of generating requirement data without the limitation of dialogue flow, the first requirement data can improve the ability of the trained language model in the expected ability direction by setting the first requirement according to the expected ability; on the other hand, the dialogue flow included in the dialogue process can improve the controllability and fluency of the data generation process, which helps to optimize the data generation process.
[0008] In some possible implementation ways, the dialogue flow indicates a reply mode of the dialogue reply; when the first response data includes dialogue reply data, the first response data satisfies the reply mode; or, when the first response data includes tool calling statement data, tool calling result data satisfies the reply mode, and the tool calling result data is determined according to the tool calling statement data.
[0009] Based on the scheme provided in the embodiments of the present application, by limiting the reply mode of the dialogue reply through the dialogue flow, the constraint on the dialogue process can be further strengthened, and the controllability and fluency of the data generation process can be further improved, which helps to improve the performance of the data generation process.
[0010] In some possible implementation ways, generating the first requirement data according to the dialogue flow includes: generating a second requirement according to the dialogue process; generating the first requirement data according to the second requirement, and the first requirement data indicating the second requirement.
[0011] In some possible implementation ways, generating the first requirement data according to the dialogue flow includes: generating a third requirement according to the dialogue process and dialogue history, the dialogue history including a completed requirement, and the third requirement representing a next requirement of a last requirement in the dialogue history; generating the first requirement data according to the third requirement, and the first requirement data indicating the third requirement.
[0012] Based on the scheme provided in the embodiments of the present application, generating the second requirement or the third requirement indicated by the first requirement data according to the dialogue flow and the dialogue history can effectively generate the first requirement in the dialogue process or the next requirement of the last requirement in the dialogue history according to the dialogue flow, which can improve the controllability and fluency of the data generation process.
[0013] In some possible implementation ways, the first response data includes tool calling statement data; generating the first response data according to the first requirement data includes: generating a plurality of tool calling statements according to the first requirement data; and determining the tool calling statement data according to the consistency of the plurality of tool calling statements.
[0014] Based on the scheme provided in the embodiments of the present application, by generating a plurality of tool calling statements according to the first requirement data, and then determining tool calling statement data including tool calling statements meeting the consistency requirement, on the one hand, training data capable of expanding the capability boundary of the trained language model can be obtained, and the trained language model can be enabled to call external tools; on the other hand, according to the consistency of the plurality of tool calling statements, appropriate tool calling statement data can be determined, and the correctness of the tool calling statement data is improved.
[0015] In some possible implementation ways, when the tool names and / or parameter arguments indicated by at least two tool calling statements in the plurality of tool calling statements are same, the tool calling statement data includes any one of the at least two tool calling statements; or when the tool names and / or parameter arguments indicated by at least two tool calling statements in the plurality of tool calling statements are not same, and the sum of the number of the plurality of tool calling statements meets the first threshold, the tool calling statement data includes any one of the plurality of tool calling statements and a first mask mark, and the first mask mark is used to indicate that the any one tool calling statement is wrong.
[0016] For example, the consistency of the plurality of tool calling statements can be determined according to the plurality of tool calling statements and the first requirement data.
[0017] Based on the scheme provided in the embodiments of the present application, by the tool names and / or parameter arguments of the plurality of tool calling statements, and / or whether the number of the plurality of tool calling statements meets the first threshold, the tool calling statement data is determined, on the one hand, it can be ensured that the training data generated for language model learning is all reasonable data verified, the correctness of the training data is improved, which is helpful for the language model to learn correct knowledge in the training process, and reduces the illusion of the language model obtained by continuing to train according to the training data; on the other hand, it can avoid generating tool calling statements infinitely.
[0018] In some possible implementation ways, the first response data includes dialogue reply data; the first response data is generated according to the first requirement data, including: generating a plurality of dialogue replies according to the first requirement data; and determining dialogue reply data according to the consistency of the plurality of dialogue replies.
[0019] Based on the scheme provided in the embodiments of the present application, on the one hand, by generating a plurality of dialogue replies according to the first demand data, and then determining the first response data including the dialogue reply data, compared with generating training data based only on tool calling scenarios, when the first demand data indicates a demand that does not require calling tools, training data that can maintain the original general ability of the trained language model can be obtained, and the diversity and authenticity of the dialogue are improved; on the other hand, according to the consistency of the plurality of dialogue replies, appropriate dialogue reply data can be determined, and the correctness of the dialogue reply data is improved.
[0020] In some possible implementation manners, when the content of at least two dialogue replies in the plurality of dialogue replies satisfies the content consistency, the dialogue reply data includes any one of the at least two dialogue replies; or when the content of at least two dialogue replies in the plurality of dialogue replies does not satisfy the content consistency, and the sum of the number of the plurality of dialogue replies satisfies the second threshold, the dialogue reply data includes any one of the plurality of dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one dialogue reply is incorrect.
[0021] For example, the consistency of the plurality of dialogue replies can be determined according to the plurality of dialogue replies and the first demand data.
[0022] Based on the scheme provided in the embodiments of the present application, by determining the dialogue reply data according to the content of the plurality of dialogue replies and / or whether the number of the plurality of dialogue replies satisfies the second threshold, on the one hand, it can be ensured that the training data generated for language model learning are all reasonable data that have been verified, the correctness of the training data is improved, which is helpful for the language model to learn correct knowledge in the training process and reduce the illusion of the language model obtained by continuing to train according to the training data; on the other hand, it can avoid generating dialogue replies infinitely.
[0023] In a second aspect, a model training method is provided, including: inputting first demand data and first response data into a language model, the first demand data being generated according to a dialogue flow, the dialogue flow being generated according to a first demand, and the dialogue flow including a dialogue process that meets the first demand; training the language model according to the first demand data and the first response data, the first response data being generated according to the first demand data.
[0024] In some possible implementation manners, the dialogue flow indicates a reply mode of the dialogue reply; when the first response data includes the dialogue reply data, the first response data satisfies the reply mode; or when the first response data includes tool calling statement data, tool calling result data determined according to the tool calling statement data satisfies the reply mode.
[0025] In some possible implementation manners, the first requirement data is generated according to a second requirement, and the first requirement data indicates the second requirement, and the second requirement is generated according to a dialogue flow.
[0026] In some possible implementation manners, the first requirement data is generated according to a third requirement, and the first requirement data indicates the third requirement, and the third requirement is generated according to the dialogue flow and a dialogue history, and the dialogue history includes requirements that have been completed, and the third requirement represents a next requirement of a last requirement in the dialogue history.
[0027] In some possible implementation manners, the first response data includes tool call statement data, and the tool call statement data is determined according to consistency of a plurality of tool call statements, and the plurality of tool call statements are generated according to the first requirement data.
[0028] In some possible implementation manners, when there are at least two tool call statements in the plurality of tool call statements that indicate same tool names and / or parameter arguments, the tool call statement data includes any one of the at least two tool call statements; or when there are not at least two tool call statements in the plurality of tool call statements that indicate same tool names and / or parameter arguments, and a sum of quantities of the plurality of tool call statements satisfies a first threshold, the tool call statement data includes any one of the plurality of tool call statements and a first mask mark, and the first mask mark is used to indicate that the any one of the tool call statements is incorrect.
[0029] In some possible implementation manners, the first response data includes dialogue reply data, and the dialogue reply data is determined according to consistency of a plurality of dialogue replies, and the plurality of dialogue replies are generated according to the first requirement data.
[0030] In some possible implementation manners, when there are at least two dialogue replies in the plurality of dialogue replies that satisfy content consistency, the dialogue reply data includes any one of the at least two dialogue replies; or when there are not at least two dialogue replies in the plurality of dialogue replies that satisfy content consistency, and a sum of quantities of the plurality of dialogue replies satisfies a second threshold, the dialogue reply data includes any one of the plurality of dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one of the dialogue replies is incorrect.
[0031] In a third aspect, a method for requirement processing is provided, including: determining a to-be-processed requirement; inputting the to-be-processed requirement into a language model to obtain a first processing result, wherein the language model is trained according to first requirement data and first response data, the first requirement data is generated according to a dialogue flow, the dialogue flow is generated according to a first requirement, the dialogue flow includes a dialogue flow process that meets the first requirement, and the first response data is generated according to the first requirement data.
[0032] In some possible implementation ways, the dialogue flow indicates a reply mode of the dialogue reply; the first response data satisfies the reply mode when the first response data comprises dialogue reply data; or the tool invocation result data satisfies the reply mode when the first response data comprises tool invocation statement data, and the tool invocation result data is determined according to the tool invocation statement data.
[0033] In some possible implementation ways, the first requirement data is generated according to a second requirement, and the first requirement data indicates the second requirement; and the second requirement is generated according to the dialogue flow.
[0034] In some possible implementation ways, the first requirement data is generated according to a third requirement, and the first requirement data indicates the third requirement; and the third requirement is generated according to the dialogue flow and a dialogue history, the dialogue history comprises a completed requirement, and the third requirement represents a next requirement of a last requirement in the dialogue history.
[0035] In some possible implementation ways, the first response data comprises tool invocation statement data, and the tool invocation statement data is determined according to consistency of a plurality of tool invocation statements, and the plurality of tool invocation statements are generated according to the first requirement data.
[0036] In some possible implementation ways, when there are at least two tool invocation statements in the plurality of tool invocation statements, and the tool names and / or parameter arguments indicated by the at least two tool invocation statements are same, the tool invocation statement data comprises any one of the at least two tool invocation statements; or when there are not at least two tool invocation statements in the plurality of tool invocation statements, and the tool names and / or parameter arguments indicated by the at least two tool invocation statements are same, and the sum of the number of the plurality of tool invocation statements satisfies a first threshold value, the tool invocation statement data comprises any one of the plurality of tool invocation statements and a first mask mark, and the first mask mark is used to indicate that the any one of the plurality of tool invocation statements is wrong.
[0037] In some possible implementation ways, the first response data comprises dialogue reply data, and the dialogue reply data is determined according to consistency of a plurality of dialogue replies, and the plurality of dialogue replies are generated according to the first requirement data.
[0038] In some possible implementation ways, when there are at least two dialogue replies in the plurality of dialogue replies, and the contents of the at least two dialogue replies satisfy content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or when there are not at least two dialogue replies in the plurality of dialogue replies, and the contents of the at least two dialogue replies satisfy content consistency, and the sum of the number of the plurality of dialogue replies satisfies a second threshold value, the dialogue reply data comprises any one of the plurality of dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one of the plurality of dialogue replies is wrong.
[0039] In a fourth aspect, a device for generating data is provided, including a processor configured to: generate a dialogue flow according to a first requirement, the dialogue flow including a dialogue procedure conforming to the first requirement; generate first requirement data according to the dialogue flow; and generate first response data according to the first requirement data, the first requirement data and the first response data being used for training a language model.
[0040] In some possible implementation manners, the dialogue flow indicates a reply manner of the dialogue reply; when the first response data includes dialogue reply data, the first response data satisfies the reply manner; or, when the first response data includes tool invocation statement data, tool invocation result data satisfying the reply manner is determined according to the tool invocation statement data.
[0041] In some possible implementation manners, the processor is specifically configured to: generate a second requirement according to the dialogue procedure; and generate the first requirement data according to the second requirement, the first requirement data indicating the second requirement.
[0042] In some possible implementation manners, the processor is specifically configured to: generate a third requirement according to the dialogue procedure and a dialogue history, the dialogue history including a completed requirement, the third requirement representing a next requirement of a last requirement in the dialogue history; and generate the first requirement data according to the third requirement, the first requirement data indicating the third requirement.
[0043] In some possible implementation manners, the first response data includes tool invocation statement data; and the processor is specifically configured to: generate a plurality of tool invocation statements according to the first requirement data; and determine the tool invocation statement data according to consistency of the plurality of tool invocation statements.
[0044] In some possible implementation manners, when at least two tool invocation statements in the plurality of tool invocation statements indicate same tool names and / or parameter arguments, the tool invocation statement data includes any one of the at least two tool invocation statements; or, when at least two tool invocation statements in the plurality of tool invocation statements do not indicate same tool names and / or parameter arguments, and a sum of a quantity of the plurality of tool invocation statements satisfies a first threshold, the tool invocation statement data includes any one of the plurality of tool invocation statements and a first mask mark, the first mask mark being used to indicate that the any one of the tool invocation statements is incorrect.
[0045] In some possible implementation manners, the first response data includes dialogue reply data; and the processor is specifically configured to: generate a plurality of dialogue replies according to the first requirement data; and determine the dialogue reply data according to consistency of the plurality of dialogue replies.
[0046] In some possible implementation manners, when the content of at least two of the multiple dialogue replies satisfies the content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or when the content of at least two of the multiple dialogue replies does not satisfy the content consistency, and the sum of the number of the multiple dialogue replies satisfies the second threshold, the dialogue reply data comprises any one of the multiple dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one dialogue reply is incorrect.
[0047] In some possible implementation manners, the apparatus is a chip.
[0048] In a fifth aspect, an apparatus for model training is provided, including a processor configured to: input first requirement data and first response data into a language model, the first requirement data being generated according to a dialogue flow, the dialogue flow being generated according to a first requirement, the dialogue flow comprising a dialogue process meeting the first requirement; and train the language model according to the first requirement data and the first response data, the first response data being generated according to the first requirement data.
[0049] In some possible implementation manners, the dialogue flow indicates a reply mode of the dialogue reply; when the first response data comprises dialogue reply data, the first response data satisfies the reply mode; or when the first response data comprises tool calling statement data, tool calling result data satisfies the reply mode, the tool calling result data being determined according to the tool calling statement data.
[0050] In some possible implementation manners, the first requirement data is generated according to a second requirement, the first requirement data indicating the second requirement, and the second requirement being generated according to the dialogue process.
[0051] In some possible implementation manners, the first requirement data is generated according to a third requirement, the first requirement data indicating the third requirement, and the third requirement being generated according to the dialogue process and a dialogue history, the dialogue history comprising a completed requirement, and the third requirement representing a next requirement of a last requirement in the dialogue history.
[0052] In some possible implementation manners, the first response data comprises tool calling statement data, the tool calling statement data being determined according to consistency of multiple tool calling statements, and the multiple tool calling statements being generated according to the first requirement data.
[0053] In some possible implementation manners, when the tool names and / or the parameter input arguments indicated by at least two of the plurality of tool calling statements are same, the tool calling statement data comprises any one of the at least two tool calling statements; or when the tool names and / or the parameter input arguments indicated by at least two of the plurality of tool calling statements are not same, and the sum of the number of the plurality of tool calling statements satisfies a first threshold, the tool calling statement data comprises any one of the plurality of tool calling statements and a first mask mark, and the first mask mark is used to indicate that the any one tool calling statement is incorrect.
[0054] In some possible implementation manners, the first response data comprises dialogue reply data, and the dialogue reply data is determined according to consistency of a plurality of dialogue replies, and the plurality of dialogue replies are generated according to the first demand data.
[0055] In some possible implementation manners, when contents of at least two of the plurality of dialogue replies satisfy content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or when contents of at least two of the plurality of dialogue replies do not satisfy content consistency, and the sum of the number of the plurality of dialogue replies satisfies a second threshold, the dialogue reply data comprises any one of the plurality of dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one dialogue reply is incorrect.
[0056] In some possible implementation manners, the apparatus is a chip.
[0057] In a sixth aspect, an apparatus for demand processing is provided, comprising a processor configured to: determine a demand to be processed; input the demand to be processed into a language model to obtain a first processing result, wherein the language model is trained according to first demand data and first response data, the first demand data is generated according to a dialogue flow, the dialogue flow is generated according to a first demand, the dialogue flow comprises a dialogue process that meets the first demand, and the first response data is generated according to the first demand data.
[0058] In some possible implementation manners, the dialogue flow indicates a reply mode of the dialogue reply; when the first response data comprises dialogue reply data, the first response data satisfies the reply mode; or when the first response data comprises tool calling statement data, tool calling result data determined according to the tool calling statement data satisfies the reply mode.
[0059] In some possible implementation manners, the first demand data is generated according to a second demand, the first demand data indicates the second demand, and the second demand is generated according to a dialogue process.
[0060] In some possible implementation, the first requirement data is generated according to a third requirement, the first requirement data indicates the third requirement, the third requirement is generated according to a dialogue flow and a dialogue history, the dialogue history comprises completed requirements, and the third requirement represents a next requirement of a last requirement in the dialogue history.
[0061] In some possible implementation, the first response data comprises tool call statement data, the tool call statement data is determined according to consistency of a plurality of tool call statements, and the plurality of tool call statements are generated according to the first requirement data.
[0062] In some possible implementation, when there are at least two tool call statements in the plurality of tool call statements indicating same tool name and / or parameter input, the tool call statement data comprises any one of the at least two tool call statements; or, when there are not at least two tool call statements in the plurality of tool call statements indicating same tool name and / or parameter input, and a sum of the number of the plurality of tool call statements satisfies a first threshold, the tool call statement data comprises any one of the plurality of tool call statements and a first mask mark, and the first mask mark is used to indicate that the any one of the plurality of tool call statements is incorrect.
[0063] In some possible implementation, the first response data comprises dialogue reply data, the dialogue reply data is determined according to consistency of a plurality of dialogue replies, and the plurality of dialogue replies are generated according to the first requirement data.
[0064] In some possible implementation, when there are at least two dialogue replies in the plurality of dialogue replies satisfying content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or, when there are not at least two dialogue replies in the plurality of dialogue replies satisfying content consistency, and a sum of the number of the plurality of dialogue replies satisfies a second threshold, the dialogue reply data comprises any one of the plurality of dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one of the plurality of dialogue replies is incorrect.
[0065] In some possible implementation, the apparatus is a chip.
[0066] In a seventh aspect, a computing apparatus is provided, comprising: a processor configured to execute computer instructions stored in a memory to cause the apparatus to perform the method in the first aspect or any possible implementation of the first aspect; or, perform the method in the second aspect or any possible implementation of the second aspect; or, perform the method in the third aspect or any possible implementation of the third aspect.
[0067] In some possible implementation manners, the processor can be a general processor, and can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, or the like; when implemented by software, the processor can be a general processor, and is implemented by reading software code stored in a memory, which can be integrated in the processor or exist independently outside the processor.
[0068] In some possible implementation manners, the apparatus further includes a memory.
[0069] In some possible implementation manners, the apparatus further includes a communication interface coupled to the processor, the communication interface configured to input and / or output information.
[0070] In some possible implementation manners, the apparatus is a chip.
[0071] An eighth aspect provides a chip or a chip system, including: a circuit configured to perform the method in the first aspect or any possible implementation manner of the first aspect; or perform the method in the second aspect or any possible implementation manner of the second aspect; or perform the method in the third aspect or any possible implementation manner of the third aspect.
[0072] A ninth aspect provides a computer program product, when a computer program in the computer program product is executed by a computing device, the method in the first aspect or any possible implementation manner of the first aspect is implemented; or the method in the second aspect or any possible implementation manner of the second aspect is implemented; or the method in the third aspect or any possible implementation manner of the third aspect is implemented.
[0073] A tenth aspect provides a computer readable storage medium, the computer readable storage medium stores a computer program or instructions, when the computer program or instructions are executed by a processor, the method in the first aspect or any possible implementation manner of the first aspect is implemented; or the method in the second aspect or any possible implementation manner of the second aspect is implemented; or the method in the third aspect or any possible implementation manner of the third aspect is implemented.
[0074] As an example, the computer readable storage includes, but is not limited to, one or more of the following: a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), a Flash memory, an electrically EPROM (EEPROM), and a hard drive.
[0075] In some possible implementation manners, the storage medium can be a nonvolatile storage medium.
[0076] The beneficial effects brought by the solutions of the second aspect to the tenth aspect can be referred to the specific description of the first aspect, and for brevity, will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0077] FIG. 1 is a schematic diagram of a training phase and an inference phase of a model.
[0078] FIG. 2 is a schematic diagram of a training process of reinforcement learning.
[0079] FIG. 3 is a schematic diagram of a supervised training manner of a deep learning model.
[0080] FIG. 4 is a schematic block diagram of a cloud scenario suitable for embodiments of the present application.
[0081] FIG. 5 is a schematic diagram of interaction between a tenant and an AI infrastructure development platform suitable for embodiments of the present application.
[0082] FIG. 6 is a schematic diagram of an agent.
[0083] FIG. 7 is a schematic diagram of a method for generating data according to an embodiment of the present application.
[0084] FIG. 8 is a schematic diagram of a flow for generating data according to an embodiment of the present application.
[0085] FIG. 9 is a schematic diagram of an apparatus for generating data or model training or demand processing according to an embodiment of the present application.
[0086] FIG. 10 is a schematic diagram of an apparatus for generating data or model training or demand processing according to an embodiment of the present application.
[0087] FIG. 11 is a schematic diagram of a chip system according to an embodiment of the present application. DETAILED DESCRIPTION
[0088] The technical solutions in the present application will be described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work shall fall within the scope of protection of the present application.
[0089] Before introducing the embodiments of the present application, the following points are first explained.
[0090] In this application, the words "example", "for example", etc. are used to mean example, illustration, or illustration. Any embodiment or design scheme described as "example" in this application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In the embodiments of this application, "of", "corresponding" and "corresponding" are sometimes mixed. It should be pointed out that when their differences are not emphasized, the meanings expressed are consistent.
[0091] The business scenarios described in the embodiments of the application are used to more clearly illustrate the technical solutions of the embodiments of the application, and do not constitute a limitation on the technical solutions provided by the embodiments of the application. Those skilled in the art can know that with the evolution of network architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the application are also applicable to similar technical problems.
[0092] In this specification, the reference "in some possible implementations" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the application. Therefore, the statements "in some possible implementations" and the like appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.
[0093] In this application, "at least one" or "at least one" means one or more, and "multiple" means two or more. The association relationship between the associated objects is described by "and / or", which means that there can be three kinds of relationships, for example, A and / or B, which can represent the following cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c, can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0094] In the present application, "first", "second", "①", "②", "#1", "#2", etc. are only for convenience of description, used to distinguish objects, and do not limit the scope of the embodiments of the present application. They are not used to describe the order or sequence of features. It should be understood that such described objects can be interchanged under appropriate circumstances to describe solutions other than the embodiments of the present application.
[0095] In order to describe the embodiments of the present application, first introduce several terms related to the embodiments of the present application.
[0096] As shown in FIG. 1, FIG. 1 shows a schematic diagram of the training phase and inference phase of the model. The basic principle of artificial intelligence (AI) is to combine massive data with super strong operation processing capability and intelligent algorithm to establish an AI model for solving a specific problem, so that the AI model can automatically induce and learn potential patterns or features from data, thereby realizing a thinking mode close to human.
[0097] The AI model, that is, the AI algorithm (or AI operator), is a general term for mathematical algorithms constructed based on the principle of artificial intelligence, and is also the basis for solving specific problems using AI. According to different specific methods and / or technologies for realizing artificial intelligence, the AI model can also be referred to as a machine learning model, a deep learning model, or a reinforcement learning model. The following specifically describes machine learning, machine learning model, deep learning, deep learning model, neural network, reinforcement learning, and reinforcement learning model.
[0098] Machine learning is a method for realizing artificial intelligence, and the goal of this method is to design and analyze some algorithms (i.e. models) that allow computers to automatically "learn". The designed algorithm is called a machine learning model.
[0099] Machine learning model is a kind of algorithm that automatically analyzes rules from data and uses rules to predict unknown data. Machine learning models include a variety of models. According to whether the model needs to rely on the label corresponding to the training data during training, the machine learning model can be divided into: 1, supervised learning model; 2, unsupervised learning model.
[0100] 1. Supervised Learning Models: These are models obtained by determining the parameters of an initial AI model based on data from a given training dataset and the labels corresponding to each data point. The process of determining the parameters of the initial AI model using the data and their labels in the training dataset is also called supervised learning (or supervised training). The labels on the data in the training dataset are usually manually labeled to indicate the correct answer for a specific task. Typical supervised learning models include: Support Vector Machines, Neural Network Models, Logistic Regression Models, Decision Trees, Naive Bayes Models, and Gaussian Discriminant Models. Supervised learning models are commonly used for classification or regression.
[0101] 2. Unsupervised Learning Models: These are models obtained by determining the parameters of an initial AI model using unlabeled data from a given training dataset. The process of determining the parameters of the initial AI model using unlabeled training data is also called unsupervised learning (or unsupervised training). Through unsupervised learning, the model can discover meaningful information and correlations in the data, thereby making predictions. There are many types of unsupervised learning models, some of the more commonly used ones being: clustering models, principal component analysis (PCA), anomaly detection models, autoencoders, and generative adversarial networks (GANs).
[0102] Deep learning is a new technological field that emerged during machine learning research. Specifically, deep learning is a method in machine learning based on deep representation learning of data. Deep learning interprets data by building neural networks that simulate the human brain's analytical learning process.
[0103] In the field of AI, deep learning is a learning technique based on deep neural network algorithms. A deep learning model consists of an input layer, hidden layers, and an output layer, and it uses multiple nonlinear transformations to process data.
[0104] In machine learning methods, almost all features need to be determined by industry experts and then encoded. However, deep learning algorithms attempt to learn features from data themselves; algorithms designed based on the principles of deep learning are called deep learning models.
[0105] A typical structure of a current deep learning model is a deep neural network. A neural network is a mathematical or computational model that simulates the structure and function of a biological neural network (the central nervous system of an animal, especially the brain). A neural network performs computation by a large number of neuron connections. A neural network can include multiple neural network layers with different functions, each layer including parameters and computation rules. Different layers in a neural network have different names according to different computation formulas or functions, for example, a layer performing convolution computation is called a convolution layer, which is often used for feature extraction of an input signal (e.g., an image). A neural network can also be composed of multiple sub-neural networks. Neural networks with different structures can be suitable for different scenarios (e.g., classification, recognition) or provide different effects when used for the same scenario. The structure of a neural network can be different in one or more of the following aspects: the number of network layers in the neural network, the order of the network layers, the weights, parameters, or computation formulas in each network layer. There are many different neural networks with high accuracy for different application scenarios such as recognition or classification in the industry. Some neural networks can be trained by a specific data set and then used alone or combined with other neural networks (or other functional modules) to complete a task.
[0106] In other words, a deep learning model is actually a machine learning model with a complex structure of a neural network. According to whether the deep learning model needs to rely on the labels corresponding to the training data during training, the deep learning model can also be divided into a supervised learning model and an unsupervised learning model, which will not be described here. Classical deep learning models include convolutional neural networks (CNN), recurrent neural networks (RNN), recursive neural networks (RNN), etc.
[0107] Reinforcement learning (RL): also known as re-education, evaluation learning or enhancement learning, used to describe and solve the problem of learning a strategy to maximize the reward or achieve a specific goal in the process of interaction between an agent and an environment.
[0108] Reinforcement learning is a learning method in which an agent learns in a trial-and-error manner, and a reward obtained by the agent interacting with an environment guides the behavior, and the goal is to make the agent obtain the maximum reward. Reinforcement learning does not require a training data set. In reinforcement learning, a reinforcement signal (i.e., a reward) provided by the environment evaluates the good or bad of an action, rather than telling the reinforcement learning system how to produce a correct action. Since the information provided by the external environment is little, the agent must learn by itself. In this way, the agent obtains knowledge in the action-evaluation (i.e., reward) environment, and improves the action scheme to adapt to the environment.
[0109] Figure 2 is a schematic diagram of a training process of reinforcement learning. As shown in Figure 2, reinforcement learning mainly includes four elements: an agent, an environment state, an action, and a reward, wherein the input of the agent is the state, and the output of the agent is the action.
[0110] In the current technology, the training process of reinforcement learning is as follows: through multiple interactions between the agent and the environment, the action, the state, and the reward of each interaction are obtained; and the multiple sets of information (action, state, and reward) are used as training data to train the agent once. The above process is used to train the agent in the next round until the convergence condition is met.
[0111] The process of obtaining the action, the state, and the reward of one interaction is shown in Figure 2. The current state s(t) of the environment is input to the agent, and the action a(t) output by the agent is obtained. The reward r(t) of this interaction is calculated according to the relevant performance indicators of the environment under the action a(t). Thus, the action a(t), the state s(t), and the reward r(t) of this interaction are obtained. The action a(t), the state s(t), and the reward r(t) of this interaction are recorded for subsequent training of the agent. The next state s(t+1) of the environment under the action a(t) is also recorded to realize the next interaction between the agent and the environment.
[0112] The agent is an entity that can think and interact with the environment. For example, the agent can be a computer system or a part of a computer system in a certain environment. The agent can perceive the environment according to its own perception, follow the existing instructions or learn autonomously, and communicate and cooperate with other agents to autonomously complete the set goal in the environment. The agent can be a software or a combination of software and hardware entity.
[0113] Any AI model needs to be trained before it is used to solve a specific technical problem. The training of an AI model refers to the process of using a specified initial model to calculate training data, adjusting the parameters in the initial model according to the calculated results, so that the model gradually learns certain rules and has a specific function. The AI model with stable function after training can be used for inference. The inference of an AI model is the process of using a trained AI model to calculate input data and obtain a predicted inference result.
[0114] The most common is supervised training of AI models. For example: most deep learning models are trained based on supervised training.
[0115] The most widely used supervised training method for deep learning models will be introduced below in conjunction with Figure 3.
[0116] Figure 3 shows a schematic block diagram of a deep learning model 100. In the training phase, a training set for the deep learning model needs to be constructed based on the target, and the training set includes multiple training data, each training data is provided with a label, and the label of the training data is the correct answer of the training data on a specific problem. The label can represent the target of training the deep learning model using the training data. For example: for training a deep learning model that can be used to identify different animals, the training set can include multiple images of different animals (i.e. training data), each image can have a label identifying the type of animal contained therein, such as cat, dog, in this example, the type of animal corresponding to each image is the label of the training data.
[0117] When training the deep learning model, the training data can be input to the deep learning model after parameter initialization in batches, and the deep learning model calculates (i.e. infers) the training data to obtain the prediction result for the training data. The prediction result obtained by inference and the label corresponding to the training data are used as data for calculating the loss according to the loss function. The loss function is a function used to calculate the gap (i.e. loss value) between the prediction result of the model for the training data and the label of the training data in the model training phase. The loss function can be implemented using different mathematical functions, and the expressions of commonly used loss functions are: mean square error loss function, logarithmic loss function, least squares method, etc.
[0118] The loss value calculated based on the loss function can be used to update the parameters of the deep learning model, and the gradient descent method is commonly used for parameter updating. The training of the model is a repeated iterative process, and each iteration infers different training data and calculates the loss value. The goal of multiple iterations is to continuously update the parameters of the deep learning model to find the parameter configuration that makes the loss value of the loss function the lowest or tends to be stable.
[0119] In the training phase, in order to make the training efficiency of the model and the performance of the model after training more optimal, some reasonable hyperparameters need to be set for training. The hyperparameters of the deep learning model refer to a class of parameters that cannot be obtained by learning training data or cannot be changed by training data driving in the training process, which is a concept relative to the parameters in the model. The hyperparameters of the deep learning model are usually set by artificial experience or experiment. The hyperparameters include: learning rate, batch size, network structure hyperparameters (such as network layer (also known as depth), network layer interaction, convolution kernel number and convolution kernel size, activation function), etc. Among them, the learning rate as a hyperparameter is used to control the amplitude of the parameter weight update of the model in the training process, which greatly affects the speed and accuracy of the training.
[0120] The trained deep learning model can be used for inference on input data. In the inference phase, the data of the actual application scene is usually used as the input data, and the inference of the trained deep learning model can obtain the inference result. The inference phase is the actual application of the trained deep learning model, which can quickly use the AI capability to solve specific technical problems. Today, there are many AI application scenarios, and the inference of the deep learning model can also be used in various application scenarios, such as personnel identification scene for access control security system, video pornography and violence detection, express delivery order number detection and identification, etc.
[0121] The above only takes the training of the most typical deep learning model as an example for introduction. The training of other types of models has slight differences, but the principle is similar. Most of them are to infer the training data, adjust the parameters in the model according to the inference result, and aim to obtain a stable parameter combination of the model performance.
[0122] Training is mainly divided into supervised training and unsupervised training. The training process of the foregoing deep learning model belongs to supervised training. Taking images as an example, the AI model is trained unsupervisedly. The training images in the training image set do not have annotations, and the training images in the training image set are input into the AI model in turn, and the AI model gradually identifies the association and potential rules between the training images in the training image set until the AI model can be used to judge or identify the type or characteristics of the input image. For example, clustering, after the AI model used for clustering receives a large number of training images, it can learn the characteristics of each training image and the association and difference between the training images, and automatically divide the training images into multiple types. Different task types can use different AI models. Some AI models can only be trained in a supervised learning manner, some AI models can only be trained in an unsupervised learning manner, and some AI models can be trained in both a supervised learning manner and an unsupervised learning manner.
[0123] A language model (LM) is a type of machine learning model that can be used to process and predict natural language data. Language models can include large language models, compact language models, or small language models, among other models for processing and predicting natural language data.
[0124] A large language model (LLM) is a type of neural network model that is trained on a large amount of data and has a large number of parameters. It can understand and generate natural language text. Specifically, a large language model is typically based on neural network technology and is trained on a large amount of text data to learn the grammar, semantics, and contextual information of the language. During the training process, the model continuously optimizes its parameters to improve its ability to understand and generate text. Due to its strong ability to understand natural language, large language models have been widely used in many fields to solve natural language understanding and generation problems. Large language models have a wide range of applications in the field of artificial intelligence, such as natural language processing, machine translation, dialogue systems, etc.
[0125] A compact language model (CLM) is a type of language model that is designed to be more compact and efficient, providing good performance in situations where computational resources are limited. CLM can be achieved through model compression, knowledge distillation, and other technical means to reduce the number of model parameters and computational complexity. Compared to LLM, CLM can achieve comparable or even better performance, especially in scenarios where computational resources are limited. Due to the reduction of computing and storage requirements, CLM is more cost-effective.
[0126] A small language model (SLM) is a type of model with relatively few parameters. They are designed to be efficient and practical, small in size, and easy to adapt to environments with limited resources and computing power, such as mobile devices or embedded systems. Compared to LLM, SLM has a much smaller number of parameters, making it more economical in terms of storage and computing resources. Due to the simplification of parameters and model structure, SLM can often respond faster when processing requests. SLM may be optimized for specific application scenarios or tasks to provide good performance under limited resources. Despite the smaller number of parameters, SLM can still achieve generalization capabilities for a variety of tasks through careful design and training.
[0127] As an example, in a possible implementation, the method provided by the embodiments of the present application can be applied to a cloud service scenario, and the method is executed by a cloud management platform in the cloud service scenario. For ease of description, the cloud service scenario is described in detail below in combination with FIG. 4.
[0128] FIG. 4 shows a schematic block diagram of a cloud scenario applicable to the embodiments of the present application. As shown in FIG. 4, the cloud scenario can include a cloud management platform 410, the Internet 420, and a client 430.
[0129] As shown in FIG. 4, the cloud management platform 410 is configured to manage an infrastructure that provides a plurality of cloud services. The infrastructure includes a plurality of cloud data centers, each of which includes a plurality of servers, and each of the servers includes cloud service resources, which provide corresponding cloud services for tenants.
[0130] The cloud management platform 410 can be located in a cloud data center, which can provide an access interface (such as an interface or an application program interface (API)). A tenant can remotely access the access interface to register a cloud account and a password at the cloud management platform 410 by using the client 430, and log in to the cloud management platform 410. After the cloud management platform 410 authenticates the cloud account and the password successfully, the tenant can further select and purchase a virtual machine of a specific specification (processor, memory, and disk) at the cloud management platform 410 by paying a fee. After the purchase is successful, the cloud management platform 410 provides a remote login account and password of the purchased virtual machine, and the client 430 can remotely log in to the virtual machine, and install and run an application of the tenant in the virtual machine. Therefore, the tenant can create, manage, log in to, and operate a virtual machine in the cloud data center through the cloud management platform 410. The virtual machine can also be referred to as an elastic compute service (ECS) or an elastic instance (different cloud service providers have different names).
[0131] It should be understood that the tenant of the cloud service can be an individual, an enterprise, a school, a hospital, an administrative organ, or the like.
[0132] The functions of the cloud management platform 410 include, but are not limited to, a user console, a computing management service, a network management service, a storage management service, an authentication service, and an image management service. The user console provides an interface or an API to interact with the tenant, the computing management service is configured to manage servers running virtual machines and containers and bare metal servers, the network management service is configured to manage network services (such as gateways and firewalls), the storage management service is configured to manage storage services (such as data bucket services), the authentication service is configured to manage the account and password of the tenant, and the image management service is configured to manage virtual machine images. The tenant uses the client 430 to log in to the cloud management platform 410 through the Internet 420, and manage the rented cloud services.
[0133] As an example, the cloud service can include, but is not limited to, an AI service. AI services and products in the cloud field not only embody the characteristics of on-demand use and purchase of cloud services, but also have the characteristics of abstraction, diversity, and wide application of AI technology. AI services in the cloud field include AI basic development platform services of the platform-as-a-service (PaaS) type.
[0134] It should be understood that the AI basic development platform service is a PaaS cloud service in the cloud management platform 410, and is a software platform that assists users (also referred to as tenants, AI developers, etc.) in building, training, and deploying AI models and developing and deploying AI applications based on a large number of underlying resources and software capabilities owned by a public cloud service provider. That is, the public cloud service provider provides an AI basic development platform to the tenants by virtue of sufficient underlying resources and upper-layer AI algorithm capabilities. The AI basic development platform has an AI development framework and various AI algorithms built in, which can be used by the tenants to quickly build and develop AI models or AI applications that meet individual needs on the AI basic development platform.
[0135] As shown in FIG. 5, FIG. 5 shows a schematic diagram of interaction between a tenant and an AI basic development platform suitable for embodiments of the present application. The interaction form of the tenant and the AI basic development platform mainly includes that the tenant logs in to the cloud management platform 410 through the web page of the client 430, selects and purchases the cloud service of the AI basic development platform in the cloud management platform 410, and after the purchase, the tenant can perform the whole-process AI development based on the functions provided by the AI basic development platform.
[0136] As an example, when a tenant develops and trains his own AI model on the AI base development platform, it is based on the base resources (mainly computing resources such as central processing units (CPUs), graphics processing units (GPUs), neural network processing units (NPUs), etc.) in the data center of the cloud service provider. Therefore, when purchasing and using the AI base development platform, the main payment is for the resources used. For example, before using the AI base development platform, the tenant needs to make a prepayment, and when prepaying, different specifications of resources can support different functions of the AI base development platform. The tenant mainly selects the name and specification of the resource according to the function of the AI base development platform that needs to be used, and the tenant can also choose the purchase duration, and the cloud management platform 410 performs pricing of the package according to the resource name, specification and purchase duration selected by the user. When the tenant purchases the prepayment package, he can use the capabilities provided by the AI base development platform and the base computing resources included in the prepayment package to build, train and deploy AI models, etc. When the resource usage exceeds the quota of the current prepayment package, the cloud management platform 410 charges for the excess resources in a pay-as-you-go manner. In fact, the base resources used by the tenant on the AI base development platform are mainly virtualized computing resources such as virtual machines and containers.
[0137] It should be understood that since the sale of the AI base development platform is actually a form of selling software capabilities together with virtualized hardware base resources, and the base resources supporting any one process in the AI base development platform may be distributed on different physical devices, that is, the hardware devices actually executing a process are usually a server cluster in the same data center, or a server cluster distributed in different data centers.
[0138] Language models (such as LLMs, SLMs, CLMs, etc.) excel in various text understanding and generation tasks through text as a medium. However, relying solely on naked models still has obvious capability boundaries, such as being unable to complete user needs such as booking tickets and hotels. By providing an external tool interface to the language model and having the language model learn to input the corresponding call statement to call the external tool, the capability boundaries of the language model can be greatly expanded.
[0139] Specifically, the method of having the language model learn to call external tools can include:
[0140] ①, put information such as how to use the tool into the model input through prompt engineering, and learn to output tool call statements based on the basic capabilities of the model itself.
[0141] Prompt engineering can make large language models perform better by finding more suitable prompts. For example, an iterative process of Thought, Action, and Observation can be introduced to enable language models to call external tools without training. This method adds the knowledge required to call tools (such as tool definitions, format requirements for calling statements, etc.) and corresponding examples to the prompt, enabling the language model to call external tools when given a user demand statement.
[0142] However, this method does not require further training of the language model and relies heavily on the understanding and reasoning ability of the language model itself, making it difficult for language models with weak basic abilities to effectively learn to call external tools. At the same time, due to the diversity and complexity of real external tools and user needs, language models based on this method are not as accurate in calling as desired.
[0143] ②. By constructing corresponding data, the model learns the corresponding knowledge and action methods in a continuous training manner.
[0144] For example, based on the collected real API and constructed examples, new tool calling demand statements can be generated, and LLM can be used to label the solution path. This method can construct a large amount of single-round demand data, which can be used to further train the language model, and the trained language model can call external tools.
[0145] However, the data constructed by this method is all single-round tool calling data, lacking interaction with the user. Language models trained based on this data will be affected in general ability and cannot maintain the original general ability of the language model and other abilities required for user interaction.
[0146] For another example, multi-agent can be introduced to generate tool calling data through interaction between multi-agent. LLM can be used as a user agent, an assistant agent, and a tool agent to form tool calling dialogue data through multi-agent interaction. Different agents can use different system prompts.
[0147] Exemplarily, the agent can use LLM as the core, including a memory module, a tool module, a planning module, and an action module, etc. The memory module is used to realize long-term memory and / or short-term memory function; the tool module includes multiple callable external tools; the planning module includes multiple planning algorithms, for example, the agent can plan the external input task according to the content in the memory; the action module supports the agent to make actions according to the planning results, such as calling tools, etc.
[0148] As shown in FIG. 6, which illustrates a schematic diagram of an agent. In an LLM-backed autonomous agent system, the LLM acts as the brain of the agent (or agent), and is supplemented by several key components:
[0149] 1) Planning, including but not limited to:
[0150] Subgoal decomposition: Agents break down large tasks into smaller, manageable subgoals, enabling them to effectively handle complex tasks. For example, through a chain of thoughts (CoT) instructing the model to “think step by step,” the model can leverage more test time computation to break down difficult tasks into smaller, simpler steps. CoT transforms large tasks into multiple manageable tasks and elucidates the explanation of the model’s thought process.
[0151] Reflection and refinement: Agents can self-criticize and reflect on past actions, learn from mistakes, and refine future steps to improve the quality of the final outcome.
[0152] 2) Memory, including but not limited to:
[0153] Short-term memory: Utilizing the model’s short-term memory to learn.
[0154] Long-term memory: Providing agents with the ability to retain and recall (indefinitely) information over long periods of time, typically through the use of external vector storage and fast retrieval.
[0155] 3) Tools, including but not limited to:
[0156] Agents learn to call external application programming interfaces to obtain additional information missing from the model’s weights (typically difficult to change after pre-training), including current information, code execution capabilities, access to proprietary information sources, etc.
[0157] 4) Action, the model performs specific tasks and records the results.
[0158] The role of prompt is mainly to give AI model prompt input information context and input model parameter information. When training supervised learning or unsupervised learning model, prompt can help model better understand the intention of input and make corresponding response. In addition, prompt can also improve the explainability and accessibility of model.
[0159] In simple terms, a prompt is a "hint" or "guide" given to an AI model to help it better understand and complete a task.
[0160] For example, a prompt is not just a user input question or query, but also contains instructions, external information (context), output indicators, and other parts. Among them, the user input or query: usually the query instruction input by the user (i.e. the prompter) into the system, which tells the model what to do, how to use external information, and how to process the query and build the output; external information (context): serves as an additional source of knowledge for the model. These can be manually inserted into the prompt, obtained through vector database retrieval (retrieval augmentation), or introduced through other means (such as API, computation, etc.); output indicator: marks the beginning of the text to be generated.
[0161] Let LLM play the roles of user agent, helper agent, and tool agent. Through the interaction between multiple agents, the process of tool invocation dialogue data is formed around a given tool set. The user agent proposes requirements; the helper agent analyzes the requirements and calls the corresponding tools to solve the user's problems; the tool agent simulates the real tool, takes the tool invocation statement generated by the helper agent as input, and outputs the corresponding tool invocation result. This method does not need to call real tools to generate corresponding tool invocation dialogue data (such as user agent requirements, helper agent generated tool invocation statements, and corresponding tool invocation results). The generated tool invocation dialogue data can be used to continue training language models, which is beneficial to the flexibility of data construction.
[0162] However, the data constructed by this method still lacks authenticity and fluency, and it is difficult to simulate the real interaction between users. At the same time, there is a lack of control over the entire dialogue process, making it difficult to stably generate data that improves the desired capabilities. Moreover, this method is also limited to tool invocation scenarios, and language models trained based on this data will be affected in terms of general capabilities, and will not be able to maintain the original general capabilities and other capabilities required for user interaction.
[0163] Therefore, the embodiments of the present application provide a method for generating data, a method and device for processing requirements, which can generate data that can be used for language model training based on the method, and the capabilities of language models trained based on the data can be improved in the direction of desired capabilities.
[0164] It should be understood that the embodiments shown below do not particularly limit the specific structure of the subject performing the method provided by the embodiments of the present application, as long as the subject can communicate according to the method provided by the embodiments of the present application by running a program in which the code of the method provided by the embodiments of the present application is recorded. For example, the subject performing the method provided by the embodiments of the present application can be a device, and in the case of no special description, the "device" in the present application can refer to the device itself (for example, an access network device, a terminal device, or a core network device, etc.), a component in the device (for example, a processor, a chip, or a chip system, etc.), or a logic module or software capable of realizing all or part of the functions of the device.
[0165] Next, a method for generating data provided by the embodiments of the present application is described in detail in combination with FIG. 7. FIG. 7 shows a schematic diagram of a method 700 for generating data provided by the embodiments of the present application. The method 700 can include:
[0166] S710, generating a dialog flow according to the first requirement, the dialog flow including a dialog process meeting the first requirement.
[0167] Before generating the data, a dialog flow can be generated according to the first requirement, the dialog flow including a dialog process. The dialog process of the generated dialog flow needs to meet the first requirement. The dialog process can be composed of multiple dialogues, and the conversion between the multiple dialogues is reasonable and natural.
[0168] The first requirement can be a requirement without calling a tool (for example, the first requirement can be a casual chat or a non-tool calling requirement, etc.), and the first requirement can also be a requirement requiring calling a tool (for example, the first requirement can be a tool calling requirement, etc.). When the first requirement is a requirement requiring calling a tool, generating a dialog flow according to the first requirement can include generating a dialog flow according to the first requirement and a tool list. The tool list can be composed of tools in a tool set (for example, the tool list can be a table or a list composed of tools in the tool set), for indicating tools that can be called or simulated (simulating calling means simulating the operation of actual calling, but not actually calling; the result obtained by simulating calling represents the result obtained by simulating actual calling, and is not the result produced by actual calling). Generating a dialog flow according to the tool list can avoid involving tools that cannot be called or simulated in the generated dialog flow, which will lead to the difficulty of implementing the generated dialog flow.
[0169] The method 700 can be performed by multiple agents, and for example, the above step S710 can be performed by a dialog planning agent. For example, the dialog planning agent can generate a dialog flow according to a given first requirement.
[0170] The dialogue flow can control the content of the dialogue, and is generated by the dialogue planning agent for a natural language string. The dialogue flow can be used as part of the input to the user agent, so that the user agent generates a new requirement according to the planning and the dialogue history.
[0171] S720, generating first requirement data according to the dialogue flow.
[0172] After generating the dialogue flow (or after obtaining the generated dialogue flow), the first requirement data can be generated according to the dialogue flow.
[0173] For example, the first requirement can be a basic requirement, and the first requirement data can be used to represent a certain sub-requirement belonging to the basic requirement. The plurality of sub-requirements meet the dialogue flow, and the conversion between the plurality of sub-requirements is reasonable and natural.
[0174] For example, the first requirement is "booking a train ticket", and the requirement of "booking a train ticket" can include a plurality of sub-requirements, such as "booking a train ticket for the trip" and "booking a train ticket for the return trip".
[0175] The scenario of needing "booking a train ticket" can include traveling, business trip, learning or visiting relatives, etc. Before leaving for the destination, not only can the round-trip train ticket be booked, but also the weather of the destination can be inquired, the local customs can be understood, or the subsequent dialogue can be started through chatting. Therefore, further, the first requirement can also include sub-requirements such as "inquiry of destination weather", "tourist suggestion" and "chitchat".
[0176] For example, the first requirement is "shopping recommendation", and the requirement of "shopping recommendation" can include a plurality of sub-requirements, such as "female shoe recommendation", "male shoe recommendation" and "jewelry recommendation".
[0177] The scenario of needing "shopping recommendation" can include purchasing items for a specific occasion. Therefore, further, the first requirement can also include sub-requirements such as "award ceremony dress suggestion" and "chitchat".
[0178] For example, the user agent can generate the first requirement data, and the first requirement data can be used to represent the specific requirement of the user agent.
[0179] S730, generating first response data according to the first requirement data, and the first requirement data and the first response data are used to train the language model.
[0180] The first response data can represent a response to the first requirement data. The first requirement data meeting the dialogue flow and the first response data representing the response to the first requirement data can constitute training data, and the training data can be used to train the language model.
[0181] Exemplarily, the step S730 can be performed by an assistant agent. For example, the assistant agent can generate the first response data.
[0182] Exemplarily, the first response data can be data including the result of calling the tool, and the first response data can also be data not including the result of calling the tool, which replies to the statement of the user agent.
[0183] Based on the scheme provided in the embodiments of the present application, by generating the first demand data according to the dialogue flow, and generating the first response data responding to the first demand data, since the dialogue flow generated according to the first demand has a limiting effect on generating the first demand data, on the one hand, training data (i.e. the first demand data and the first response data) capable of improving the ability of the trained language model in the direction of the first demand can be obtained, compared with the process of generating demand data without the limitation of dialogue flow, the first demand data can improve the ability of the trained language model in the expected ability direction by setting the first demand according to the expected ability; on the other hand, the dialogue flow included in the dialogue flow can improve the controllability and fluency of the process of generating data, which helps to optimize the process of generating data.
[0184] In some possible implementation ways, the first response data includes tool calling statement data. The step S730 can include:
[0185] S730a, generating a plurality of tool calling statements according to the first demand data.
[0186] S730b, determining the tool calling statement data according to the consistency of the plurality of tool calling statements.
[0187] The format of the tool calling statement and the non-tool calling statement (for example, dialogue reply) generated according to the first demand is different. After generating the data responding to the first demand according to the first demand, it can be judged according to the format of the generated data responding to the first demand whether it is a tool calling (if it is judged according to the format of the generated data responding to the first demand that the generated data responding to the first demand is a tool calling statement, it is considered to be a tool calling; if it is judged according to the format of the generated data responding to the first demand that the generated data responding to the first demand is a non-tool calling statement, it is considered to be not a tool calling), whether it is a tool calling or not, the consistency of the plurality of data responding to the first demand (i.e. the plurality of tool calling statements or the plurality of dialogue replies) can be further judged.
[0188] Exemplarily, the steps S730a and S730b can be performed by an assistant agent, and the consistency of the plurality of tool calling statements can be determined by a verification agent.
[0189] For example, taking the existence of assistant agent #1 and verification agent #1 as an example, assistant agent #1 can receive the first request data generated by the user agent. Based on the first request data, assistant agent #1 can generate multiple tool call statements. Verification agent #1 can receive the multiple tool call statements generated by assistant agent #1 and determine the consistency of the multiple tool call statements. After determining the tool call statement used to respond to the first request based on the consistency of the multiple tool call statements, assistant agent #1 can generate tool call statement data based on the tool call statement used to respond to the first request.
[0190] It should be understood that the embodiments of this application do not limit the order in which an assistant agent generates multiple data responses to the first requirement. An assistant agent can generate multiple data responses to the first requirement simultaneously or sequentially.
[0191] For example, step S730a can be performed by an assistant agent and a verification agent, and step S730b can be performed by an assistant agent. The consistency of multiple tool call statements can be determined by the verification agent.
[0192] For example, taking the existence of assistant agent #1 and verification agent #1 as an example, assistant agent #1 and verification agent #1 can receive first request data generated by the user agent. Based on the first request data, assistant agent #1 can generate one or more first tool invocation statements, and verification agent #1 can generate one or more second tool invocation statements. Verification agent #1 can receive the first tool invocation statements generated by assistant agent #1 and determine the consistency between the first tool invocation statements and the second tool invocation statements. After determining the tool invocation statement used to respond to the first request based on the consistency between the first tool invocation statements and the second tool invocation statements, assistant agent #1 can generate tool invocation statement data based on the tool invocation statement used to respond to the first request.
[0193] It should be understood that the embodiments of this application do not limit the order in which the verification agent and the assistant agent generate the data in response to the first requirement. The verification agent and the assistant agent can generate the data in response to the first requirement simultaneously or sequentially.
[0194] For example, step S730a can be executed by multiple assistant agents, and step S730b can be executed by one assistant agent. The consistency of multiple tool call statements can be determined by the verification agent.
[0195] For example, taking the existence of an agent #1, an agent #2 and an agent #1 as an example, the agent #1 and the agent #2 can receive the first requirement data generated by the user agent. According to the first requirement data, the agent #1 and the agent #2 can generate a plurality of first tool invocation statements. The agent #1 can receive the plurality of first tool invocation statements generated by the agent #1 and the agent #2, and determine the consistency of the plurality of first tool invocation statements. After determining the tool invocation statement for responding to the first requirement according to the consistency of the plurality of first tool invocation statements, the agent #1 or the agent #2 can generate tool invocation statement data according to the tool invocation statement for responding to the first requirement.
[0196] It should be understood that when there are a plurality of agents, the plurality of agents can generate a plurality of first tool invocation statements, the agent #1 can receive the plurality of first tool invocation statements, and determine the consistency of the plurality of first tool invocation statements. After determining the tool invocation statement for responding to the first requirement according to the consistency of the plurality of first tool invocation statements, an optional one of the plurality of agents or a specified one of the plurality of agents can be selected, and the optional one of the plurality of agents or the specified one of the plurality of agents can generate tool invocation statement data according to the tool invocation statement for responding to the first requirement.
[0197] It should be understood that the present application does not limit the order in which the plurality of agents generates data for responding to the first requirement, and the plurality of agents can generate data for responding to the first requirement simultaneously or sequentially.
[0198] For example, the above step S730a can be performed by a plurality of agents and an agent, S730b can be performed by an agent, and the consistency of a plurality of tool invocation statements can be determined by an agent.
[0199] For example, taking the existence of an agent #1, an agent #2 and an agent #1 as an example, the agent #1, the agent #2 and the agent #1 can receive the first requirement data generated by the user agent. According to the first requirement data, the agent #1 and the agent #2 can generate a plurality of first tool calling statements, and the agent #1 can generate one or more second tool calling statements. The agent #1 can receive the first tool calling statements generated by the agent #1 and the agent #2, and determine the consistency of the first tool calling statements and the second tool calling statements. After determining the tool calling statements for responding to the first requirement according to the consistency of the first tool calling statements and the second tool calling statements, the agent #1 or the agent #2 can generate tool calling statement data according to the tool calling statements for responding to the first requirement.
[0200] It should be understood that when there are a plurality of agents, the plurality of agents can generate a plurality of first tool calling statements, the agent #1 can receive the plurality of first tool calling statements, and determine the consistency of the plurality of first tool calling statements and the second tool calling statements. After determining the tool calling statements for responding to the first requirement according to the consistency of the plurality of first tool calling statements and the second tool calling statements, an optional one of the plurality of agents or a specified one of the plurality of agents can be selected, and the optional one of the plurality of agents or the specified one of the plurality of agents can generate tool calling statement data according to the tool calling statements for responding to the first requirement.
[0201] It should be understood that the embodiments of the present application do not limit the order of the data generated by the agent for responding to the first requirement and the data generated by the plurality of agents for responding to the first requirement, and the plurality of agents and the agent can generate the data for responding to the first requirement simultaneously or sequentially.
[0202] When the first requirement is a requirement for calling a tool, a plurality of tool calling statements can be generated. The tool calling statements can be used to call or simulate calling the tool calling statements to indicate calling the tool, wherein the calling tool can be real or simulated. The embodiments of the present application do not limit this.
[0203] Different agents can use different system prompts. Enhancing the compliance of the dialogue flow in the form of prompts can ensure that the dialogue between the plurality of agents is in accordance with the dialogue flow.
[0204] Based on the scheme provided in the embodiments of the present application, by generating a plurality of tool calling statements according to the first requirement data, and then determining tool calling statement data including tool calling statements meeting the consistency requirement, on the one hand, training data capable of expanding the capability boundary of the trained language model can be obtained, and the trained language model can be enabled to call external tools; on the other hand, according to the consistency of the plurality of tool calling statements, appropriate tool calling statement data can be determined, and the correctness of the tool calling statement data can be improved.
[0205] In some possible implementation manners, when the tool names and / or parameter arguments indicated by at least two tool calling statements in the plurality of tool calling statements are the same, the tool calling statement data includes any one of the at least two tool calling statements; or when the tool names and / or parameter arguments indicated by at least two tool calling statements in the plurality of tool calling statements are not the same, and the sum of the number of the plurality of tool calling statements meets a first threshold, the tool calling statement data includes any one of the plurality of tool calling statements and a first mask mark, and the first mask mark is used to indicate that the any one tool calling statement is incorrect.
[0206] In the context of programming and software development, "parameter argument" (or simply "argument") refers to the values or data passed to a function, method, or procedure. These values are received when the function is called and are used or processed inside the function. Parameter arguments are part of the function definition, specifying the data types, number of arguments, and how they are used by the function. "Tool name" can refer to a function library, class library, framework, or other tools that can be called by code to perform specific tasks. Parameter arguments and tool names are two different concepts that often appear together in programming and software development. When calling a function or method, we need to specify the corresponding parameter arguments, which may be passed directly through code or through an interface provided by a certain tool or framework.
[0207] For example, when using a certain library or framework, we may need to specify which specific tool to use by passing the library name or framework name as a parameter. In addition, when setting parameters in configuration files or environment variables, the specification of tool names is often involved.
[0208] For example, the verification agent#1 can take the first requirement data and the plurality of tool calling statements as inputs, perform consistency voting on the plurality of tool calling statements, and determine the consistency of the plurality of tool calling statements.
[0209] For example, a plurality of tool call statements can be generated first, and the tool name and / or parameter argument of the plurality of tool call statements can be determined. If the tool name and / or parameter argument of at least two tool call statements are the same, one of the at least two tool call statements can be selected as the most reasonable tool call statement, and the tool call statement data includes the most reasonable tool call statement. Alternatively, a first tool call statement can be generated first, and the tool name and / or parameter argument of the first tool call statement can be determined. Then, a second tool call statement can be generated, and the tool name and / or parameter argument of the second tool call statement can be determined. If the tool name and / or parameter argument of the first and second tool call statements are the same, it is considered that the first and second tool call statements are consistent and meet the consistency, and one of the at least two tool call statements can be selected as the most reasonable tool call statement, and the tool call statement data includes the most reasonable tool call statement. Alternatively, if the tool name and / or parameter argument of the first and second tool call statements are not the same, a third tool call statement can be generated, and the tool name and / or parameter argument of the third tool call statement can be determined. The tool name and / or parameter argument of the first, second, and third tool call statements are compared until at least two tool call statements with the same tool name and / or parameter argument are determined, and one of the at least two tool call statements with the same tool name and / or parameter argument is selected as the most reasonable tool call statement. For example, when the plurality of tool call statements cannot determine the most reasonable tool call statement, a new tool call statement can be generated and consistency can be determined.
[0210] Further, since the tool name and / or parameter argument of the at least two tool call statements are the same, it can be considered that the at least two tool call statements have a higher probability of successfully calling / simulating calling the tool required to process the first demand, and it can be considered that the consistency of the plurality of tool call statements can improve the accuracy of the tool call statement.
[0211] When the number of generated tool call statements reaches the first threshold, and there is still no tool call statement with at least two consistent tool names and / or parameter arguments, one of the plurality of generated tool call statements can be selected and marked with a first mask, and the selected tool call statement and the first mask are determined as the tool call statement data. Subsequently, if the selected tool call statement is used to train the language model, the training data including the tool call statement can be skipped according to the first mask, so as to ensure that the training data used for language model learning are all reasonable data verified.
[0212] The tool invocation statement data can be used to invoke / simulate invocation of a tool, and the result of the invocation / simulated invocation of the tool can be generated according to the result of the invocation / simulated invocation of the tool.
[0213] For example, the tool agent can generate the result of the invocation / simulated invocation of the tool according to the tool invocation statement data, the assistant agent can summarize the result of the invocation of the tool generated by the tool agent, and reply to the user agent based on the summarized result. The tool invocation statement data, the result of the invocation / simulated invocation of the tool, and the data of the reply of the assistant agent to the user agent after summarizing the result of the invocation of the tool can all be used to train the language model.
[0214] Based on the scheme provided in the embodiments of the present application, the tool invocation statement data is determined based on the tool name and / or the parameter input of the plurality of tool invocation statements, and / or whether the number of the plurality of tool invocation statements meets the first threshold. On the one hand, it can be ensured that the training data generated for the language model learning is reasonable data that has been verified, which improves the correctness of the training data, helps the language model to learn correct knowledge in the training process, and reduces the illusion of the language model obtained by continuing to train based on the training data. On the other hand, it can avoid generating tool invocation statements infinitely.
[0215] In some possible implementation manners, the first response data includes dialogue reply data. S730 can include:
[0216] S730d, generating a plurality of dialogue replies according to the first demand data.
[0217] S730e, determining dialogue reply data according to the consistency of the plurality of dialogue replies.
[0218] When it is determined that the format of the data (e.g., the plurality of dialogue replies) responding to the first demand generated according to the first demand meets the format of non-tool invocation, the dialogue reply data can be determined according to the consistency of the dialogue replies.
[0219] For example, when the first demand is a demand that does not require invocation of a tool, dialogue replies can be generated, and the dialogue reply data can be determined according to the consistency (e.g., content consistency) of the dialogue replies.
[0220] For example, the steps S730d and S730e described above can be performed by an assistant agent, and the consistency of the plurality of dialogue replies can be determined by a verification agent.
[0221] For example, taking the existence of assistant agent #1 and verification agent #1 as an example, assistant agent #1 can receive the first request data generated by the user agent. Based on the first request data, assistant agent #1 can generate multiple dialogue responses. Verification agent #1 can receive the multiple dialogue responses generated by assistant agent #1 and determine the consistency of the multiple dialogue responses. After determining the dialogue response used to respond to the first request based on the consistency of the multiple dialogue responses, assistant agent #1 can generate dialogue response data based on the dialogue response used to respond to the first request.
[0222] It should be understood that the embodiments of this application do not limit the order in which an assistant agent generates multiple data responses to the first requirement. An assistant agent can generate multiple data responses to the first requirement simultaneously or sequentially.
[0223] For example, step S730d can be executed by an assistant agent and a verification agent, S730e can be executed by the assistant agent, and the consistency of multiple tool call statements can be determined by the verification agent.
[0224] For example, assuming there are assistant agent #1 and verification agent #1, both can receive first request data generated by the user agent. Based on the first request data, assistant agent #1 can generate one or more first dialogue responses, and verification agent #1 can generate one or more second dialogue responses. Verification agent #1 can receive the first dialogue responses generated by assistant agent #1 and determine the consistency between the first and second dialogue responses. After determining the dialogue response used to respond to the first request based on the consistency between the first and second dialogue responses, assistant agent #1 can generate dialogue response data based on this dialogue response used to respond to the first request.
[0225] It should be understood that the embodiments of this application do not limit the order in which the verification agent and the assistant agent generate the data in response to the first requirement. The verification agent and the assistant agent can generate the data in response to the first requirement simultaneously or sequentially.
[0226] For example, step S730d can be performed by multiple assistant agents, and step S730e can be performed by one assistant agent. The consistency of multiple dialogue responses can be determined by the verification agent.
[0227] For example, taking the existence of an agent #1, an agent #2 and an agent #1 as an example, the agent #1 and the agent #2 can receive the first demand data generated by the user agent. According to the first demand data, the agent #1 and the agent #2 can generate a plurality of first dialogue replies. The agent #1 can receive the plurality of first dialogue replies generated by the agent #1 and the agent #2, and determine the consistency of the plurality of first dialogue replies. After determining the dialogue reply for responding to the first demand according to the consistency of the plurality of first dialogue replies, the agent #1 or the agent #2 can generate dialogue reply data according to the dialogue reply for responding to the first demand.
[0228] It should be understood that when there are a plurality of agents, the plurality of agents can generate a plurality of first dialogue replies, the agent #1 can receive the plurality of first dialogue replies, and determine the consistency of the plurality of first dialogue replies. After determining the dialogue reply for responding to the first demand according to the consistency of the plurality of first dialogue replies, an optional one of the plurality of agents or a specified one of the plurality of agents can be selected, and the optional one of the plurality of agents or the specified one of the plurality of agents can generate dialogue reply data according to the dialogue reply for responding to the first demand.
[0229] It should be understood that the present application does not limit the order in which the plurality of agents generates data for responding to the first demand, and the plurality of agents can generate data for responding to the first demand simultaneously or sequentially.
[0230] For example, the above step S730d can be performed by a plurality of agents and an agent, and S730e can be performed by an agent, and the consistency of the plurality of dialogue replies can be determined by the agent.
[0231] For example, taking the existence of an agent #1, an agent #2 and an agent #1 as an example, the agent #1, the agent #2 and the agent #1 can receive the first demand data generated by the user agent. According to the first demand data, the agent #1 and the agent #2 can generate a plurality of first dialogue replies, and the agent #1 can generate one or more second dialogue replies. The agent #1 can receive the first dialogue replies generated by the agent #1 and the agent #2, and determine the consistency of the first dialogue replies and the second dialogue replies. After determining the dialogue reply for responding to the first demand according to the consistency of the first dialogue replies and the second dialogue replies, the agent #1 or the agent #2 can generate dialogue reply data according to the dialogue reply for responding to the first demand.
[0232] It should be understood that when there are a plurality of agents, the plurality of agents can generate a plurality of first dialogue replies, the agent #1 can receive the plurality of first dialogue replies, and determine the consistency of the plurality of first dialogue replies and the second dialogue replies. After determining the dialogue reply for responding to the first demand according to the consistency of the plurality of first dialogue replies and the second dialogue replies, an optional one of the plurality of agents or a specified one of the plurality of agents can be selected, and the optional one of the plurality of agents or the specified one of the plurality of agents can generate dialogue reply data according to the dialogue reply for responding to the first demand.
[0233] It should be understood that the embodiments of the present application do not limit the order of the agent generating data for responding to the first demand and the plurality of agents generating data for responding to the first demand, and the plurality of agents and the agent can generate data for responding to the first demand at the same time or sequentially.
[0234] Based on the scheme provided by the embodiments of the present application, on the one hand, by generating a plurality of dialogue replies according to the first demand data, and then determining the first response data including dialogue reply data, compared with generating training data based only on tool calling scenarios, when the first demand data indicates a demand that does not need to call a tool, training data that can maintain the original general ability of the trained language model can be obtained, and the diversity and authenticity of the dialogue are improved. On the other hand, according to the consistency of the plurality of dialogue replies, appropriate dialogue reply data can be determined, and the correctness of the dialogue reply data is improved.
[0235] In some possible implementations, when the content of at least two of the multiple dialogue replies satisfies the content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or when the content of at least two of the multiple dialogue replies does not satisfy the content consistency, and the sum of the number of the multiple dialogue replies satisfies the second threshold, the dialogue reply data comprises any one of the multiple dialogue replies and a second mask mark, and the second mask mark is used to indicate that the any one dialogue reply is incorrect.
[0236] For example, the verification agent #1 can take the first requirement data and the multiple dialogue replies as input, perform consistency voting on the multiple dialogue replies, and determine the consistency of the multiple dialogue replies.
[0237] For example, after the data responding to the first requirement is generated, it can be determined whether the generated data responding to the first requirement satisfies the non-tool call (because the assistant agent generates the first response data, it cannot be known in advance whether the current dialogue is a tool call or a non-tool call, and it needs to be determined according to the format of the dialogue reply). The multiple data responding to the first requirement can be generated first, when at least two of the generated data responding to the first requirement satisfy the non-tool call, it can be considered that the data responding to the first requirement satisfying the non-tool call is the dialogue reply, the content of the multiple dialogue replies is determined, and the content consistency of the multiple dialogue replies is determined; or, the first dialogue reply can be determined according to the format of the first generated data responding to the first requirement, the content of the first dialogue reply is determined, then the second dialogue reply and its content are determined according to the format of the second generated data responding to the first requirement, and the content consistency of the two dialogue replies is compared.
[0238] For example, after determining that two dialog replies satisfying the non-tool call are generated (hereinafter distinguished as dialog reply #A and dialog reply #B), the content of the dialog reply #A and the dialog reply #B can be determined, and the content consistency of the dialog reply #A and the dialog reply #B can be compared. When the content of the dialog reply #A and the dialog reply #B is consistent, either of the dialog reply #A and the dialog reply #B can be selected as the most reasonable dialog reply. When the content of the dialog reply #A and the dialog reply #B is inconsistent, a third dialog reply, i.e., dialog reply #C, can be generated, the content of the dialog reply #C can be determined, and the content consistency of the dialog reply #A, the dialog reply #B and the dialog reply #C can be compared. When the content of two of the dialog reply #A, the dialog reply #B and the dialog reply #C is consistent, for example, the content of the dialog reply #A and the dialog reply #C is consistent, either of the dialog reply #A and the dialog reply #C can be selected as the most reasonable dialog reply. When the content of the dialog reply #A, the dialog reply #B and the dialog reply #C is inconsistent, the dialog reply #D can be continuously generated, the content of the dialog reply #D can be determined, and the content consistency can be compared. When at least two dialog replies with consistent content cannot be determined from the dialog reply #A, the dialog reply #B, the dialog reply #C, the dialog reply #D, … and the dialog reply #X, the new dialog reply can be continuously generated and the content consistency can be compared until at least two dialog replies with consistent content can be determined, or until the number of the generated dialog replies is greater than or equal to the second threshold value.
[0239] It can be understood that the above content consistency is not limited to the complete consistency of the content. When the content of two dialog replies is generally consistent, expresses similar meaning or can achieve the same result, etc., it can be considered that the content is consistent. The above at least two dialog replies with consistent content satisfy the consistency requirement.
[0240] When the number of the generated dialog replies is greater than or equal to the second threshold value, and there is still no dialog reply satisfying the above content consistency requirement, one of the generated dialog replies can be selected and the second mask mark can be performed, the selected dialog reply and the second mask mark are determined as the dialog reply data. If the selected dialog reply is used for training the language model in the future, the training data can be skipped according to the second mask mark, so as to ensure that the training data used for the language model learning is all reasonable data verified.
[0241] Based on the scheme provided in the embodiments of the present application, the dialogue reply data is determined by whether the content of the multiple dialogue replies and / or the number of the multiple dialogue replies satisfies the second threshold. On the one hand, it can be ensured that the generated training data for language model learning is all reasonable data that has been verified, the correctness of the training data is improved, which helps the language model to learn correct knowledge in the training process and reduces the illusion of the language model obtained by continuing to train according to the training data; on the other hand, it can avoid generating dialogue replies infinitely.
[0242] In some possible implementation manners, the dialogue flow indicates a reply manner of the dialogue reply; when the first response data includes the dialogue reply data, the first response data satisfies the reply manner; or when the first response data includes the tool calling statement data, the tool calling result data satisfies the reply manner, and the tool calling result data is determined according to the tool calling statement data.
[0243] In the method 700, the generated first demand data needs to meet the constraint of the dialogue flow. In the step S730, the generated first response data and / or the tool calling result data determined according to the tool calling statement data in the first response data can also need to meet the constraint of the dialogue flow.
[0244] For example, in addition to the dialogue flow, the dialogue flow can also include / indicate a reply manner of the dialogue reply (for example, the dialogue flow constrains the language style of the reply to be serious, witty or lovely, etc.), and the tool agent can return the result of the tool calling of the calling / simulated calling tool according to the tool calling statement data. The assistant agent can summarize the result of the tool calling and send the tool calling result data for representing the summarized result of the tool calling to the user agent. The tool calling result data for representing the summarized result of the tool calling sent by the assistant agent to the user agent can meet the constraint of the dialogue flow (the tool calling result data needs to meet the language style of the reply constrained by the dialogue flow).
[0245] For example, in addition to the dialogue flow, the dialogue flow can also include / indicate a reply manner of the dialogue reply (for example, the dialogue flow constrains the language style of the reply, whether to increase the rhetorical question, etc.), and the dialogue reply data can also need to meet the reply manner of the dialogue flow.
[0246] Based on the scheme provided in the embodiments of the present application, by limiting the reply manner of the dialogue reply by the dialogue flow, the constraint on the dialogue flow can be further strengthened, the controllability and fluency of the data generation process can be further improved, and the performance of the data generation process can be improved.
[0247] In some possible implementation manners, the S720 can include:
[0248] S720a, generating a second requirement according to the dialogue flow.
[0249] S720b, generating first requirement data according to the second requirement, the first requirement data indicating the second requirement.
[0250] In some possible implementation manners, S720 can include:
[0251] S720c, generating a third requirement according to the dialogue flow and the dialogue history, the dialogue history including requirements that have been completed, the third requirement representing a next requirement of a last requirement in the dialogue history.
[0252] S720d, generating first requirement data according to the third requirement, the first requirement data indicating the third requirement.
[0253] For example, the first requirement can include a plurality of sub-requirements, and the dialogue flow generated according to the first requirement can include a dialogue flow composed of dialogues based on the plurality of sub-requirements. Before the first requirement data is generated according to the dialogue flow, a second requirement can be generated according to the dialogue flow, the second requirement being a first requirement in the plurality of sub-requirements if any requirement in the dialogue flow has not been completed before the first requirement data is generated according to the dialogue flow. If one or more requirements in the dialogue flow have been completed but all requirements in the dialogue flow have not been completed before the first requirement data is generated according to the dialogue flow, a last requirement in the dialogue history can be determined according to the dialogue flow and the dialogue history, and a third requirement can be generated according to the last requirement in the dialogue history, the third requirement and the last requirement satisfying the dialogue flow (for example, a dialogue flow includes requirements a, b, c, d, and e in sequence, when the last requirement in the dialogue history is determined to be requirement c, a next requirement of requirement c, i.e., requirement e, can be generated, and requirement e can be the third requirement).
[0254] According to the scheme provided in the embodiments of the present application, the second requirement or the third requirement indicated by the first requirement data generated according to the dialogue flow and the dialogue history can effectively generate a first requirement in the dialogue flow or a next requirement of a last requirement in the dialogue history according to the dialogue flow, and the controllability and fluency of the process of generating data can be improved.
[0255] Some possible toolset construction manners are introduced as follows:
[0256] ① The names and brief introductions of a large number of tools that can be obtained from the Internet, which may be valuable (without requiring the API of these tools to be functional, nor requiring them to have structured documents that LLMs can directly use), and using the text generation capabilities of LLMs, a comprehensive tool set can be built, and a standardized document format can be built for each tool. The documentation of tools with structured documents can describe the functions and usage of the tools in detail. In this way, a diversified and structured tool set similar to real scenarios can be built.
[0257] ② Existing tool sets can also be obtained, which can have structured documents that LLMs can directly use; or, existing tool sets can not have structured documents that LLMs can directly use, and LLMs can optimize or supplement existing tool sets based on their own capabilities, ultimately obtaining a tool set with structured documents that LLMs can directly use.
[0258] ③ A large number of real tools corresponding to names, introductions, descriptions, function documents, OpenAPI specifications, etc. can also be obtained, and a tool set of real tools can be built based on the above content.
[0259] Tool usage examples: In order to obtain tool usage examples in the above tool set, a simulation environment can be designed to simulate the interaction between language models, users, and tools. For example, LLMs can play different agents, each agent has a specific prompt, and multiple agents can implement the interaction between language models, users, and tools through the above method 700 or the following process 800. In this way, a large number of tool usage examples can be generated without any human intervention. Each tool usage example can consist of three key elements: {user's instructions, operations and their corresponding tool outputs, and final responses}.
[0260] It can be understood that when the user agent's demand does not involve the calling of tools, the helper agent can generate a reply directly without calling tools. At this time, the user agent's instructions and the helper agent's final response can constitute a non-tool usage example.
[0261] By LLMs playing the agents in the above method 700, tool usage examples and / or non-tool usage examples can be obtained. Tool usage examples can be used to train SLMs, CLMs, and other language models, so that SLMs, CLMs, and other language models can learn corresponding knowledge and action patterns, and expand the capability boundaries of SLMs, CLMs, and other language models; Non-tool usage examples can be used together with tool usage examples to train SLMs, CLMs, and other language models, so that SLMs, CLMs, and other language models can learn to call tools while maintaining their original general capabilities.
[0262] It can be understood that the first demand data and the first response data can be taken as training data, and a language model can be trained by a model training method. The trained language model can be used for demand processing by a demand processing method.
[0263] For example, the first demand data and the first response data can be taken as training data, and the training data can be input into the language model. The training of the language model can be implemented by fine-tuning the language model.
[0264] For example, when the first response data meets the reply mode specified by the dialogue flow, the language model can be trained using the first response data, and the language model can be required to have the same reply mode. For example, the same reply mode requirement can be implemented by adding a corresponding prompt in the first response data.
[0265] The specific implementation of the method for training the language model can refer to related technologies, and the embodiments of the present application will not be repeated.
[0266] For example, after determining the demand to be processed, the demand to be processed can be input into the language model, and the language model is trained by taking the first demand data and the first response data as training data. The demand to be processed can be processed by inputting the demand to be processed into the language model, and the demand processing method can be used for demand processing to obtain a first processing result. The specific implementation of the demand processing method can refer to related technologies, and the embodiments of the present application will not be repeated.
[0267] For example, the demand of "booking a train ticket" input by the user can be determined as the demand to be processed, and the language model can process the demand and output the result of booking a train ticket. The result of booking a train ticket is one possible implementation of the first processing result.
[0268] It can be understood that the above-mentioned data generation method, model training method or demand processing method can be implemented by a device. For example, the above-mentioned data generation method, model training method or demand processing method can be implemented by a processor in the device. The specific implementation of the above-mentioned method by the processor in the device can refer to the summary, and will not be repeated here.
[0269] FIG. 8 is a schematic diagram of a data generation flow 800 provided by an embodiment of the present application. The flow 800 can be used to implement the above-mentioned method 700, and the flow 800 can be used to generate tool usage instances and / or non-tool usage instances.
[0270] As shown in (a) of FIG. 8, the dialogue planning agent can generate a dialogue flow according to the given demand and a tool list, etc., and the dialogue flow can constitute constraints on the dialogue process. The dialogue process planned in the dialogue flow needs to meet the given demand and each demand transition is reasonable and natural.
[0271] The following will take the given demand of "booking a ticket" as an example to illustrate the process 800 provided by the embodiments of the present application in detail. It should be understood that the given demand of "booking a ticket" is only an example for the purpose of understanding the present application and should not be construed as a limitation of the present application. The given demand can also be a demand such as "booking a hotel", "booking a ticket", "booking a meeting", "tourist suggestion", etc. The given demand can be determined according to the aspect that is intended to be trained, or can be determined randomly according to a certain rule in a demand library that stores a plurality of demands, and the embodiments of the present application do not limit this.
[0272] For example, when it is intended to train the tool calling capability of the aspect of booking a ticket of the language model, the given demand can be "booking a ticket"; or, among the demands stored in the demand library, a demand can be randomly selected as the given demand.
[0273] When the given demand is "booking a ticket", the dialogue planning agent can generate a dialogue flow according to "booking a ticket" and a tool list, etc. For example, the dialogue flow planned by the dialogue planning agent can include tool calling demands, casual conversations or non-tool calling demands; the tool list can be composed of tools in the tool set, and is used to represent tools that can be called or simulated by the assistant agent.
[0274] The tool calling demand refers to a demand that a user needs to use a certain tool or system to complete a specific task or obtain specific information. The tool calling demand has a clear purpose, and the user hopes to achieve a specific function or obtain a specific result through tool calling. For example, the user hopes to book a ticket through tool calling.
[0275] The non-tool calling demand refers to a demand that a user does not need to use a specific tool, but needs to obtain certain information, suggestion or perform certain non-tool interaction. The non-tool calling demand is more flexible and diversified, and can involve the user's emotions, opinions, suggestions or general information query. For example, the user can inquire about tourist suggestions, seek commodity recommendations, discuss news events, etc., which do not need the support of specific tools, but need to interact and discuss with the user more widely.
[0276] Idle chat refers to the need of users to have casual and non-purposeful conversation with each other. Idle chat is casual and non-purposeful, and users can just want to have friendly communication with others, share ideas or feelings. For example, daily chat between users like friends, casual communication on social media, etc. The conversation in these scenarios usually has no specific purpose or need, but just to enhance mutual understanding and friendship.
[0277] For example, the dialogue flow can plan a dialogue process based on the following needs:
[0278] Need ①: The user raises the need of booking a ticket.
[0279] Need ②: The user raises idle chat.
[0280] Need ③: The user asks for travel suggestions.
[0281] Need ④: The user adds the need of booking a ticket.
[0282] Need ⑤: The user raises the need of booking a return ticket.
[0283] The above five needs are only illustrative and do not limit the present application. The above five needs are reasonably and naturally converted, and the dialogue process based on the above five needs meets the given needs.
[0284] The user agent generates needs that meet the above dialogue process.
[0285] Specifically, after the dialogue planning agent generates the dialogue flow, the user agent needs to determine which need is currently in the dialogue flow plan according to the dialogue flow and the dialogue history, and then generate the corresponding need. The generated corresponding need is the next need of the last completed need in the dialogue process.
[0286] For example, the user agent can generate need ① according to the dialogue history if no need is completed. The user agent can generate the next need of need ①, i.e. need ②, according to the dialogue history if the last completed need is need ①. The user agent can generate the next need of need ②, i.e. need ③, according to the dialogue history if the last completed need is need ②. The user agent can generate the next need of need ③, i.e. need ④, according to the dialogue history if the last completed need is need ③. The user agent can generate the next need of need ④, i.e. need ⑤, according to the dialogue history if the last completed need is need ④. The user agent can end the generation of data according to the dialogue history if the last completed need is need ⑤, since there is no other need after need ⑤.
[0287] If the requirement generated by the user agent does not conform to the above dialogue flow, the user agent needs to re-generate the requirement.
[0288] In some possible implementation manners, in addition to the dialogue planning agent and the user agent, a helper agent for generating response data responding to the requirement generated by the user agent can be further arranged.
[0289] The helper agent in the scheme can be one helper agent as shown in FIG. 8, which can receive the interaction requirement / reply follow-up data and the like sent by each user agent, and generate follow-up data / complete requirement data for responding to the user agent. The helper agent in the scheme can also be multiple helper agents (not shown in FIG. 8), which can respectively receive the interaction requirement / reply follow-up data and the like sent by each user agent, and generate multiple follow-up data / complete requirement data for responding to the user agent (for example, multiple helper agents can respectively generate multiple response data responding to the first requirement, and the multiple response data responding to the first requirement are used for responding to the user agent). The embodiments of the present application do not limit this.
[0290] For example, the helper agent can generate response data responding to the requirement generated by the user agent. After the response data is generated, whether the response data involves tool calling can be determined according to the format of the response data. If the response data involves tool calling, the response data is considered as a tool calling statement; if the response data does not involve tool calling, the response data is considered as a dialogue reply.
[0291] In some possible implementation manners, in addition to the dialogue planning agent, the helper agent and the user agent, a tool agent for calling tools can be further arranged.
[0292] In some possible implementation manners, in addition to the dialogue planning agent, the helper agent and the user agent, a verification agent for verifying whether the response data meets consistency can be further arranged. After the helper agent generates one / multiple dialogue replies or one / multiple tool calling statements, the generated one / multiple dialogue replies or one / multiple tool calling statements can be sent to the verification agent, and whether the generated one / multiple dialogue replies or one / multiple tool calling statements meet consistency can be verified by the verification agent.
[0293] For example, the verification agent can verify whether the multiple dialogue replies / multiple tool invocation statements generated by one helper agent satisfy consistency, as shown in FIG. 8. The verification agent can also verify whether the multiple dialogue replies / multiple tool invocation statements generated by multiple helper agents satisfy consistency. The verification agent can further verify whether the multiple dialogue replies / multiple tool invocation statements generated by one / multiple helper agents and the verification agent itself satisfy consistency (for example, the verification agent can receive the interaction demand / reply follow-up question data sent by each user agent, and generate follow-up question / complete demand data for responding to the user agent (not shown in FIG. 8). The verification agent can verify whether the response data generated by itself and the received response data satisfy consistency, and the received response data can be generated by one / multiple helper agents according to the demand data sent by the user agent).
[0294] After the user agent generates demand ① according to the dialogue flow, the demand ① needs to be sent to the helper agent (and can also be sent to the verification agent). If the demand ① involves using a tool, the response data of the demand ① generated by the helper agent and / or the verification agent should include tool invocation statements. The helper agent can send tool invocation statement data including the generated tool invocation statements to the tool agent to obtain tool invocation results. If multiple tool invocation statements generated by the helper agent and / or the verification agent satisfy consistency, any one of the multiple tool invocation statements satisfying consistency can be determined as tool invocation statement data, and the tool invocation statement data is considered correct. The tool agent can send the tool invocation results to the helper agent, and the helper agent can summarize the tool invocation results and reply to the user agent. If multiple tool invocation statements generated by the helper agent and / or the verification agent do not satisfy consistency and the number of the multiple tool invocation statements is greater than or equal to a preset threshold (for example, a first threshold), any one of the multiple tool invocation statements not satisfying consistency can be marked with a mask, and the any one tool invocation statement and the mask are determined as tool invocation statement data.
[0295] The specific details of determining the tool invocation statement data according to the consistency of the multiple tool invocation statements can refer to the related description of the above step S730.
[0296] After the user agent generates the requirement ③ according to the dialogue flow, the requirement ③ needs to be sent to the assistant agent (and can also be sent to the verification agent). The requirement ③ does not involve the use of tools, so the data generated by the assistant agent and / or the verification agent in response to the requirement ③ should include a dialogue reply. After the assistant agent generates the dialogue reply, the dialogue reply data can be sent to the user agent according to the generated dialogue reply.
[0297] For example, the assistant agent and / or the verification agent can generate 2 dialogue replies, reply #A and reply #B, the verification agent performs consistency verification on reply #A and reply #B, when reply #A and reply #B meet the condition of substantially consistent content or consistent content, it can be determined that reply #A / reply #B is the most reasonable reply, and the assistant agent sends the dialogue reply data including the most reasonable reply to the user agent; when reply #A and reply #B do not meet the condition of substantially consistent content or consistent content, the assistant agent and / or the verification agent can generate reply #C, and the verification agent performs consistency verification, until the most reasonable reply is determined among multiple dialogue replies; or, a threshold (for example, a second threshold) can be preset, when the number of dialogue replies is greater than or equal to the threshold and there is no most reasonable reply among the generated multiple dialogue replies, any one of the multiple dialogue replies can be selected to be marked with a mask, and the any one dialogue reply and the mask are sent to the user agent as the dialogue reply (i.e., the any one dialogue reply and the mask are determined as the dialogue reply data).
[0298] The specific details of determining the dialogue reply data according to the consistency of multiple dialogue replies can refer to the related description of step S730 described above.
[0299] The requirement generated by the user agent according to the dialogue flow, the reply sent by the assistant agent to the user agent, or the tool invocation statement and / or dialogue reply generated by the assistant agent and / or the verification agent described above can all be used as data for continuing to train the language model, so that the ability of the trained language model can be improved in the direction of the expected ability. When the data for continuing to train the language model has a mask mark, it can be skipped in training, thereby reducing the errors of the assistant agent. When the user's requirement involves tool invocation, the reply sent by the assistant agent to the user agent can include the result of the tool invocation.
[0300] As shown in (b) of FIG. 8, the user agent needs to generate a demand according to the dialogue flow and the dialogue history. For example, the user agent judges that the last demand that has been completed at present is the user's demand for adding a ticket reservation according to the dialogue history, and the user agent can generate the next demand of the user's demand for adding a ticket reservation, i.e., the demand for reserving a return ticket, according to the dialogue flow.
[0301] In some possible implementation manners, no special requirements can be made to the helper agent and / or the verification agent, as long as the helper agent and / or the verification agent can accurately answer the questions of the user agent and the verification agent can verify the consistency.
[0302] In another possible implementation manner, the reply manner of the reply generated by the helper agent and / or the verification agent needs to conform to the constraint of the dialogue flow.
[0303] For example, the reply manner of the reply generated by the helper agent and / or the verification agent can also be constrained by the dialogue flow, for example, the dialogue flow can limit the reply manner of the reply generated by the helper agent and / or the verification agent (such as the style of the reply of the reply generated by the helper agent and / or the verification agent, whether it is necessary to add a counter-question, etc.). At this time, the user agent and the helper agent and / or the verification agent need to comply with the dialogue flow when generating. If the reply manner of the reply generated by the helper agent and / or the verification agent does not conform to the dialogue flow, the helper agent and / or the verification agent need to regenerate the reply.
[0304] For example, one or more of the dialogue planning agent, the user agent, the helper agent, the tool agent, and the verification agent can be an LLM agent, and based on the data obtained by the LLM agent, the SLM and the CLM language model can be trained to expand the capability boundary of the SLM and the CLM language model. The tool agent can also be a real tool interface.
[0305] The method 700 or the flow 800 can be used for content generation of specific demands, construction of dialogue robots, etc. For example, the method 700 or the flow 800 can be used for constructing a script, and the expected plot development can be taken as a dialogue flow, different roles of LLM are played, and a script conforming to the expectation is generated based on the dialogue flow.
[0306] FIG. 9 is a schematic diagram of an apparatus 900 for generating data or model training or demand processing according to an embodiment of the present application. As shown in FIG. 9, the apparatus 900 can be a data generation device with a data generation function, a model training device with a model training function, or a demand processing device with a demand processing function, or a component (for example, a unit, a module, a chip, or a chip system) configured in the data generation device, the model training device, or the demand processing device. The apparatus 900 includes a processing unit 920, and optionally, a transceiver unit 910. The transceiver unit 910 can be configured to implement a transceiving function. The transceiver unit 910 can also be referred to as a communication interface or a communication unit. The processing unit 920 can be configured to process a demand to be processed or first demand data or first demand.
[0307] Optionally, the apparatus 900 can further include a storage unit, which can be configured to store instructions and / or data. The processing unit 920 can read the instructions and / or data in the storage unit, so that the apparatus implements the foregoing method embodiments.
[0308] For example, the apparatus 900 can be a data generation device with a data generation function, or can be applied to a data generation device, or be used with a data generation device, or be a data generation apparatus capable of implementing the method executed by a data generation device, such as a chip, a chip system, or a circuit. For details, refer to the related description of the chip system shown in FIG. 11.
[0309] For example, the apparatus 900 can be a model training device with a model training function, or can be applied to a model training device, or be used with a model training device, or be a model training apparatus capable of implementing the method executed by a model training device, such as a chip, a chip system, or a circuit. For details, refer to the related description of the chip system shown in FIG. 11.
[0310] For example, the apparatus 900 can be a demand processing device with a demand processing function, or can be applied to a demand processing device, or be used with a demand processing device, or be a demand processing apparatus capable of implementing the method executed by a demand processing device, such as a chip, a chip system, or a circuit. For details, refer to the related description of the chip system shown in FIG. 11.
[0311] As a design, the apparatus 900 can be configured to execute the steps or procedures of the method embodiments of FIG. 7. The processing unit 920 is configured to execute the processing-related operations (for example, the steps S710-S730 described above) in the method embodiments of FIG. 7. The transceiver unit 910 can be configured to receive data, and the processing unit 920 can take the data received by the transceiver unit 910 as the demand to be processed or the first demand or the first demand data.
[0312] As a design, the processing unit 920 can be used to input the to-be-processed demand into the trained language model to obtain a first processing result.
[0313] It should be understood that the specific process of each unit performing the corresponding steps described above has been described in detail in the method embodiments described above, and for the sake of brevity, will not be repeated here.
[0314] It should also be understood that the apparatus 900 herein is embodied in the form of functional units. The term "unit" herein can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combination of logic circuitry and / or other suitable components supporting the described functions.
[0315] The apparatus 900 of each of the above schemes can have the functions of implementing the corresponding steps in the method 700 described above. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functions described above; for example, the transceiver unit can be replaced by a transceiver (for example, the transmitting unit in the transceiver unit can be replaced by a transmitter, and the receiving unit in the transceiver unit can be replaced by a receiver), and other units, such as the processing unit, can be replaced by a processor, which respectively performs the transceiving operations and related processing operations in each method embodiment.
[0316] In addition, the transceiver unit 910 described above can also be a transceiver circuit (for example, it can include a receiving circuit and a transmitting circuit), and the processing unit can be a processing circuit.
[0317] It should be noted that the apparatus in FIG. 9 can be the data or model generation training or demand processing device in the foregoing embodiments, or can be a chip or a chip system, such as a system on chip (SoC). Among them, the transceiver unit can be an input / output circuit, a communication interface; the processing unit is a processor or microprocessor or integrated circuit integrated on the chip. This is not limited here.
[0318] FIG. 10 is a schematic diagram of an apparatus 1000 for generating data or model training or demand processing according to an embodiment of the present application. As shown in FIG. 10, the apparatus 1000 includes a processor 1010. Optionally, the apparatus 1000 further includes a communication interface 1030 for receiving and / or sending signals or data, and the processor 1010 is configured to process the received signals or data. For example, the processor 1010 is configured to process the signals or data received and / or sent by the communication interface 1030.
[0319] Optionally, the apparatus 1000 further includes a memory 1020, and the processor 1010 is coupled to the memory 1020. The memory 1020 is configured to store programs or instructions and / or data for generating data or model training or demand processing. The processor 1010 is configured to execute the programs or instructions for generating data or model training or demand processing stored in the memory 1020, or read the data stored in the memory 1020, to perform the methods in the above method embodiments.
[0320] Optionally, the processor 1010 is one or more.
[0321] Optionally, the memory 1020 is one or more.
[0322] Optionally, the memory 1020 and the processor 1010 are integrated together, or are separately arranged.
[0323] For example, the processor 1010 can have the functions of the processing unit 920 shown in FIG. 9, the memory 1020 can have the functions of the storage unit, and the communication interface 1030 can have the functions of the transceiver unit 910 shown in FIG. 9.
[0324] For example, the apparatus 1000 can be a data generating device with a data generating function, or can be a data generating device applied to a data generating device, or matched with a data generating device, and can implement the method executed by the data generating device, such as a chip, a chip system or a circuit. For details, please refer to the related description of the chip system shown in FIG. 11.
[0325] For example, the apparatus 1000 can be a model training device with a model training function, or can be a model training device applied to a model training device, or matched with a model training device, and can implement the method executed by the model training device, such as a chip, a chip system or a circuit. For details, please refer to the related description of the chip system shown in FIG. 11.
[0326] For example, the apparatus 1000 can be a demand processing device with a demand processing function, or can be a demand processing device applied to a demand processing device, or matched with a demand processing device, and can implement the method executed by the demand processing device, such as a chip, a chip system or a circuit. For details, please refer to the related description of the chip system shown in FIG. 11.
[0327] As a design, the apparatus 1000 is configured to perform the steps or flows in the method embodiments of FIG. 7 above, the communication interface 1030 is configured to perform the operations related to transceiving in the method embodiments above, and the processor 1010 is configured to perform the operations related to processing in the method embodiments of FIG. 7 above (for example, determining the to-be-processed demand; inputting the to-be-processed demand into the language model after continued training to obtain the first processing result, etc.).
[0328] It should be understood that the processor mentioned in the embodiments of the present application can be a device or a part of circuit for processing function in the following devices: a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), ASICs, field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware code processing executed by the processor, or executed by a combination of hardware and software modules in the code processing. The software module can be located in a storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the storage 1020, and the processor 1010 reads the information in the storage 1020, and combines the hardware to complete the functions required by the units included in the electronic device, or executes the data processing method of the method embodiments of the present application.
[0329] The communication interface uses a transceiving device such as but not limited to a transceiver to realize the communication between the apparatus 1000 and other devices or communication networks. For example, the to-be-processed data can be obtained through the communication interface.
[0330] The bus can include a path for transmitting information between the various components (for example, the storage 1020, the processor 1010, the communication interface 1030) of the apparatus 1000.
[0331] Next, the chip system in the data generation device / demand processing device / model training device will be described in combination with FIG. 11.
[0332] FIG. 11 is a schematic diagram of a chip system 1100 according to an embodiment of the present application. The chip system 1100 (or also referred to as a processing system) includes a logic circuit 1110 and an input / output interface 1120.
[0333] The logic circuit 1110 can be a processing circuit in the chip system 1100. The logic circuit 1110 can be coupled with a storage unit, and invoke instructions in the storage unit, so that the chip system 1100 can implement the methods and functions of the embodiments of the present application. The input / output interface 1120 can be an input / output circuit in the chip system 1100, and output information processed by the chip system 1100, or input data or signaling information to be processed by the chip system 1100.
[0334] For example, if the demand processing device installs the chip system 1100, the logic circuit 1110 is coupled with the input / output interface 1120, and the input / output interface 1120 can input the input information to the logic circuit 1110 for processing.
[0335] The embodiments of the present application provide a computer readable storage medium, which stores computer instructions for implementing the method executed by the data generation device / demand processing device / model training device in the above-mentioned method embodiments.
[0336] For example, the computer program is executed by a computer, so that the computer can implement the method executed by the data generation device / demand processing device / model training device in the above-mentioned method embodiments.
[0337] The embodiments of the present application provide a computer program product, which contains instructions, and the instructions are executed by a computer to implement the method executed by the data generation device / demand processing device / model training device in the above-mentioned method embodiments.
[0338] The explanations and beneficial effects of the related contents in any of the above-mentioned devices can refer to the corresponding method embodiments provided above, and will not be repeated here.
[0339] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0340] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-mentioned system, device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0341] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0342] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0343] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be a physically independent unit, or two or more units can be integrated into a unit.
[0344] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0345] The above is merely specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of generating data, characterized by, The method comprises: generating a dialogue flow according to a first requirement, the dialogue flow comprising a dialogue procedure conforming to the first requirement; generating first requirement data according to the dialogue flow; generating first response data according to the first requirement data, the first requirement data and the first response data being used to train a language model.
2. The method of claim 1, wherein: the dialogue flow indicates a reply mode of a dialogue reply; when the first response data comprises dialogue reply data, the first response data satisfies the reply mode; or when the first response data comprises tool call statement data, tool call result data satisfying the reply mode is determined according to the tool call statement data.
3. The method according to claim 1 or 2, characterized in that, The first requirement data is generated according to the dialogue flow, comprising: generating a second requirement according to the dialogue procedure; generating the first requirement data according to the second requirement, the first requirement data indicating the second requirement; and / or generating a third requirement according to the dialogue procedure and a dialogue history, the dialogue history comprising completed requirements, the third requirement representing a next requirement of a last requirement in the dialogue history; generating the first requirement data according to the third requirement, the first requirement data indicating the third requirement.
4. The method of any one of claims 1 to 3, wherein: the first response data comprises tool call statement data; generating the first response data according to the first requirement data, comprising: generating a plurality of tool call statements according to the first requirement data; determining the tool call statement data according to consistency of the plurality of tool call statements.
5. The method of claim 4, wherein: when at least two tool call statements in the plurality of tool call statements indicate the same tool name and / or parameter input, the tool call statement data comprises any one of the at least two tool call statements; or when at least two tool call statements in the plurality of tool call statements do not indicate the same tool name and / or parameter input, and the sum of the number of the plurality of tool call statements satisfies a first threshold, the tool call statement data comprises any one of the plurality of tool call statements and a first mask mark, the first mask mark being used to indicate that the any one tool call statement is incorrect.
6. The method of any one of claims 1 to 3, wherein: the first response data comprises dialogue reply data; generating the first response data according to the first requirement data, comprising: generating a plurality of dialogue replies according to the first requirement data; determining the dialogue reply data according to consistency of the plurality of dialogue replies.
7. The method of claim 6, wherein: when at least two dialogue replies in the plurality of dialogue replies satisfy content consistency, the dialogue reply data comprises any one of the at least two dialogue replies; or When the content of at least two of the plurality of dialogue replies does not satisfy the content consistency, and the sum of the number of the plurality of dialogue replies satisfies a second threshold, the dialogue reply data comprises any one of the plurality of dialogue replies and a second mask mark, the second mask mark being used to indicate that the any one of the dialogue replies is incorrect.
8. A method of demand processing, characterized by, Comprising: determining a demand to be processed; inputting the demand to be processed into a language model to obtain a first processing result, wherein the language model is trained according to first demand data and first response data, the first demand data is generated according to a dialogue flow, the dialogue flow is generated according to a first demand, the dialogue flow comprises a dialogue process that meets the first demand, and the first response data is generated according to the first demand data.
9. The method of claim 8, wherein, the dialogue flow indicates a reply mode of a dialogue reply; when the first response data comprises dialogue reply data, the first response data satisfies the reply mode; or when the first response data comprises tool invocation statement data, tool invocation result data satisfies the reply mode, and the tool invocation result data is determined according to the tool invocation statement data.
10. The method of claim 8 or 9, wherein, the first demand data is generated according to a second demand, the first demand data indicates the second demand, and the second demand is generated according to the dialogue process; or the first demand data is generated according to a third demand, the first demand data indicates the third demand, and the third demand is generated according to the dialogue process and a dialogue history, the dialogue history comprises a demand that has been completed, and the third demand represents a next demand of a last demand in the dialogue history.
11. An apparatus for generating data, characterized by Comprising a processor, the processor being configured to: generate a dialogue flow according to a first demand, the dialogue flow comprising a dialogue process that meets the first demand; generate first demand data according to the dialogue flow; generate first response data according to the first demand data, the first demand data and the first response data being used to train a language model.
12. The apparatus of claim 11, wherein, The device is a chip.
13. An apparatus for demand processing, characterized by Comprising a processor, the processor being configured to: determine a demand to be processed; input the demand to be processed into a language model to obtain a first processing result, wherein the language model is trained according to first demand data and first response data, the first demand data is generated according to a dialogue flow, the dialogue flow is generated according to a first demand, the dialogue flow comprises a dialogue process that meets the first demand, and the first response data is generated according to the first demand data.
14. The apparatus of claim 13, wherein, The device is a chip.
15. A chip or chip system, characterized by Comprising: a circuit, the circuit being configured to perform the method of any one of claims 1 to 7, or the method of any one of claims 8 to 10.
16. A computer program product, characterised in that, When the computer program in the computer program product is executed by a computing device, the method of any one of claims 1 to 7, or the method of any one of claims 8 to 10 is implemented.
17. A computer readable storage medium characterized by: The storage medium stores a computer program or instructions, which, when executed by a computing device, implement the method of any one of claims 1-7, or the method of any one of claims 8-10.
Citation Information
Patent Citations
Task solution-oriented training method and use method of generative large language model
CN116756564A
Dialogue processing method and device, dialogue model training method and device, equipment and medium
CN117216212A
Enhancing dialogue management systems using fact fetchers
US20240086434A1
Natural language processing applications using large language models
US20240095463A1
Systems and Methods for Analysis of Home Telematics Using Generative AI
US20240289596A1