Multi-agent API calling data generation method and system based on large language model
By dynamically generating API call data through a multi-agent framework, the problem of insufficient training data in existing technologies is solved, high-quality and diverse API call training data generation is achieved, the adaptability and reliability of large language models are improved, manual intervention is reduced, and development efficiency is improved.
Patent Information
- Application Number
- CN202510404203.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, obtaining high-quality and diverse API call training datasets relies on manual annotation or limited rule sets, which is difficult to meet the needs of complex application scenarios and limits the adaptability and reliability of large language models (LLMs) in practical applications.
It adopts a multi-agent framework based on a large language model, builds a structured Schema knowledge base, and uses a multi-agent collaborative mechanism to dynamically generate API call data, including user questions, question expansion, API recall, and result checking agents, to simulate real-world interaction processes and generate standardized and widely representative API call training data.
It significantly improves the quality and diversity of API call data, reduces manual intervention, enhances the adaptability and reliability of large language models in different scenarios, and improves development efficiency and resource utilization.
Smart Images

Figure CN120670835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method and system for generating multi-agent API call data based on a large language model, which is mainly used to automatically generate application program interface (API) call training data. Background Art
[0002] In recent years, with the rapid development of large language models (LLMs), the ability of models to automatically select and correctly call external APIs based on context to perform specific tasks has become an important research topic to improve the practicality and interactivity of models. In AI-driven applications, automated task processing relies not only on the model's understanding of user intent, but also requires it to accurately identify when and how to call appropriate API services to complete these tasks.
[0003] However, a major challenge in enhancing LLM's ability to call APIs is obtaining high-quality and diverse API call training datasets. Traditional methods usually rely on manual annotation or a limited set of rules to generate training data, which limits the quality and generalization ability of the generated data and makes it difficult to meet the growing needs in complex application scenarios. Summary of the Invention
[0004] To address the above issues, the present invention proposes a multi-agent API call data generation method and system based on a large language model to address the problem of insufficient high-quality training data in specific scenarios. By introducing a multi-agent framework, a more efficient and flexible API call training data generation mechanism is achieved. This aims to simulate real-world interaction processes, batch-generate standardized and widely representative API call instances, and enhance the adaptability and reliability of LLM in specific application scenarios. The above-mentioned inventive objectives of the present invention are achieved through the following technical solutions:
[0005] The present invention provides a method for generating multi-agent API call data based on a large language model, comprising:
[0006] Step S1: sort out the API set in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, classify API associations based on the similarity, and form a scenario knowledge base;
[0007] Step S2: Design a multi-agent collaborative mechanism based on the scenario knowledge base to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, an API matching agent, and a result checking agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, the API recall agent simulates multiple recall situations, and the matching agent maps the questions to specified API calls, and the result checking agent verifies the correctness and consistency of the API calls.
[0008] In step S3, a sampling inspection is performed based on the generated API call data, and the sampling results are used as samples. The result inspection agent is used to re-label the data and select the API call data that meets the requirements as the final training data set.
[0009] Furthermore, step S1 includes:
[0010] Organize the API set in the target application scenario according to the predefined format and build the corresponding structured Schema knowledge system. The structured Schema knowledge includes the API interface name, interface description, API input parameters, and API output parameters.
[0011] Furthermore, in step S1, the similarity between APIs is calculated by embedding technology, and the API association relationships are classified according to the similarity to form a scenario knowledge base, including:
[0012] Through embedding technology, the interface descriptions in the API set and the corresponding question list are converted into vector representations. Based on the vector representation, the similarity score between each pair of APIs is calculated using cosine similarity. The APIs are then classified according to the similarity score and the classification criteria of the association relationship, including similar categories, unrelated categories, and general similar categories.
[0013] Based on the classification results, a scenario knowledge base containing structured Schema knowledge is generated.
[0014] Furthermore, in step S2, the user asks the method agent to generate a corresponding random question, and the method expansion agent performs a preset variant operation on the question, including:
[0015] Step S21: The user questioning agent randomly generates context-appropriate questions based on the scenario knowledge base and the large language model as the initial question set.
[0016] Step S22, based on the initial question set, the intelligent agent is expanded through questioning to perform variant operations, the variant operations include synonym conversion and grammatical deformation, and also includes analyzing and generating equivalent but differently expressed questions in the initial question set through natural language processing technology to expand the initial question set.
[0017] Furthermore, after simulating multiple recall situations through the API recall agent, the matching agent maps the problem to a specified API call and the result checking agent verifies the correctness and consistency of the API call, including:
[0018] Step S23: Based on the questioning in the initial question set and the structured information in the Schema knowledge, identify the keywords and core intent involved in the question; perform semantic matching through the natural language understanding capabilities of the large language model, and search for matching APIs based on the classification results in step S1; and form multiple sets of question-API list-API recall instructions for recall situations, including unrecalled APIs, recalled APIs with errors, and recalled multiple similar APIs;
[0019] In step S24, the API call data is checked according to the preset standards, and a general Prompt template is constructed based on each set of questions that meet the preset standards and the corresponding API list, which is input into the large language model and the corresponding API call instructions are output.
[0020] Furthermore, the inspection agent verifies the instance of the API call according to the preset criteria, including:
[0021] Step S25: Verify the syntax specification, input parameter rationality, parameter integrity and response structure correctness of the API call data. The result checking agent verifies the API call data according to the preset Schema knowledge standard. The verification process is formalized as follows:
[0022] Among them, V(C i , S) represents API call data C i Whether it meets the Schema knowledge standard, 1 means it meets the standard, and 0 means it does not meet the standard;
[0023] In step S26, based on the standard API call data that conforms to the Schema knowledge, through the question set expanded by the expanded agent in step S22, a Prompt template for consistency detection is constructed to call the large model reasoning for consistency verification and mark it.
[0024] Furthermore, we conduct sampling inspection based on the generated API call data, and use the sampling results as samples. The result inspection agent re-labels the data and selects the API call data that meets the requirements as the final training data set, including:
[0025] After randomly sampling samples based on the labeled API call data for verification, a set of compliant samples and a set of non-compliant samples are constructed based on the sampling results. The sets are input into the Prompt template for consistency detection in step S26 for re-evaluation, and the non-compliant samples are removed based on the evaluation results to generate the final training data set.
[0026] Based on the same inventive concept, the present invention also provides a multi-agent API call data generation system based on a large language model, which adopts the multi-agent API call data generation method as described above, including:
[0027] The scenario knowledge base construction module is used to sort out the API collection in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, and classify API associations based on the similarity to form a scenario knowledge base;
[0028] The agent collaborative generation module is used to design a multi-agent collaborative mechanism based on the scenario knowledge base to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, and an API matching agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, and the API recall agent simulates multiple recall situations. The matching agent maps the questions to specified API calls, and the result checking agent verifies the correctness and consistency of the API calls.
[0029] The data verification and screening module is used to perform sampling inspection based on the generated API call data, and re-label the sample based on the sampling results through the result inspection agent, and screen the qualified API call data as the final training data set.
[0030] Furthermore, the scenario knowledge base construction module includes,
[0031] The Schema knowledge construction unit is used to organize the API set in the target application scenario according to a predefined format and build a corresponding structured Schema knowledge system. The structured Schema knowledge includes the API interface name, interface description, API input parameters, and API output parameters.
[0032] The API association classification unit converts the interface descriptions in the API set and the corresponding question list into vector representations through embedding technology; based on the vector representation, the similarity score between each pair of APIs is calculated through cosine similarity, and the APIs are classified according to the similarity score and the classification criteria of the association relationship, including similar categories, unrelated categories and general similar categories; based on the classification results, a scenario knowledge base containing structured Schema knowledge is generated.
[0033] Furthermore, the agent collaborative generation module includes:
[0034] The user question agent module is used to randomly generate context-appropriate questions based on the scenario knowledge base and the large language model as the initial question set;
[0035] The question expansion agent module is used to perform variant operations based on the initial question set. Variant operations include synonym conversion and grammatical inflection. It also includes using natural language processing technology to analyze and generate equivalent but differently expressed questions in the initial question set to expand the initial question set.
[0036] The API recall agent module is used to identify the keywords and core intent involved in the questions based on the questioning and structured information in the schema knowledge in the initial question set. It performs semantic matching through the natural language understanding capabilities of the large language model and searches for matching APIs based on the classification results in step S1.
[0037] The API matching agent module is used to generate multiple sets of problem-API list-API recall instructions for recall situations, including unrecalled APIs, recalled API errors, and recalled multiple similar APIs;
[0038] The result checking agent module is used to verify the API call data according to the preset standards, and build a general prompt template based on each set of questions that meet the preset standards and the corresponding API list, input it into the large language model, and output the corresponding API call instructions.
[0039] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0040] This invention automatically generates API call training data that complies with standards and is broadly representative through a multi-agent collaborative mechanism based on a large language model. While improving data quality and diversity, it significantly reduces manual intervention and greatly improves development efficiency and resource utilization. The generated high-quality training data enhances the adaptability and reliability of the large language model in different scenarios, thereby significantly improving the practical application value of the model.
[0041] (1) The present invention uses a multi-agent collaborative mechanism based on a large language model to batch generate widely representative API call instances in a short period of time. The mechanism can dynamically adjust the data generation process to ensure that each batch of data not only meets the predefined specifications but also reaches a high level in diversity and coverage, thereby effectively improving the overall quality of API call data.
[0042] (2) By leveraging the division of labor and collaboration among multi-agent modules, the present invention is able to simulate question-answering tasks in various real-world environments. This capability enables the generated data to cover a wide range of types, from simple queries to complex scenarios, helping large language models to be trained in a wider range of application scenarios. By interacting with a wider range of tasks, the model's adaptability and reliability are significantly enhanced, and it can still perform well when faced with complex semantics or unseen scenarios.
[0043] (3) Traditional API calls for training data typically require extensive manual effort, including data collation, labeling, and verification. However, this invention significantly reduces the need for manual intervention through multi-agent collaboration. The generation process is automatically completed by the system, requiring only a small amount of sampling verification at key nodes. This automated generation method not only significantly saves time and labor costs, but also makes the generation of training data faster and more stable, thereby accelerating the cycle of model iteration and optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flowchart of the steps of the method for generating multi-agent API call data based on a large language model of the present invention;
[0045] Figure 2 This is an overall flow chart of the method for generating API call data in an embodiment of the present invention;
[0046] Figure 3 This is a table of recall results and output strategies for the API recall agent recall in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0048] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0049] First embodiment
[0050] This paper provides a multi-agent API call data generation method based on a large language model, which is used to solve the problem of insufficient training data in the existing technology. By introducing a multi-agent framework, a more efficient and flexible API call training data generation mechanism is achieved. This method not only helps to improve the accuracy of the model's function calls, but also lays a solid foundation for the development and optimization of future intelligent systems, such as Figure 1 and Figure 2 As shown, the specific embodiments of the present invention are as follows, including:
[0051] Step S1: sort out the API set in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, classify API associations according to the similarity, and form a scenario knowledge base; wherein, step S1 includes:
[0052] Step S11, sort out the API set in the target application scenario according to a predefined format, and build a corresponding structured Schema knowledge system, wherein the structured Schema knowledge includes the API interface name, interface description, API input parameters, and API output parameters.
[0053] Specifically, this step involves comprehensively reviewing all API sets involved in a specific application scenario. We collaborate with business experts and API development technicians to clarify and organize the predefined formats of each API, forming a rigorous structured schema knowledge system. This structured schema knowledge includes the following key elements:
[0054] (1) API interface name: Expressed in English, ensure that the name has a clear correspondence with its function and ensure the uniqueness of the API name;
[0055] (2) Interface Description: Detailed description of the API's functional purpose, input parameters, and output parameters. For APIs with multiple complex uses, multiple functional descriptions should be designed. A corresponding question template can also be constructed to guide the subsequent matching of questions with the API.
[0056] (3) API input parameters: List the input parameters of each API in detail, including but not limited to the input parameter name, Chinese interpretation, encoding value description and its type, such as array[integer], integer(int64), string, etc., and identify the required parameter list.
[0057] (4) API output parameters: define the output parameters in the API response result, including the output parameter name, Chinese meaning and value type.
[0058] By constructing the above-mentioned structured Schema, metadata knowledge is provided for data generation, and clear rules and boundaries are set for subsequent interactions between intelligent agents, ensuring the quality and standardization of API call instances.
[0059] Taking the operation of catering stores as an example, first, we collect and confirm the API set involved based on the business needs of the specific scenario, and work closely with business experts and technical experts to comprehensively collect and organize all the API interfaces involved in the scenario. The APIs involved include: product listing and delisting interface, store ordering interface, inventory query interface, operating indicator query interface, etc.
[0060] For each API, provide a detailed documentation description in a unified predefined format, covering the main functions, applicable scenarios, input and output parameter descriptions, and create a corresponding question template.
[0061] Ultimately, a strictly structured Schema knowledge system is formed. The Schema specifies the parameter type, structure, and possible value range of each API. The following is a Schema example:
[0062] schema:
[0063] business-overview-config:
[0064] desc: Operation indicator query interface, by entering the start date, end date, currency, brand ID, channel source, return turnover, operating income, expenditure amount, number of valid orders, average price before discount, non-operating income, refund amount, number of refund orders and other indicators
[0065] method:POST
[0066] templates: ["What was the total revenue yesterday?", "How many orders did you receive from the Meituan channel this week?"]
[0067] input:
[0068] client:[client app|pos|pc|jsc,string]
[0069] currency:[currency A, currency B, currency C, string]
[0070] sellerId:[brand id, integer(int64)]
[0071] startDate: [Statistical date (start) format: YYYY-MM-DD, string]
[0072] endDate: [Statistical date (end) format: YYYY-MM-DD, string]
[0073] orderSource:[Channel Source 1: Alipay; 2: Meituan; 3: Ele.me; 4: Pos, array[integer]]
[0074] orderType:[Order type 1: Purchase order; 2: Deposit order; 3: Product order; 4: Live shopping, array[integer]]
[0075] ouput:
[0076] indexMeaning:[index meaning, string]
[0077] indexName:[index name, string]
[0078] isHour:[1-include hourly period data, 2-excluding hourly period data, 3-do not display period and trend data, integer(int32)]
[0079] rate:[percentage, string]
[0080] rateTb:[year-on-year percentage, string]
[0081] tag:[index English identifier, string]
[0082] tagName:[index Chinese identifier, string]
[0083] theDate:[date, string]
[0084] theDateTb:[year-on-year date, string]
[0085] value:[number, string]
[0086] query_product:
[0087] info: Query product lists based on specific brands, stores, channels, and shelf status ...
[0089] In this Schema example, (1) the function and purpose of the interface are defined. Business-overview-config is used to query business indicators, supporting multiple statistical ranges and multiple channel sources. Query_product is used to query product lists based on brand, store, channel, and shelf status. (2) The request method of the interface is defined, such as POST, and sample templates. The sample templates are common user questions to help users understand the actual application scenarios of the interface: for example, "What was the total revenue yesterday?" or "How many orders did the Meituan channel have this week?"; (3) The input parameters of the interface are defined, including the field name, type, description, value range, and format requirements of the input parameters. It is also noted whether each parameter is required and its meaning and purpose are explained. The input parameters in this example include: ①client: client type, string, possible values include app, pos, pc, jsc; ②currency: currency type, string, optional values include currency A, currency B, currency C, etc.; ③startDate and endDate: statistical start and end dates, formatted as YYYY-MM-DD; ④orderSource and orderType: supported channel sources and order types, array type, integer value. (4) Define the output fields of the interface, determine the field name, type, and meaning of the data returned by the interface, for example: ①indexMeaning: indicator meaning, ②value: numerical result, ③rateTb: year-on-year percentage, and then describe the purpose of the output fields and how to parse the returned data structure. (5) Organize all defined information into a unified structured schema, and save all interface information, including input parameters, output fields, and template examples, as part of the scenario knowledge base to facilitate subsequent calls and maintenance.
[0090] Step S12: The interface descriptions in the API set and the corresponding question list are converted into vector representations through embedding technology; the similarity scores between each pair of APIs are calculated through cosine similarity based on the vector representation, and the APIs are classified according to the similarity scores and the classification criteria of the association relationship, including similar categories, unrelated categories, and general similar categories; based on the classification results, a scenario knowledge base containing structured Schema knowledge is generated.
[0091] What needs to be explained specifically is that this step calculates the similarity between the APIs in advance, and classifies the relationships between the APIs according to the similarity results. The embedding technology is used to vectorize the interface description of each API and the corresponding question template, and then the similarity score between each pair of APIs is calculated based on these vectors. According to the obtained similarity scores, the association between APIs is divided into three categories, including: ① Similarity Category A: The TopN APIs that are most similar to the given API are selected as their closely related API sets. ② Unrelated Category C: When the similarity between two APIs is lower than the set threshold, they are classified into the unrelated category. ③ General Similarity Category B: The remaining APIs are regarded as intermediate categories that have certain associations but do not belong to the extreme cases of high or low similarity. The following is a detailed explanation of step S12:
[0092] (1) Pre-calculate the similarity between the APIs and classify the relationships between the APIs according to the similarity results. The specific operations are as follows:
[0093] Vectorized representation: Embedding technology is used to vectorize the interface description of each API and the question in the corresponding question template. In this embodiment, deep learning models such as BERT and RoBERTa are selected as the preferred solution due to their excellent natural language understanding capabilities, including:
[0094] Choose a pre-trained language model, such as BERT or RoBERTa, to capture semantic information in the text;
[0095] Perform standardized text preprocessing on each API interface description and the questions in its related question templates, such as removing spaces and converting full-width characters to half-width characters, to ensure the quality and consistency of the input text.
[0096] Input the preprocessed text into the selected embedding model to generate a fixed-dimensional vector representation v API , where v API A vector group representing the interface description and question list.
[0097] Based on the vector representation generated above, the similarity score between each pair of APIs is calculated, including:
[0098] Cosine Similarity is selected as the similarity measurement method, and the formula is defined as:
[0099]
[0100] Among them, v i and v jThey represent the vector representation of the two APIs respectively, · represents the vector dot product, and ||v|| represents the modulus of the vector.
[0101] For each API, there are multiple question template vectors. To measure the similarity between two APIs, the maximum similarity score of all possible combinations between them is selected as the final score:
[0102] S max (API i , API j )=max{similarity(v i,k , v j,l )|k∈K i , l∈K j}
[0103] Among them, K i and K j A set of vectors representing API i and j respectively.
[0104] According to the obtained similarity score S max (API i , API j ), the association relationships between APIs are divided into three categories, and the specific classification standards are as follows:
[0105] Similarity Category A: Select the top N APIs that are most similar to a given API as its closely related API set. The specific value of Top N can be flexibly adjusted according to the actual application requirements and the selected inference model. Generally, the range of Top N can vary from 3 to 10. In this invention, it is set to 5, that is, N = 5;
[0106] Irrelevant Category C: When the similarity between two APIs falls below a set threshold, they are classified as irrelevant. This threshold is determined based on the opinions of domain experts, the characteristics of the embedding model, and technical verification results to ensure the accuracy and rationality of the classification.
[0107] Category B: The remaining APIs are considered to be in an intermediate category, showing some relevance but not falling into the extremes of high or low similarity. While not as clearly similar as Category A, these APIs may still demonstrate some potential for interaction in certain contexts. Therefore, when generating training data, appropriately considering these APIs can help improve data diversity and authenticity.
[0108] Step S2, based on the scenario knowledge base, a multi-agent collaborative mechanism is designed to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, an API matching agent, and a result checking agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, the API recall agent simulates multiple recall situations, and the matching agent maps the questions to specified API calls, and the result checking agent verifies the correctness and consistency of the API calls; including,
[0109] Step S21: The user questioning agent randomly generates context-appropriate questions based on the scenario knowledge base and the large language model as the initial question set.
[0110] Step S22: Based on the initial question set, the agent performs variant operations by using the questioning method. The variant operations include synonym conversion and grammatical transformation. In addition, the agent generates equivalent questions in the initial question set but with different expressions by using natural language processing technology to expand the initial question set.
[0111] Step S23: Based on the questioning in the initial question set and the structured information in the Schema knowledge, identify the keywords and core intent involved in the question; perform semantic matching through the natural language understanding capabilities of the large language model, and search for matching APIs based on the classification results in step S1; and form multiple sets of question-API list-API recall instructions for recall situations, including unrecalled APIs, recalled APIs with errors, and recalled multiple similar APIs;
[0112] Step S24: Verify the API call data according to the preset standards, and construct a universal prompt template based on each set of questions that meet the preset standards and the corresponding API list, input it into the large language model, and output the corresponding API call instructions;
[0113] Step S25: Verify the syntax specification, input parameter rationality, parameter integrity and response structure correctness of the API call data. The result checking agent verifies the API call data according to the preset Schema knowledge standard. The verification process is formalized as follows:
[0114] Among them, V(C i , S) represents API call data C i Whether it meets the Schema knowledge standard, 1 means it meets the standard, and 0 means it does not meet the standard;
[0115] In step S26, based on the standard API call data that conforms to the Schema knowledge, through the question set expanded by the expanded agent in step S22, a Prompt template for consistency detection is constructed to call the large model reasoning for consistency verification and mark it.
[0116] Specifically, combined Figure 2 We designed a multi-agent system to generate API call training data, where each agent is responsible for performing a specific task, simulating the process of real-world user annotation of training data. This approach dynamically generates a large number of high-quality API call examples, and each agent's underlying macro model and prompt template are independent of each other, including:
[0117] (1) Construct a user-asking agent. Design an agent that can simulate real users asking questions, hereinafter referred to as the "user-asking agent". This agent is based on a pre-built scenario knowledge base and uses the natural language generation (NLG) capability of LLM to construct contextual questions. The user-asking agent will refer to the API description, input and output parameter information in the structured schema knowledge, and combine it with actual application scenarios to create a variety of question expressions. In addition, the agent also has a certain degree of randomness to ensure that the generated questions are both close to reality and widely representative.
[0118] In this step, large models such as ChatGLM, Qwen, and BaiChuan are used to conduct in-depth analysis of the information implied by each API in turn, and automatically generate natural language questions closely related to these APIs. In this intelligent agent module, a general script is generated by writing questions to simulate real roles in the catering industry, such as brand merchants or store managers. Prompt is constructed in the script and LLM is called. For different APIs, a list of questions can be generated by simply changing the API description part in Prompt. Preferably, the questions in the question template provided by data experts can be used as a few-shot example of Prompt. The task description, API description, API parameter structure, API question example, and constraint description of the given generated question are organized into a prompt word Prompt, and the prompt word is given to the large model to generate a list of questions.
[0119] Here is an example prompt:
[0120] You are an intelligent assistant that is good at generating natural language questions based on the API structure. Do your best to generate questions based on the following API structure information. Questions can be obtained by calling the API. The following is the API structure information:
[0121] ```
[0122] business-overview-config:
[0123] desc: Operation indicator query interface, by entering the start date, end date, currency, brand ID, channel source, return turnover, operating income, expenditure amount, number of valid orders, average price before discount, non-operating income, refund amount, number of refund orders and other indicators
[0124] input:
[0125] client:[client app|pos|pc|jsc,string]
[0126] currency:[currency A, currency B, currency C, string]
[0127] sellerId:[brand id, integer(int64)]
[0128] startDate: [Statistical date (start) format: YYYY-MM-DD, string]
[0129] endDate: [Statistical date (end) format: YYYY-MM-DD, string]
[0130] orderSource:[Channel Source 1: Alipay; 2: Meituan; 3: Ele.me; 4: Pos, array[integer]]
[0131] orderType:[Order type 1: Purchase order; 2: Deposit order; 3: Product order; 4: Live shopping, array[integer]]
[0132] ouput:
[0133] indexMeaning:[index meaning, string]
[0134] indexName:[index name, string] ...
[0136] ```
[0137] The following are some examples of reference questions:
[0138] / *What was the total revenue yesterday? ;* /
[0139] / *How many orders did Meituan have this week? * /
[0140] / *Which store has the highest operating income this month* /
[0141] Require:
[0142] -You need to simulate the role of a real brand merchant to ask questions;
[0143] - Directly provide a list of questions without any explanation. The list of questions is in list format;
[0144] - Questions require diversity, randomly combining various inputs to ask questions, and there must be differences between questions.
[0145] For the interface "Business Indicator Query Interface", the large model may generate questions such as "How many live shopping orders are there this month?" or "What are the average price and expenditure amount per order in the past month?" as the initial question set.
[0146] (2) Constructing a question expansion agent. This involves introducing an agent specifically designed to expand the user's original questions, referred to as the "question expansion agent." The agent's primary task is to perform synonym conversion, grammatical transformation, or other forms of variation operations on the questions generated by the user question agent, thereby increasing the diversity of training data. Using natural language processing techniques such as word embedding and syntactic analysis, a series of new questions that are semantically equivalent to the initial questions but expressed differently are generated while maintaining the intent. This enriches the number of training samples and enhances the understanding ability of the subsequent fine-tuning model.
[0147] Get the initial set of questions Finally, the question-based augmentation agent is constructed, enriching the diversity of the training set through semantic or structural similarity. The initial question is parsed using the large model to extract the intent, key entities, and input parameters. While maintaining the intent, the entities and input parameters are reorganized, expanded, and grammatically transformed to generate a new question representation.
[0148] Specific operations include:
[0149] 1.1. Initial question analysis: Using large language models (LLMs) such as ChatGLM, Qwen, and BaiChuan, combined with carefully designed prompt templates, we conduct in-depth analysis of each initial question. The analysis process aims to extract key elements from the question, including but not limited to:
[0150] Intent I: Indicates the core purpose or need of the user asking the question.
[0151] Key entities E and input parameters P: refer to the specific objects or parameters involved in the problem, such as specific fields or variables in the API interface description.
[0152] The parsed output can be formalized as a triple (I, E, P), where I represents the intent, E represents the key entity set, and P represents the input parameter set. To achieve this goal, the present invention uses the following Prompt template:
[0153] You are an intelligent assistant that excels at parsing natural language questions. Based on the following question, extract and return its intent, key entities, and input parameters. The format is as follows:
[0154] Intent: [Description of Intent]
[0155] Key entities: [entity list]
[0156] Input parameters: [input parameter list]
[0157] Here’s the question to be solved: “What is the number of live shopping orders this month?”
[0158] 1.2. Question reorganization: Under the premise of ensuring that the intent I remains unchanged, various operations are performed on the key entity E and input parameter P to generate diversified question expressions. Specific methods include:
[0159] Entity replacement: Use synonyms or related concepts to replace the original entity, keeping the question intent unchanged. For example, "live shopping" can be replaced with "online sales activities";
[0160] Input parameter combinations: Randomly combine different input parameter values to create new questions. For example, changing parameters such as time range, currency, or channel can generate similar but not identical questions.
[0161] Grammatical transformation: Adjusting the sentence structure or using a different sentence structure to express the same idea. For example, switching from a declarative sentence to a question, or changing the subject, predicate, and object order.
[0162] These operations can be implemented by writing a series of rules or templates, or they can be automatically generated using large models. To ensure the quality of generated questions, natural language processing techniques such as dependency parsing can be introduced to ensure grammatical correctness and semantic coherence.
[0163] This embodiment adopts a large model automatic generation solution based on Prompt. The following is a Prompt example:
[0164] You are an intelligent assistant that excels at generating diverse questions. Based on the given intent, key entities, and input parameters, generate at least five different but related questions. Be mindful to preserve the original intent while maximizing the diversity of the questions.
[0165] Intent: Query order quantity
[0166] Key Entity: [Live Shopping]
[0167] Participation: [Time range: this month]
[0168] 1.3 Questioning Expansion and Transformation
[0169] Further expand and transform existing questioning methods to increase the diversity and coverage of training data. Specific measures include:
[0170] Context extension: Combine other information in the API description to add more background details or qualifications to the question. For example, add specific store location, brand name, etc. to the original question.
[0171] Multi-turn conversation simulation: This simulates real-world multi-turn conversations and generates a continuous sequence of questions. For example, it starts with a general question and then asks more detailed sub-questions based on the results.
[0172] Diversified types: Design different types of questions, such as open-ended questions, multiple-choice questions, true-or-false questions, etc., to cover a wider range of application scenarios.
[0173] The Prompt example of the present invention performing this step is:
[0174] You are an intelligent assistant that excels at generating complex questions. Based on the given intent, key entities, and input parameters, combined with API information, generate at least three questions with additional context or qualifications. Design a multi-turn conversation, including one summary question and two follow-up questions.
[0175] The API structure is as follows: ...
[0177] Intent: Query order quantity
[0178] Key Entity: [Live Shopping]
[0179] Participation: [Time range: this month]
[0180] Finally, after the above series of operations, a new batch of question expressions are expanded. At the same time, the original question and these similar questions are marked to prepare for the subsequent result checking agent.
[0181] (3) Constructing an API recall agent. In the actual process of using LLM to execute API calls, because there are too many APIs in the real production environment, it is often necessary to combine Schema information to recall the APIs that match the user's questions. Due to the complexity of the real environment, there are many situations in the API recall results. For example, how to deal with the situation when the question is not within the API support range and the API cannot be recalled; how to deal with the situation when the API that does not match the actual intention of the question is recalled; how to deal with the situation when multiple similar APIs cannot accurately recall the API with the actual intention of the question. In response to the above situations, an agent is specially designed to simulate the results of recalling APIs, hereinafter referred to as the "API recall agent". At the same time, the question list after the agent operation is expanded based on the question and the association between the APIs obtained in step S1 is combined to simulate the recall results of APIs for the situations of not recalling APIs, recalling incorrect APIs, recalling multiple similar APIs, recalling accurate APIs, recalling some related APIs but some irrelevant APIs, etc., to form multiple sets of question-API list-API recall instructions.
[0182] After the above steps, a list of questions has been formed. Since the method generation operation is performed on each API in turn, APIA i And every question The mapping relationship between them is clear. Specifically, each question q j There is a corresponding APIA i(j) , where i(j) represents the question q j The corresponding API index.
[0183] By using the obtained associations between APIs, we simulate various situations of API recall in real scenarios. The following recall situations are defined:
[0184] Negative Sample (NS): The question cannot match any API;
[0185] Recall the wrong API (Negative Sample, NS): The question matches an incorrect API;
[0186] Recall multiple similar APIs (Positive Sample, PS): The question matches multiple similar APIs, but includes the correct API;
[0187] Recalling accurate API (Positive Sample, PS): The question is accurately matched to a correct API;
[0188] Recall partially relevant APIs (Positive Sample, PS): The question matches some relevant APIs, including irrelevant APIs and correct APIs.
[0189] In order to describe these situations more accurately, we define positive samples (PS) and negative samples (NS):
[0190] Positive samples (PS): include recalling multiple similar APIs, recalling accurate APIs, and recalling some related APIs, etc., which contain the correct API;
[0191] Negative samples (NS): include cases where the correct API is not included, such as the recalled API and the recalled incorrect API.
[0192] List all questions Randomly divide into positive sample question lists and negative sample question list And divide it in the ratio of 9:1. Further, list all the positive sample questions Negative Sample Question List Divide into more specific recall situations according to a certain ratio. For example, the following ratio can be used:
[0193] Recall multiple similar APIs: 50% of positive samples;
[0194] Recall precision API: 30% of positive samples;
[0195] Recall some related APIs: accounting for 20% of positive samples;
[0196] Unrecalled API: 50% of negative samples;
[0197] Recall error API: 50% of negative samples;
[0198] For each question q j Construct the generation of API recall data. The specific operations are:
[0199] Select a question: From the question list Select one question in turn j ;
[0200] Get recall status: Determine the question q based on the pre-set recall status distribution j Recall type T j , where T j ∈{NS,PS};
[0201] Follow the recall situation definition: According to T jThe definition of , randomly selects a specific API instance as the final recall result A r For example, if T j If it is "recall multiple similar APIs", then i(j) Randomly select multiple APIs from a similar API set; if it is a "recall precision API", directly return A i(j) .
[0202] Through the above operations, the API recall agent can systematically simulate the API recall results in different situations, forming multiple sets of questions-API lists-API recall instructions. For a more intuitive understanding of the above implementation steps, please refer to Figure 3 The specific instance and output strategy in the inference model refer to how to return when encountering this situation.
[0203] (4) Design an agent to perform question-API call matching. Based on the API recall agent step, further design an agent specifically responsible for mapping user questions to corresponding API calls, hereinafter referred to as the "API matching agent". The core function of this agent is to identify the implicit intention in the user's question and, combined with the API list given to it, convert it into a specific API call result. By constructing a general Prompt template, for each set of questions and its corresponding API list, the knowledge in the corresponding API Schema document is associated, such as function description, input parameters, output parameters, etc., to form a Prompt instance, which is then handed over to the LLM for call reasoning to generate the final API call result.
[0204] Based on the question-API recall list instance constructed in the previous step, for each set of questions and their corresponding API lists, this invention will construct a general API reasoning prompt template and API function set. This template will integrate the knowledge in the schema document corresponding to the associated API, such as function description, input parameters, output parameters, etc., to form a complete prompt instance. Specifically:
[0205] - Question method: natural language questions asked by users.
[0206] -API list: recall API results constructed based on the API recall agent;
[0207] -Schema document: A structured document corresponding to each API, including function description, input parameters, output parameters, required parameter descriptions, and other information.
[0208] The general API reasoning prompt template is designed to ensure that the LLM can understand the intention of the question. The Function Calling function of the LLM large model is used to encapsulate the API list into a function set, so that the LLM can generate accurate API call instructions based on the API function description and parameter requirements.
[0209] The following is an example of an API inference prompt for reference (some large models can use the tools parameter to input the API set):
[0210] You are a text-to-API call generator. Your main goal is to help users match the input text to the correct API to call and extract the key information in the input text into the parameters of the API.
[0211] The following is the API structure information
[0212] ```
[0213] API1: business-overview-config (business indicator query interface, by entering the start date, end date, currency, brand ID, channel source, return turnover, operating income, expenditure amount, number of valid orders, average price before discount, non-operating income, refund amount, number of refunded orders and other indicators)
[0214] Input parameters:
[0215] Parameter 1: client (client app|pos|pc|jsc), Parameter 2: currency (currency A; currency B; currency C), Parameter 3: sellerId (brand id), Parameter 4: startDate (statistical date (start)), Parameter 5: endDate (statistical date (end)), Parameter 6: orderSource (channel source), Parameter 7: orderType (order type)
[0216] Output parameters:
[0217] Parameter 1: indexMeaning (index meaning), Parameter 2: indexName (index name), Parameter 3: isHour (1-includes hourly period data, 2-excluding hourly period data, 3-does not display period and trend data), Parameter 4: rate (percentage), Parameter 5: rateTb (year-on-year percentage), Parameter 6: tag (indicator English label), Parameter 7: tagName (indicator Chinese label), Parameter 8: theDate (date), Parameter 9: theDateTb (year-on-year date), Parameter 10: value (value)
[0218] ```
[0219] Q: What was yesterday's turnover? Yesterday refers to "2024-04-18".
[0220] Answer:{"api":"business-overview-config","body":{"startDate":"2024-04-18","endDate":"2024-04-18"}}
[0221] Question: How many orders did Meituan have yesterday?
[0222] answer:
[0223] The above prompt is mainly for positive samples. When a negative sample is located, it is directly output according to the output strategy in the API recall agent.
[0224] Each question is then reasoned through the large model to obtain the result.
[0225] (5) Constructing a result checking agent. To ensure the quality of generated data, we design an agent responsible for verifying the correctness of the API call results, referred to as the “result checking agent”. This agent mainly checks the following two aspects:
[0226] ① Based on the pre-set schema knowledge standards and API recall, the generated API call instances (i.e., API call results) are submitted to the LLM for comprehensive inspection. The inspection process includes, but is not limited to, checking whether the API call complies with syntax specifications, whether parameter settings are reasonable, and whether the required parameters are returned reasonably.
[0227] ② Based on the similar questions in the expanded intelligent agent's expansion of the original question method, in theory, the API call results of the similar questions should also be similar. The LLM determines whether the call conditions between similar questions are consistent, such as whether the input parameters are consistent, and whether there are only differences between the input values.
[0228] If any non-compliance with the standard is found, it is marked as non-compliant; otherwise, it is marked as compliant.
[0229] The correctness of the result-checking agent is related to the quality of the final training data generation. Therefore, the underlying large model of the agent can choose a more advanced large model with larger parameters, such as Qwen-Turbo / Qwen-Max / GPT-4o.
[0230] After obtaining the inference results for all questions, the result checking agent first conducts a comprehensive check on the API call instance to ensure that it meets the preset standards. The specific inspection content includes the following aspects:
[0231] Syntax specification: Check whether the API call follows the correct syntax format.
[0232] Parameter setting rationality: Verify whether the input parameters are reasonable and conform to the API definition.
[0233] Required parameter integrity: Confirm that all required parameters are returned correctly.
[0234] Correctness of response structure: Ensure that the data structure of the API response is consistent with expectations.
[0235] For each API call instance C i , the result checking agent verifies according to the preset Schema knowledge standard. The verification process can be formalized as: Among them, V(C i , S) represents the API call instance C i Whether it conforms to the Schema knowledge standard, 1 means it conforms, 0 means it does not conform.
[0236] The result checking agent is also responsible for evaluating the consistency of call results between similar questions. In theory, the API call results for similar questions should be similar. To this end, the agent compares the call results of similar questions to determine whether there is consistency between them. The specific implementation steps are as follows:
[0237] Determine a set of similar questions: Mark questions with similar intent in step 2.2
[0238] Evaluate input parameter consistency: For questions that only involve synonym replacement or word order adjustment without changing the input parameters, check whether their input parameters are completely consistent. For example, if two questions "How many live shopping orders are there this month?" and "What is the total number of live shopping orders this month?" point to the same API, then their input parameters should be completely consistent.
[0239] Evaluate function consistency: Verify that all similar questions call the same API function. That is, even though the input parameters may differ slightly, the API function should be the same. For example, "What is the number of live shopping orders this month?" and "What is the total number of live shopping orders in the past month?" describe different time ranges, but both call the operating indicator query interface.
[0240] Evaluate output consistency: Check whether the output results for similar queries are consistent. Even if there are reasonable differences in the input parameters, the output results should be logically consistent. For example, two queries querying the number of live shopping orders "this month" and "the past month" should theoretically return the same result.
[0241] If any non-compliant situation is found, it will be marked as non-compliant; otherwise, it will be marked as compliant. Result evaluation This invention mainly adopts the method of building a prompt to call the large model inference. The following is a prompt example:
[0242] You are an intelligent assistant that excels at evaluating the consistency of similar query call results. Based on the similar queries and their API call results provided below, evaluate whether the calls between these queries are consistent. Pay special attention to the following points:
[0243] 1. **Input consistency**: For questions that only involve synonym replacement or word order adjustment without changing the input parameters, check whether their input parameters are completely consistent.
[0244] 2. **Function consistency**: Confirm whether the API functions called by all similar questions are consistent.
[0245] 3. **Output Result Consistency**: Check whether the output results of similar questions are consistent. Even if there are reasonable differences in the input parameters, the output results should remain logically consistent.
[0246] Similar questions and their API call results:
[0247] Question 1: “How many live shopping orders did you receive this month?”
[0248] API name: Business indicator query interface
[0249] Calling parameters:
[0250] - Time range (start date: 2024-10-01, end date: 2024-10-31)
[0251] -Brand ID: 123456
[0252] -Order Type: Live Shopping
[0253] Return result:
[0254] -Order quantity: 500
[0255] Question 2: "What is the total number of live shopping orders this month?"
[0256] API name: Business indicator query interface
[0257] Calling parameters:
[0258] - Time range (start date: 2024-10-01, end date: 2024-10-31)
[0259] -Brand ID: 123456
[0260] -Order Type: Live Shopping
[0261] Return result:
[0262] -Order quantity: 500
[0263] Please evaluate the following and return the result:
[0264] 1. Are the input parameters consistent?
[0265] 2. Are the called API functions consistent?
[0266] 3. Are the output results consistent?
[0267] 4. If there is a difference, directly output the evaluation result, yes or no, without providing other irrelevant text.
[0268] Step S3, perform sampling inspection based on the generated API call data, and use the sampling results as samples, re-label them through the result inspection agent, and select the API call data that meets the requirements as the final training data set, including:
[0269] After randomly sampling samples based on the labeled API call data for verification, a set of compliant samples and a set of non-compliant samples are constructed based on the sampling results. The sets are input into the Prompt template for consistency detection in step S26 for re-evaluation, and the non-compliant samples are removed based on the evaluation results to generate the final training data set.
[0270] Specifically, based on the above step S2, all generated question-API-call API result instances are marked and then handed over to business experts and technical experts for sampling inspection. Each question is carefully read to determine whether it conforms to the questioning habits in real scenarios, and whether the API call results obtained after matching with the recalled API results designed for it meet the actual results. The results of manual verification are used as samples, and the result inspection agent is again handed over to readjust the marking results of all generated instances. Finally, only the instances marked as conforming are retained as the final training data set.
[0271] The labeled instance set, that is, all labeled question-API-API call result instances, is submitted to business experts and technical experts for sampling inspection. Samples are randomly selected from the instance set, and each question is carefully read and evaluated for its compliance.
[0272] Construct a new tag set based on expert verification results Input the qualified samples and the unqualified samples into Prompt as Few-Shot. The input example is:
[0273] You are an intelligent assistant that excels at evaluating the consistency of call results for similar queries. Based on the similar queries and their API call results provided below, evaluate whether the calls between these queries are consistent. ...
[0275] The following is example output:
[0276] Question 1: ...
[0277] Question 2: ...
[0278] Result: consistent ...
[0280] Please evaluate the following and return the result:
[0281] 1. Are the input parameters consistent?
[0282] 2. Are the called API functions consistent?
[0283] 3. Are the output results consistent?
[0284] 4. If there are any differences, please explain the reasons for the differences.
[0285] Based on the updated results, the agent relabels the results in step 2.4 and removes samples marked as non-compliant, retaining only compliant samples.
[0286] In summary, the multi-agent collaboration framework in this embodiment constructs an API call data set, and constructs a collaborative process and network composed of multiple agents. Each agent has its own role and task. This multi-agent collaboration architecture is not limited to the implementation scheme mentioned in the above invention. For example, in terms of specific implementation details, such as the question-based expansion agent in step 2.2, you can also choose to use a small model solution for training intent classification and entity recognition to replace the extraction of question intent and entities. The question-based expansion agent can greatly enrich the diversity of training data, and combined with the result evaluation agent, it ensures that the generated data is highly realistic and representative; and by pre-calculating the similarity and classifying the API, various recall situations in real scenarios are simulated, thereby improving the response capability of subsequent LLM in actual applications.
[0287] Second embodiment
[0288] Based on the same inventive concept, the present invention also provides a multi-agent API call data generation system based on a large language model, which adopts the multi-agent API call data generation method as described above, including:
[0289] The scenario knowledge base construction module is used to sort out the API collection in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, and classify API associations based on the similarity to form a scenario knowledge base;
[0290] The agent collaborative generation module is used to design a multi-agent collaborative mechanism based on the scenario knowledge base to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, and an API matching agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, and the API recall agent simulates multiple recall situations. The matching agent maps the questions to specified API calls, and the result checking agent verifies the correctness and consistency of the API calls.
[0291] The data verification and screening module is used to perform sampling inspection based on the generated API call data, and re-label the sample based on the sampling results through the result inspection agent, and screen the qualified API call data as the final training data set.
[0292] Furthermore, the scenario knowledge base construction module includes,
[0293] The Schema knowledge construction unit is used to organize the API set in the target application scenario according to a predefined format and build a corresponding structured Schema knowledge system. The structured Schema knowledge includes the API interface name, interface description, API input parameters, and API output parameters.
[0294] The API association classification unit converts the interface descriptions in the API set and the corresponding question list into vector representations through embedding technology; based on the vector representation, the similarity score between each pair of APIs is calculated through cosine similarity, and the APIs are classified according to the similarity score and the classification criteria of the association relationship, including similar categories, unrelated categories and general similar categories; based on the classification results, a scenario knowledge base containing structured Schema knowledge is generated.
[0295] Furthermore, the agent collaborative generation module includes:
[0296] The user question agent module is used to randomly generate context-appropriate questions based on the scenario knowledge base and the large language model as the initial question set;
[0297] The question expansion agent module is used to perform variant operations based on the initial question set. Variant operations include synonym conversion and grammatical inflection. It also includes using natural language processing technology to analyze and generate equivalent but differently expressed questions in the initial question set to expand the initial question set.
[0298] The API recall agent module is used to identify the keywords and core intent involved in the questions based on the questioning and structured information in the schema knowledge in the initial question set. It performs semantic matching through the natural language understanding capabilities of the large language model and searches for matching APIs based on the classification results in step S1.
[0299] The API matching agent module is used to generate multiple sets of problem-API list-API recall instructions for recall situations, including unrecalled APIs, recalled API errors, and recalled multiple similar APIs;
[0300] The result checking agent module is used to verify the API call data according to the preset standards, and build a general prompt template based on each set of questions that meet the preset standards and the corresponding API list, input it into the large language model, and output the corresponding API call instructions.
[0301] The various prompt examples described in this embodiment are only preferred implementations of the present invention. It should be pointed out that, without departing from the principles of the present invention, several adaptation improvements and refinements may be made to the prompt, which should also be considered within the scope of protection of the present invention.
[0302] The implementation method of the solution described in this embodiment is to generate training data for processing API calls. It should be noted that without departing from the principles of the present invention, this technical architecture can also support the generation of training data such as NL2SQL and document question and answer, which should also be considered as the scope of protection of the present invention.
[0303] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
[0304] It should be noted that the above embodiments can be freely combined as needed. The above description is only a preferred embodiment of the present invention. It should be pointed out that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for generating multi-agent API call data based on a large language model, characterized in that: include, Step S1: sort out the API set in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, classify the API association relationships according to the similarity, and form a scenario knowledge base; Step S2: design a multi-agent collaborative mechanism based on the scenario knowledge base to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, an API matching agent, and a result checking agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, the API recall agent simulates multiple recall situations, and the matching agent maps the questions to the specified API calls, and the result checking agent verifies the correctness and consistency of the API calls. In step S3, a sampling inspection is performed based on the generated API call data, and the sampling results are used as samples. The intelligent agent is re-labeled through the result inspection, and the API call data that meets the requirements are selected as the final training data set.
2. The multi-agent API call data generation method according to claim 1, characterized in that: The step S1 includes: The API set in the target application scenario is sorted out according to a predefined format, and the corresponding structured Schema knowledge system is constructed, wherein the structured Schema knowledge includes API interface name, interface description, API input parameters and API output parameters.
3. The multi-agent API call data generation method according to claim 2, characterized in that: In step S1, the similarity between APIs is calculated by embedding technology, and the API association relationships are classified according to the similarity to form a scenario knowledge base, including: By using the embedding technology, the interface description in the API set and the corresponding method list are converted into vector representation; Calculating the similarity score between each pair of the APIs by cosine similarity based on the vector representation, and classifying the APIs according to the similarity score and the classification criteria of the association relationship, including similarity category, unrelated category and general similarity category; Based on the classification results, a scenario knowledge base containing the structured Schema knowledge is generated.
4. The multi-agent API call data generation method according to claim 3, characterized in that: In step S2, the user questioning agent generates a corresponding random question, and the questioning expansion agent performs a preset variant operation on the question, including: Step S21: the user questioning agent randomly generates the questions that are in line with the context based on the scenario knowledge base and the large language model as an initial question set; Step S22, based on the initial question set, the variant operation is performed by the question expansion agent through the question method, and the variant operation includes synonym conversion and grammatical deformation. At the same time, it also includes analyzing and generating equivalent but differently expressed questions in the initial question set through natural language processing technology to expand the initial question set.
5. The multi-agent API call data generation method according to claim 4, characterized in that: After simulating multiple recall situations through the API recall agent, the matching agent maps the problem to the specified API call and the result check agent verifies the correctness and consistency of the API call, including: Step S23: identifying keywords and core intents involved in the questions based on the questioning in the initial question set and the structured information in the Schema knowledge; Perform semantic matching through the natural language understanding capability of the large language model, and search for the matching API based on the classification result in step S1; and form multiple sets of problem-API list-API recall description instance sets for recall situations including unrecalled APIs, recalled API errors, and recalled multiple similar APIs; Step S24, checking the API call data according to the preset standard, and constructing a universal Prompt template according to each group of questions that meet the preset standard and the corresponding API list, inputting it into the large language model, and outputting the corresponding API call instruction.
6. The multi-agent API call data generation method according to claim 5, characterized in that: The checking agent checks the instance of the API call according to preset standards, including: Step S25: Verify the syntax specification, input parameter rationality, parameter integrity and response structure correctness of the API call data; wherein, the result check agent verifies the API call data according to the preset Schema knowledge standard, and the verification process is formalized as follows: ; in, Represents the API call data Whether it meets the standards of the schema knowledge, 1 means it meets the standards, and 0 means it does not meet the standards; Step S26, based on the API call data that meets the standards of the Schema knowledge, through the question set expanded by the expanded agent in step S22, construct the Prompt template for consistency detection to call the large model reasoning for consistency verification and mark it.
7. The method for generating multi-agent API call data according to claim 6, characterized in that: Perform sampling inspection based on the generated API call data, and use the sampling results as samples, re-label the intelligent agent through the result inspection, and select the API call data that meets the requirements as the final training data set, including: After randomly sampling samples based on the labeled API call data for verification, a set of compliant samples and a set of non-compliant samples are constructed based on the sampling results, and the Prompt template for the consistency detection in step S26 is input for re-evaluation, and the non-compliant samples are removed based on the evaluation results to generate the final training data set.
8. A multi-agent API call data generation system based on a large language model, using the multi-agent API call data generation method according to any one of claims 1 to 7, characterized in that: include, The scenario knowledge base construction module is used to sort out the API set in the target application scenario, build structured Schema knowledge, calculate the similarity between APIs through embedding technology, classify the API association relationships based on the similarity, and form a scenario knowledge base; An agent collaborative generation module is used to design a multi-agent collaborative mechanism based on the scenario knowledge base to dynamically generate API call data associated with the target application scenario; the multi-agent includes a user questioning agent, a question expansion agent, an API recall agent, and an API matching agent. The user questioning agent generates corresponding random questions, the question expansion agent performs preset variant operations on the questions, the API recall agent simulates multiple recall situations, and the matching agent maps the questions to the specified API calls, and the result checking agent verifies the correctness and consistency of the API calls; The data verification and screening module is used to perform sampling inspection based on the generated API call data, and re-label the intelligent agent based on the sampling results as samples, and screen the API call data that meets the requirements as the final training data set.
9. The multi-agent API call data generation system according to claim 8, characterized in that: The scenario knowledge base construction module includes: A Schema knowledge construction unit is used to sort out the API set in the target application scenario according to a predefined format and construct the corresponding structured Schema knowledge system, wherein the structured Schema knowledge includes the API interface name, interface description, API input parameters, and API output parameters; The API association classification unit converts the interface descriptions in the API set and the corresponding question lists into vector representations through the embedding technology; calculates the similarity scores between each pair of APIs through cosine similarity based on the vector representation, and classifies the APIs according to the similarity scores and the classification criteria of the association relationships, including similar categories, irrelevant categories and general similar categories; and generates a scenario knowledge base containing the structured Schema knowledge based on the classification results.
10. The multi-agent API call data generation system according to claim 9, characterized in that: The intelligent agent collaborative generation module includes: A user questioning agent module is used to randomly generate the questions that are in line with the context based on the scenario knowledge base and the large language model as an initial question set; a question expansion agent module, configured to perform variant operations based on the initial question set, wherein the variant operations include synonym conversion and grammatical inflection, and further include analyzing and generating equivalent but differently expressed questions in the initial question set through natural language processing technology, thereby expanding the initial question set; An API recall agent module is configured to identify keywords and core intents involved in the questions based on the questioning in the initial question set and the structured information in the Schema knowledge; Performing semantic matching through the natural language understanding capability of the large language model, and searching for the matching API based on the classification result in step S1; The API matching agent module is used to generate multiple sets of problem-API list-API recall instructions for recall situations, including unrecalled APIs, recalled incorrect APIs, and recalled multiple similar APIs; The result checking intelligent agent module is used to check the API call data according to preset standards, and to construct a universal Prompt template based on each group of questions that meet the preset standards and the corresponding API list, input it into the large language model, and output the corresponding API call instructions.