Calling method and device of language tool, electronic equipment and storage medium

By constructing an index data structure and a multi-stage routing strategy, the problems of low efficiency and high maintenance when accessing large-scale language models via large-scale APIs are solved, achieving efficient and accurate API calls and automated management, which is suitable for complex cloud platforms.

CN121597828BActive Publication Date: 2026-04-07JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies suffer from low API call efficiency and high operation and maintenance costs when integrating large-scale APIs into large language models. This is mainly due to the inability to accurately select and call APIs and the inability to automatically adapt to dynamic changes in APIs.

Method used

By constructing an indexed data structure, tool cards are generated based on the user's original language query, and the target language tool is selected. Multi-stage routing strategies and parameter completion modules are used to ensure accurate and automated calls, thereby reducing operation and maintenance costs.

Benefits of technology

It achieves efficient and low-cost conversion from natural language queries to API calls, improves API call efficiency and accuracy, reduces system operation and maintenance costs, and is suitable for private cloud platforms with a large number of interfaces and complex functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597828B_ABST
    Figure CN121597828B_ABST
Patent Text Reader

Abstract

The application discloses a language tool calling method and device, electronic equipment and storage medium. The method comprises the following steps: constructing an index data structure of a protocol document, quickly responding to a user's input original language, screening a plurality of candidate language tools from a large-scale language tool, and reducing the cost of a large language model processing full information; the generation mechanism of the tool card further reduces the key information of the candidate language tool. Finally, based on the execution request generated by the target language tool and the target parameter corresponding to the target tool card, the target language tool is called through the execution request. Through the method, the technical problems of low API calling efficiency and high operation and maintenance cost caused by accessing a large-scale API to a large language model in the related art are solved, accurate mapping from natural language query to API calling is realized, and the technical effects of improving the efficiency and accuracy of language tool calling are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for invoking a language tool. Background Technology

[0002] With the capabilities demonstrated by Large Language Models (LLMs) in natural language understanding and task planning, combining them with language tools (such as Application Programming Interfaces) to achieve automated operations and intelligent services has become a significant trend in technological development. APIs are a means of controlling and monitoring massive resources such as computing, networking, and storage. Therefore, integrating large-scale APIs into LLMs enables cloud resource management through natural language, greatly reducing the operational complexity for maintenance personnel. While LLMs have shown great potential, accurately selecting and calling from a large number of APIs is a challenging task when integrating them into LLM's natural language processing. The speed and accuracy of selection are difficult to guarantee. Furthermore, due to the high frequency of API updates, manually developed adapter functions or tools relying on unstructured text retrieval cannot automatically adapt to API changes, resulting in high maintenance costs.

[0003] In related technologies, solutions based on general search enhancement or fixed adapters are commonly used to alleviate the problems of excessively long context windows and maintenance. However, these methods only focus on solving the problems of excessively long context windows and maintenance, without considering the efficiency and accuracy of API calls. Therefore, this leads to low API call efficiency and high maintenance costs when integrating large-scale APIs into large language models. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for invoking language tools, to at least solve the problems of low API call efficiency and high operation and maintenance costs caused by connecting large-scale APIs to large language models in related technologies.

[0005] This application provides a method for invoking a language tool, comprising: constructing an index data structure based on a specification document; querying the index data structure based on the original language input by the user to determine multiple candidate language tools; generating multiple tool cards for the multiple candidate language tools, and determining the target language tool corresponding to the target tool card from the multiple tool cards, wherein one tool card is generated for each candidate language tool; generating an execution request based on the target language tool and target parameters, and invoking the target language tool based on the execution request.

[0006] This application also provides a language tool invocation device, comprising: a construction module for constructing an index data structure according to a specification document; a query module for querying the index data structure based on the user-inputted raw language to determine multiple candidate language tools; a generation module for generating multiple tool cards for the multiple candidate language tools and determining the target language tool corresponding to the target tool card from the multiple tool cards, wherein one tool card is generated for each candidate language tool; and an invocation module for generating an execution request based on the target language tool and target parameters, and invoking the target language tool based on the execution request.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the calling method of any of the above-described language tools when executing the computer program.

[0008] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of calling the language tool described above.

[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of calling any of the above-described language tools.

[0010] This application, by constructing an index data structure for specification documents, enables rapid response to the user's raw language input and precise selection of a candidate set (multiple candidate language tools) from a large-scale set of tools, reducing the high cost of processing full-scale information in large language models. Furthermore, the tool card mechanism further narrows down the key information of candidate language tools, allowing them to make optimal decisions within a limited context and reducing token consumption during invocation. Finally, based on the target language tool and target parameters corresponding to the target tool card, an execution request is generated to invoke the target language tool, ensuring accuracy and automation while reducing system maintenance costs and improving the efficiency and accuracy of language tool invocation. Therefore, it solves the technical problems of low API call efficiency and high maintenance costs caused by integrating large-scale APIs into large language models, achieving accurate mapping from natural language queries to API calls, thereby improving cloud resource management efficiency and reducing maintenance costs. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a hardware structure block diagram of a computer terminal for a method of invoking a language tool according to an embodiment of this application;

[0013] Figure 2 This is a system architecture diagram of a cloud platform-based language tool invocation according to an embodiment of this application;

[0014] Figure 3 This is a flowchart of a method for invoking a language tool according to an embodiment of this application;

[0015] Figure 4 This is a schematic diagram of data flow and control flow based on language tool calls according to an embodiment of this application;

[0016] Figure 5 This is a schematic diagram illustrating lifecycle management and hot update based on language tool invocation according to an embodiment of this application;

[0017] Figure 6 This is a schematic diagram of an exception handling and rollback strategy based on language tool invocation according to an embodiment of this application;

[0018] Figure 7 This is a structural block diagram of a language tool invocation device according to an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] The specific application environment architecture or specific hardware architecture on which the execution of the calling methods of the language tools depends is described here.

[0023] First, some nouns or terms that appear in the explanation of the embodiments of this application shall be interpreted as follows:

[0024] (1) API (Application Programming Interface): A contract used to connect different software components or systems, defining the types, methods, and required data formats for calls or requests that can be made.

[0025] (2) LLM (Large Language Model): A deep learning model that can understand and generate natural language after being trained on large-scale data.

[0026] (3) Swagger / OpenAPI: An API specification standard for describing, producing, consuming and visualizing RESTful Web services. In this application embodiment, it is mainly used to understand the structure and function of API.

[0027] (4) HTTP (Hypertext Transfer Protocol): The most widely used network protocol on the Internet, used for communication between clients and servers.

[0028] (5) UUID (Universally Unique Identifier): A 128-bit number used to uniquely identify information in a computer system.

[0029] (6) IP (Internet Protocol): A protocol used to exchange data packets between networks, forming the foundation of the Internet.

[0030] (7) CIDR (Classless Inter-Domain Routing): An address deduction method for assigning IP addresses to users and for effectively routing IP packets on the Internet.

[0031] (8) JSON (JavaScript Object Notation): A lightweight data exchange format that is easy to read and write, and easy for machines to parse and generate.

[0032] (9) REST (Representational State Transfer): A software architecture style for distributed hypermedia systems, and the mainstream design pattern for current Web services.

[0033] (10) RAG (Retrieval-Augmented Generation): A technique that combines external knowledge base retrieval with large language model generation to improve the accuracy and timeliness of model responses.

[0034] The methods and embodiments provided in this application can be executed on server devices, mobile terminals, computer terminals, or similar computing devices. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for a method of invoking a language tool according to an embodiment of this application. Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor unit (MPU), or programmable logic device (PLD)) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0035] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the language tool invocation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thus implementing the above-described methods. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0037] Before explaining the technical solution of this application, we will first briefly introduce the technical problems that exist when integrating large-scale APIs into LLM.

[0038] Currently, when integrating large-scale API ecosystems (such as private cloud platforms with thousands of interfaces) into LLM, the following technical challenges are encountered:

[0039] (1) Context window limitation and token cost issues: Current LLMs (such as the GPT (Generative Pre-trained Transformer) series) all have limitations on the length of the context window. For example, the API documentation of a private cloud platform (such as Swagger / OpenAPI documentation) may contain tens of thousands or even hundreds of thousands of tokens. Inputting all the large-scale API information into the single request context of the LLM will not only exceed the window limit, but also generate extremely high token costs;

[0040] (2) API selection and invocation efficiency and accuracy issues: Even if API information is provided to the LLM, the efficiency and accuracy of the LLM in selecting from a massive number of APIs are difficult to guarantee. When faced with hundreds of APIs with similar functions but different parameters, the LLM is prone to "illusion", making incorrect API selections, confusing parameters, or "hesitating" between multiple options, resulting in task failure or inefficiency. As the number of APIs increases, the probability of task failure increases exponentially.

[0041] (3) The dynamic nature and maintenance cost of APIs: With the continuous and rapid iteration of cloud platform technology stack, APIs are frequently added, modified or deprecated. Manual API access, documentation maintenance or adaptation layer code writing will bring huge maintenance burdens. Related technologies often lack an automated mechanism to adapt to such changes, resulting in the API information on which LLM depends being outdated and the system being unreliable.

[0042] The following two types of solutions are mainly adopted for the above-mentioned technical problems:

[0043] 1. A scheme based on general search enhancement: This scheme treats the descriptive documents of APIs (such as function summaries and parameter descriptions) as unstructured text and constructs a vector database. When a user submits a query, the system first vectorizes the user's query, performs a semantic similarity search in the database, finds the most relevant API text fragments, and then submits these API text fragments along with the user's query to the LLM. The LLM then selects and invokes the API based on these "enhanced" contexts.

[0044] Limitations: While this approach alleviates the problem of excessively long contexts to some extent, its core weakness lies in "flattening" the structured API descriptive documentation into unstructured text, resulting in the loss of a significant amount of crucial information. For example, it fails to accurately understand structured constraints such as parameter types, required status, enumeration value ranges, and mutual exclusion relationships between parameters. This makes LLM prone to errors during the parameter population phase. Furthermore, semantic similarity matching has limited ability to differentiate APIs with highly overlapping functions but subtle differences in key parameters (e.g., one queried by ID, the other by name), leading to a lower quality API candidate set for route recall.

[0045] 2. Fixed Adapter or SDK (Software Development Kit) Based Solution: This solution pre-develops a fixed set of high-level SDK or adapter functions for the LLM. The LLM does not directly interact with the underlying API but instead calls the encapsulated adapter functions. For example, frameworks such as LangChain and LlamaIndex allow developers to manually define Tool objects.

[0046] Shortcomings: This solution requires manual modification or redevelopment of the adapter whenever the underlying API changes. For systems with hundreds or thousands of APIs, this amount of manual encapsulation is unacceptable. It shifts the complexity of maintenance from the context management of the LLM to the software engineering management of the adaptation layer, without fundamentally solving the technical problem.

[0047] In summary, existing technologies either suffer from insufficient accuracy due to their inability to fully utilize the structured information of APIs, or suffer from poor scalability and maintainability due to their reliance on manual encapsulation. Therefore, there is an urgent need for a solution that can automatically parse and index large-scale API specifications and accurately map natural language queries to API calls through efficient and low-cost routing strategies.

[0048] In view of the above problems, this application provides a method for selecting and invoking natural language tools on a cloud platform based on Swagger / OpenAPI indexing and multi-stage methods. By constructing a two-layer tool invocation system, it decouples the atomic APIs at the bottom layer of the cloud platform from external tools facing the upper-layer agent. The agent is a large language model or proxy program responsible for understanding user intent and performing task planning. This two-layer tool invocation system ensures that the agent only needs to interact with a small number of stable and semantically clear external tools, without needing to be aware of the specific information of hundreds or even thousands of underlying APIs. The complex problem of "selecting a large language model from a large number of APIs" is decomposed into a series of processes: "lightweight model / stepwise filtering" and "selecting a large language model from a small number of APIs." This method, based on Swagger / OpenAPI indexing and multi-stage cloud platform natural language tool selection and invocation, not only significantly reduces the cost of a single call to a large language model, but also ensures the system's adaptability to dynamic changes in cloud platform APIs through automated indexing and continuous learning. This addresses issues such as long context windows, high token consumption, low API call efficiency, and high operation and maintenance costs that arise when connecting large-scale APIs to large language models in related technologies.

[0049] Specifically, this application provides a natural language tool invocation system for cloud platforms. Through an automated API indexing mechanism and a multi-stage intelligent routing strategy, it achieves efficient and low-cost conversion from natural language queries to accurate API calls, and is applicable to private cloud platforms with a large number of interfaces and complex functions. Figure 2 This is a system architecture diagram of a cloud platform-based language tool invocation according to an embodiment of this application, such as... Figure 2 As shown, the system architecture includes: an index building module, a tool routing module, a parameter completion module, a tool execution module, a caching and strategy module, and a monitoring and learning module.

[0050] (1) Index building module

[0051] The index building module forms the foundation for automated API access and understanding. It can automatically parse the API specification documents of the cloud platform (such as Swagger / OpenAPI format specification documents). Through deep parsing of the API specification documents, it generates a series of index data structures for subsequent retrieval, selection, and invocation. Specifically, when the system is deployed for the first time or when an update to the API specification document is detected, the index building module automatically parses and generates the following index data structures:

[0052] a. Atomic Interface Registry: A structured database that records specific API information parsed and extracted from the Swagger / OpenAPI specification document, including but not limited to HTTP request methods (such as GET, POST), request paths, function summaries, detailed descriptions, parameters (such as name, type, required status, description, enumeration values, etc.), and response models. This atomic interface registry is the "single source of truth" for internal system operations. To handle naming conflicts that may arise from multiple Swagger sources, the API ID (tool_id) is normalized using the "service.tag.operationId" format, and conflict resolution rules are set; furthermore, to support version management, the atomic interface registry stores the "schema_version" and "api_version" for each API.

[0053] b. Inverted Index: A keyword retrieval data structure that stores the mapping between keywords extracted from API details (such as function summary, request path, detailed description, parameters, etc.) and their corresponding API identifiers. For example, the keyword "virtual machine" can be mapped to multiple API identifiers such as "compute.list_instances" and "compute.get_instance_details". Using this inverted index, the system can quickly find relevant candidate APIs based on keywords in the user's natural language query, avoiding a full traversal.

[0054] c. Domain Dictionary Mapping Table: A semantic mapping table that associates informal and diverse business terms in user natural language queries with standard cloud platform service domains or entities. For example, expressions such as "cloud host," "virtual machine," and "VM" in user natural language can be mapped to the standard "Compute Domain" through the domain dictionary mapping table, which helps to quickly narrow down the search scope for APIs in the early stages of API selection.

[0055] d. Regular Expression Library: A predefined set of regular expressions used to identify and extract entities with fixed formats from user natural language query text, such as IP addresses (IPv4 / IPv6), CIDR blocks, UUIDs, operating system names (e.g., CentOS, Ubuntu), resource statuses (e.g., running, stopped), etc.

[0056] (2) Tool routing module

[0057] The API routing module accurately filters a small number of candidate APIs that meet the user's natural language query requirements. It employs a multi-stage, progressively filtering "funnel-like" routing strategy, replacing the traditional approach of directly providing all API information to a large language model for routing. Specifically, it includes the following stages:

[0058] Phase 1: Service Domain and Tag Filtering. This phase uses a domain dictionary mapping table and a regular expression library to initially filter user-input natural language queries, identifying the core service domains (such as computing, networking, and storage) and key tags involved in the query. This phase is implemented using a lightweight classification model (e.g., a Naive Bayes model) or pre-defined decision rules, narrowing down the candidate API set from hundreds to dozens or even just a few.

[0059] Phase Two: Inverted Index Recall and Lightweight Reordering. Within the "candidate API set" selected in Phase One, keyword recall is performed using an inverted index table. Specifically, after segmenting the natural language query, matching APIs are searched in the inverted index table. The recall strategy combines intersection-first (i.e., matching multiple keywords for high precision) and union-last (i.e., matching any keyword for high recall). The recalled API list is sorted using a lightweight reordering algorithm. This algorithm calculates a relevance score for each candidate API in the list and sorts them according to their relevance scores. The system can select K (K∈[3,5]) candidate APIs and set a confidence threshold τ∈[0.8, 1.0]. If the highest relevance score of a candidate API is still lower than the confidence threshold, it is marked as low confidence and a clarification suggestion is generated.

[0060] Phase 3: Generation of Candidate Tool Cards. The tool routing module selects the K candidate APIs with the highest relevance scores and generates a standardized "tool card" for each candidate API. The tool card can contain the following fixed fields: "{tool_id, title, one_line, required_params, score, why, notes:{regex:{}}, low_confidence, clarify_suggestion}". This tool card is a highly condensed summary of the API information, containing the minimum information set required for agent decision-making.

[0061] (3) Parameter completion module

[0062] After the agent selects a tool (API) from K candidate API tool cards, the parameter completion module automatically fills in the parameters required by the tool from the original user language query.

[0063] Automatic parameter completion: Using a domain dictionary mapping table and a regular expression library, the parameter completion module extracts values ​​from the user's language query text that match the selected API parameter. For example, if the API requires an "ip_address" parameter, the module uses a regular expression for IP addresses to find a matching value in the user's language query. The module also handles parsing time expressions (such as "last week" or "today 9 AM - 6 PM") and numerical ranges (such as "CPU ≥ 4"). The built-in parser uses deterministic regular expressions / dictionaries and follows set time zone rules; if parsing fails, it falls back to a clarification interaction. It's important to note that when a parameter has preset enumerated values ​​(e.g., the "status" parameter can be active / down), the module only maps synonyms or approximate values ​​within the enumeration set. If the confidence level of the approximate mapping is below a preset threshold, automatic completion fails, triggering a clarification interaction to ensure parameter accuracy.

[0064] Clarifying Interaction: After auto-filling parameters, if any required parameters are still missing, the parameter completion module does not immediately return failure. Instead, it intelligently generates a single clarifying question for the user, avoiding inefficient multi-round dialogues. For example, if the `create_instance` interface lacks the "image_id" parameter, the parameter completion module generates the question: "Which image would you like to use to create the cloud server? You can provide the image ID or name," and can optionally include several commonly used images as options. This proactive guidance significantly improves the user experience and task success rate. It should be noted that clarification is applied to required parameters that cannot be auto-filled; if the field has enumerated values, up to 5 candidate values ​​can be provided during the clarification process; when multiple fields are missing, a single clarifying question is generated based on the highest priority for gain.

[0065] Confidence assessment: The system generates a confidence score for each filled parameter. This confidence score is determined based on the filling method (e.g., the confidence score of an exact match using a regular expression is higher than that of a fuzzy match using synonyms). This score can be used as a basis for the agent to determine whether it is necessary to confirm with the user again.

[0066] (4) Tool execution module

[0067] The tool execution module is the endpoint that interacts with the underlying API of the cloud platform. It receives an execution request containing the API identifier and filled parameters, and executes the following process:

[0068] Parameter validation and permission checks: Before execution, the parameter type, format, and value range are finalized and validated. The system also checks whether the user initiating the call has the necessary permissions to access the API. The validation logic is based on SwaggerSchema, supporting not only query parameters but also comprehensive validation of complex parameter structures and types, such as "application / json" bodies, "multipart / form-data," arrays, and nested objects.

[0069] REST Calls and Pagination Processing: HTTP requests are constructed and sent to the cloud platform's API gateway. For query APIs that return list data, the tool execution module automatically handles pagination logic and has built-in explicit control mechanisms, such as setting `page_limit` (maximum page count), `max_items` (maximum number of items), and `early_stop` (early stop condition). In rate-limiting scenarios such as "429 Too Many Requests," the system employs exponential backoff and jitter strategies for retrying. Simultaneously, a total timeout and concurrency budget are set across the entire chain. This budget is dynamically adjusted based on user quotas and backend service SLOs to prevent resource abuse and ensure system availability. When budget protection is triggered, the system returns a structured error containing "error_code: RATE_LIMIT" and "suggestion" fields (e.g., suggesting narrowing filter conditions or using asynchronous tasks).

[0070] Multi-tool aggregation execution: Supports the execution of predefined "high-level aggregation tools". These high-level aggregation tools encapsulate a complex workflow consisting of multiple atomic API calls. For example, a request to query "the usage of all IPs in the network where a certain cloud host is located" may correspond to an aggregation tool. This tool first calls the API to query the network port information of the cloud host, then calls the API to query the subnet information based on the network port information, and finally summarizes the data and outputs the execution result.

[0071] To ensure consistency across services, execution results adopt a unified structure, such as {ok: boolean, data: {items: [], stats: {…}}, meta: { source: [tool_id…], page, pageSize, total, elapsed_ms, trace_id}, error: { error_code, message, trace_id}}. The "source" field records the identifier of the underlying / aggregation tool involved in generating the execution results, for traceability.

[0072] (5) Caching and Strategy Module

[0073] To improve system response speed and further reduce costs, an intelligent caching mechanism was designed.

[0074] Static content caching: Long-term caching of API descriptions, tool cards, parameter filling suggestions, and other content that does not change frequently;

[0075] Dynamic content caching: API call results are cached temporarily. To ensure data isolation and security in a multi-user environment, the cache key explicitly includes user and permission information. When verifying a token, the system extracts the policy digest "policy_hash", and the cache key is expanded to "tenant_id|role|policy_hash|tool_id|params_hash|schema_version".

[0076] When the "hash" changes, the associated cache becomes invalid immediately to prevent unauthorized cache hits. Any cache hit across users or roles is considered invalid to prevent data pollution. The cache's TTL (Time-To-Live) can be adaptively adjusted according to the frequency of API data changes.

[0077] (6) Monitoring and Learning Module

[0078] The monitoring and learning module records key logs during system operation, including but not limited to user queries, user final selections, parameter completion success rates, API execution time and results, etc. Through offline analysis of these key log data, the system can achieve self-optimization and evolution, for example:

[0079] Automatically discover new synonyms or business terms and add them to the domain dictionary mapping table;

[0080] Adjust the scoring weights of the reordering algorithm in the tool's routing module based on user selection patterns;

[0081] Identify a list of frequently combined APIs to provide data support for creating new high-level aggregation tools.

[0082] Through the collaborative work of the aforementioned modules, this embodiment constructs a closed loop from parsing specification documents, tool routing, parameter completion to tool execution. Specifically, the system exports three core tools to the external MCP (Model Context Protocol) Server side: tool_router.select, tool_param.fill, and tool_exec.run, along with a few necessary higher-level aggregation tools. The return value of the "list_tools()" interface is strictly constrained within this scope; the underlying atomic API interfaces exist only in the internal registry and cannot be directly enumerated or called by the client, thus achieving a minimally simplistic external interface and complex internal management.

[0083] The following example uses the method of calling language tools. Figure 3 This is a flowchart of a method for invoking a language tool according to an embodiment of this application, such as... Figure 3 As shown, the specific steps of this method include the following:

[0084] Step S302: Construct the index data structure according to the specification document;

[0085] In step S302 of this embodiment, the specification document (API specification document) includes, but is not limited to, the Swagger / OpenAPI document of the cloud platform, and the index data structure includes an atomic interface data table, an inverted index table, a domain dictionary mapping table, and a regular expression library. Through deep parsing of the specification document, a series of index data structures are generated for subsequent retrieval, selection, and invocation.

[0086] Step S304: Based on the original language input by the user, query the index data structure to determine multiple candidate language tools;

[0087] In step S304 of this embodiment, the index data structure is queried according to the original language input by the user (i.e., the original natural language query) to accurately filter out multiple candidate language tools (candidate APIs) that meet the requirements.

[0088] Step S306: Generate multiple tool cards for multiple candidate language tools, and determine the target language tool corresponding to the target tool card from the multiple tool cards, wherein one tool card is generated for each candidate language tool;

[0089] In step S306 of this embodiment, multiple tool cards are generated for multiple candidate APIs. For example, for the candidate language tools determined by the above query, the following candidate tool cards can be generated:

[0090] Card 1:

[0091] Tool ID: "compute.list_instances".

[0092] Function: Query and list all cloud servers that meet the criteria.

[0093] Required parameter: None.

[0094] Recommendation reason: Matched the keyword "cloud server".

[0095] Card 1 will be submitted to the agent, which will make a tool selection decision. The agent will select a target language tool (target API) from multiple tool cards and then automatically fill in the parameters required by the tool from the original natural language query.

[0096] Because the number of tool cards is small and the information is highly condensed, the context length and number of tokens required to pass to large language models are effectively controlled.

[0097] Step S308: Generate an execution request based on the target language tool and target parameters, and call the target language tool based on the execution request.

[0098] In step S308 of this embodiment, the ID of the target API and the target parameters (filled parameters) are encapsulated to generate an execution request, and the target API is called based on the execution request.

[0099] Through steps S302 to S308, by constructing an index data structure for the specification document, the system can quickly respond to the user's input in the original language and accurately filter out a candidate set (multiple candidate language tools) from a large-scale tool library, reducing the high cost of processing all information in a large language model. Furthermore, the tool card mechanism further narrows down the key information of candidate language tools, enabling them to make optimal decisions within a limited context and reducing token consumption during invocation. Finally, based on the target language tool and target parameters corresponding to the target tool card, an execution request is generated to invoke the target language tool, ensuring the accuracy and automation of the invocation while reducing system maintenance costs and improving the efficiency and accuracy of language tool invocation. Therefore, this solution addresses the technical problems of low API invocation efficiency and high maintenance costs caused by integrating large-scale APIs into large language models, achieving accurate mapping from natural language queries to API calls, thereby improving cloud resource management efficiency and reducing maintenance costs.

[0100] In one exemplary embodiment, constructing an index data structure based on a specification document includes: parsing the specification document of the cloud platform to obtain the specification structure and specification format of the corresponding language tool; and constructing an index data structure based on the specification structure and specification format.

[0101] In this embodiment, when the system is deployed for the first time or when an API specification document update is detected, the index building module is automatically executed to pull and parse the Swagger / OpenAPI document from the cloud platform. The specification structure and specification format correspond to the structure and format of the Swagger / OpenAPI document, as detailed below:

[0102] Generate an atomic interface registry: parse the API path (e.g., " / v2.1 / servers / detail"), method ("GET"), operationId ("listServers"), summary ("List servers"), description, parameters, and detailed attributes (type, required, description, etc.), and store them in a structured database. For example, the listServers interface is recorded in detail.

[0103] Build an inverted index table: Extract keywords and establish mappings. For example, the keywords "cloud host" and "server" are mapped to compute.servers.listServers; the keywords "status" and "status" are mapped to API.

[0104] Populate the domain dictionary mapping table: map terms such as "cloud host" and "virtual machine" to "computing domain".

[0105] Load the regular expression library: Ensure that your system contains regular expressions for identifying operating system versions (such as "CentOS 7.9"), IP addresses, UUIDs, etc.

[0106] Through the above embodiments, by deeply analyzing the Swagger / OpenAPI documentation and combining it with inverted index tables, domain dictionary mapping tables, and regular expression libraries, more accurate tool routing is achieved, thus improving the accuracy of tool routing.

[0107] In an exemplary embodiment, a query is performed on an index data structure based on the user-inputted raw language to determine multiple candidate language tools, including: performing a first query on a domain dictionary mapping table and a regular expression library based on the user-inputted raw language to obtain multiple initial language tools; and performing a second query on an inverted index table based on the multiple initial language tools to determine multiple candidate language tools.

[0108] In this embodiment, the query is performed based on the user's original language input, and a small number of candidate APIs that meet the requirements are efficiently and accurately filtered. A multi-stage, step-by-step filtering "funnel" routing strategy is adopted, replacing the traditional method of directly providing all API information to a large language model for routing.

[0109] The first query performs preliminary filtering on the user's original language query using a domain dictionary mapping table and a regular expression library, resulting in multiple initial APIs.

[0110] The second query involves using an inverted index table to retrieve keywords from the "multiple initial APIs" filtered by the first query, and then identifying multiple candidate APIs.

[0111] In an exemplary embodiment, based on the user-inputted raw language, a first query is performed on a domain dictionary mapping table and a regular expression library to obtain multiple initial language tools, including: preprocessing the raw language to obtain multiple keywords, wherein the preprocessing includes at least one of the following: letter conversion, mixed word segmentation, name splitting, and word count sliding window; determining the service domain of the raw language query based on the multiple keywords and the domain dictionary mapping table; and performing a format signal query based on the service domain and the regular expression library to obtain multiple initial language tools.

[0112] In this embodiment, Figure 4 This is a schematic diagram of data flow and control flow based on language tool calls according to an embodiment of this application, such as... Figure 4 As shown, the query processing flow includes standardization and signal extraction. When the system receives the user's raw language input (e.g., "find all online cloud hosts for CentOS 7.9"), it performs preprocessing of the raw language text, including converting letters to lowercase, segmenting Chinese and English words, splitting camelCase naming, and applying 2-3 character sliding windows for Chinese characters to avoid unreproducible parsing results, thus obtaining multiple keywords.

[0113] Based on keywords in the preprocessed text, the service domain to which the query belongs is determined by querying the domain dictionary mapping table. For example, the term "cloud host" points to "computing domain". This step narrows the scope of candidate APIs from the entire platform's API set to a subset of APIs only related to the computing domain, greatly reducing the computational complexity of subsequent steps.

[0114] At the same time, the entity "CentOS 7.9" is identified by the regular expression library, and "signals" with a clear format are extracted, such as the operating system "CentOS 7.9" and the status "online" in this embodiment, to obtain multiple initial APIs. The above signals will serve as an important basis for subsequent routing and parameter filling.

[0115] In an exemplary embodiment, based on multiple initial language tools, a second query is performed through an inverted index table to determine multiple candidate language tools, including: based on multiple initial language tools, a keyword query is performed through an inverted index table to generate a language tool list, wherein the language tool list includes multiple language tools; the multiple language tools are scored and sorted to determine multiple candidate language tools.

[0116] In this embodiment, keyword retrieval is performed using an inverted index table within the "multiple initial APIs" filtered by the first query. Specifically, after segmenting the natural language query, words such as "CentOS," "online," and "cloud host" are used to search for matching APIs in the inverted index table. The retrieved API list is sorted using a lightweight reordering algorithm, which calculates a relevance score for each candidate API in the API list. This score includes factors such as: keyword matching (API summary matching score is higher than parameter description matching), phrase matching bonus (continuous phrases from the original language query also appear consecutively in the API description), and regular expression signal hit bonus (API parameters match signals extracted from the query).

[0117] Through the above embodiments, after multi-stage screening, a lightweight model is used in the early stage to filter out most irrelevant APIs, providing a small number of highly relevant candidate APIs for LLM. This reduces the size of the contexts submitted to LLM from tens of thousands or even hundreds of thousands to a few hundred, achieving an order-of-magnitude reduction.

[0118] In an exemplary embodiment, scoring and ranking multiple language tools to determine multiple candidate language tools includes: weighting the multiple language tools according to a rearrangement formula to obtain a relevance score for each language tool; and ranking the language tools corresponding to the relevance scores if the relevance scores meet a first threshold to determine multiple candidate language tools.

[0119] In this embodiment, the rearrangement formula includes weighted scoring of at least one of the following: word frequency, keyword hit, abstract hit, parameter hit, word reward, and regularity reward.

[0120] In this embodiment, the recalled API list is sorted using a lightweight reordering algorithm. This algorithm calculates a relevance score for each candidate API in the API list, and the relevance score is determined by the following formula:

[0121] score = α × BM25 + β1 × name hit rate + β2 × summary hit rate + β3 × param hit rate + β4 × word reward + β5 × regular expression reward

[0122] Wherein, score is the relevance score for each candidate API, and α, β1, β2, β3, β4, and β5 are the initial weight values ​​for each scoring metric, for example, α=1.0, β1=2.0, β2=1.2, β3=1.5, β4=1.0, and β5=1.0.

[0123] The system can set a confidence threshold τ∈[0.8, 1.0], and sort the APIs that meet the confidence threshold according to their relevance scores. The system can select K (K∈[3,5]) candidate APIs.

[0124] Through the above embodiments, the reordering algorithm comprehensively considers multiple dimensions of indicators such as keywords, abstracts, and parameters, effectively distinguishing APIs with similar functions but different details, and ensuring that the candidate APIs provided to LLM are of high quality.

[0125] In one exemplary embodiment, generating an execution request based on a target language tool and target parameters includes: populating the target parameters required by the target language tool based on the user's input in the original language; encapsulating the identifier of the target language tool and the target parameters to generate the execution request.

[0126] In this embodiment, as Figure 4 The parameter completion process shown involves selecting a target API from multiple candidate APIs, automatically filling in the target parameters required by the target API from the original user language query, encapsulating the target API ID and target parameters, generating an HTTP request (execution request), and then standardizing the returned results before displaying them to the user.

[0127] For example, for a simple query, the system encapsulates the target API ID and target parameters and sends them to the tool execution module. After the tool execution module completes parameter validation and permission checks, it sends an HTTP request to the cloud platform API gateway and then processes the returned results in a standardized manner before displaying them to the user.

[0128] For complex queries, such as "list the number of cloud hosts running in each network", the tool's routing module can directly match a preset cloud.query.host_network_summary aggregation tool. The execution logic of this aggregation tool is encapsulated within the tool's execution module, enabling it to autonomously and sequentially call multiple atomic APIs (e.g., first calling "network.list_networks" to retrieve all networks, then calling "compute.list_instances" for each network and passing in the network ID and status=running parameter for filtering), and completing data aggregation on the server side before returning the results all at once.

[0129] In the absence of advanced aggregation tools, if the system determines that a single API cannot meet the user's needs, it attempts to perform dynamic chained call planning. Based on the API dependencies defined in the atomic interface registry, the system uses a graph search algorithm to generate a multi-step API call plan. This planning process follows explicit boundary conditions, including but not limited to maximum call depth, maximum fan-out, and total timeout. For failed operations that support retries, the system ensures safe retries through idempotent keys, and the generated plan is then submitted to the agent for confirmation or executed step-by-step by the tool execution module.

[0130] Through the above embodiments, the process of encapsulating the target API ID and target parameters to generate an execution request ensures the accuracy and security of the call, avoids parameter confusion or omission during transmission, and enhances the reliability and integrity of the execution request.

[0131] In one exemplary embodiment, populating the target parameters required by the target language tool based on the user-inputted raw language includes: determining the target value based on the user-inputted raw language through a domain dictionary mapping table and a regular expression library, and populating the target value into the target parameters required by the target language tool.

[0132] In this embodiment, the target value matching the target parameter is extracted from the user's original language query text using a domain dictionary mapping table and a regular expression library. For example, in the example of "find all online cloud hosts of CentOS 7.9", if the "compute.list_instances" tool has two optional parameters, os_name and status, the system will fill the target value "CentOS 7.9" into os_name and the target value "online" (mapped to the internal status running) into status.

[0133] In an exemplary embodiment, the method for invoking the language tool further includes: in the event that the target value is missing from the target parameter, initiating a clarification mechanism, wherein the clarification mechanism generates a single question and answer for the user to obtain the target value; determining the target value according to the clarification mechanism, and filling the target value into the target parameter required by the target language tool.

[0134] In this embodiment, if there are still missing required parameters after automatic parameter filling, i.e., the target parameter is missing its target value, the parameter completion module does not directly return failure, but intelligently generates a single clarification question for the user.

[0135] For example, if a required parameter has no corresponding information in the original language query, the system will initiate an interactive clarification mechanism. For instance, if the user's original language is "create a cloud server", and the "flavor_id (specification ID)" parameter of the compute.create_instance tool is required, the system will generate a single clarification question: "What specification of cloud server do you need to create? For example: 2 cores and 4GB of memory.", and provide up to 5 common specifications as options for the user to confirm.

[0136] In an exemplary embodiment, the method for invoking the language tool is as follows: based on the user's input of the original language, the target value is determined through a domain dictionary mapping table and a regular expression library; when there are multiple enumerated values ​​corresponding to the target parameter, the target value is matched according to the synonym mapping or near-synonym mapping, and the target value is filled into the target parameter required by the target language tool.

[0137] In this embodiment, when the target parameter has a preset enumeration value (such as the "status" parameter being active or down), the system performs a mapping match of synonyms or approximate values ​​within the enumeration set to fill the target value into the target parameter required by the target API. If the confidence of the approximate match is lower than a preset threshold, it will not be filled, triggering the subsequent clarification process to ensure the accuracy of the parameter.

[0138] It should be noted that during the parameter completion process, the system prioritizes the use of rules, dictionaries, and lightweight models with lower computational costs. When the above low-cost methods fail to produce high-confidence results, the system selectively calls large language models for deeper semantic understanding or reasoning, thereby not only ensuring the accuracy of the results but also greatly reducing resource consumption costs.

[0139] Through the above embodiments, the interactive clarification mechanism greatly improves the user experience. When information is insufficient, it does not simply return failure, but intelligently generates a single, focused clarification question, which can be accompanied by options, thereby enhancing the user interaction experience.

[0140] In an exemplary embodiment, the identifier of the target language tool and the target parameters are encapsulated to generate an execution request, including: verifying the target parameters and authenticating the user identity token; if the target parameters are successfully verified and the user identity token is successfully authenticated, the identifier of the target language tool and the target parameters are encapsulated to generate an execution request.

[0141] In this embodiment, as Figure 4The tool execution and multi-tool execution methods shown validate the target parameter type, format, and value range before execution. Based on the user initiating the call, the user's identity token is parsed to confirm the user's permission to call the API. The validation logic is based on Swagger Schema, supporting not only query parameters but also fully covering the validation of complex parameter structures and types such as "application / json" bodies, "multipart / form-data," arrays, and nested objects. If the target parameters are successfully validated and the user's identity token is authenticated, the target API ID and target parameter package are encapsulated. After validation and permission checks, an HTTP request is generated and sent to the cloud platform API gateway.

[0142] Through the above embodiments, real-time authentication of user identity tokens, combined with format and compliance verification of target parameters, effectively prevents unauthorized access and malicious calls, protects the resource security of the cloud platform, and maintains the overall stability of the system.

[0143] In an exemplary embodiment, after invoking the target language tool based on the execution request, the method further includes: generating an execution result for the original language based on the target language tool and in conjunction with a large language model; standardizing the execution result to obtain a result list; and returning the result list to the user.

[0144] In this embodiment, based on the target API being called and combined with LLM, execution results for the original language are generated. After standardizing the execution results, they are displayed to the user in the form of a list.

[0145] In an exemplary embodiment, the execution results are standardized to obtain a result list, including: pagination based on the execution results and a control mechanism to obtain multiple paginated results, wherein the control mechanism is a pre-set number of pages, number of items per page, and execution stop conditions; and the multiple paginated results are aggregated to obtain a result list.

[0146] In this embodiment, the tool execution module can automatically process the pagination logic based on the execution results and has a built-in explicit control mechanism, such as setting page_limit (maximum number of pages), max_items (maximum number of items) and early_stop (early stop condition), so that multiple pagination results are obtained after pagination processing.

[0147] Aggregate multiple paginated results into a list and wrap it in a unified response structure before returning it to the user.

[0148] Through the above embodiments, pagination can control the amount of data requested in a single request, avoiding performance bottlenecks and resource waste caused by loading too much data at once, and improving overall operating efficiency and stability.

[0149] In one exemplary embodiment, multiple tool cards generated for multiple candidate language tools are cached for a long time, while the target language tool invoked based on the execution request is cached for a short time.

[0150] In this embodiment, as Figure 4 The caching and performance optimization shown here is an intelligent caching mechanism designed to improve system response speed and further reduce costs.

[0151] Tool card caching: For frequently queried queries, the generated tool cards are cached long-term. When the same query occurs again, the system can directly retrieve the tool card from the cache, skipping the first few steps of routing.

[0152] Query result caching: API call results are cached for a short period of time, and different TTLs are set according to the volatility of the data. If a user repeats the same query within a short period of time, the cache will be hit directly, without having to call the underlying API again. The cache TTL can be adaptively adjusted according to the frequency of changes in API call results.

[0153] Through the above embodiments, the refined design of the multi-level caching mechanism ensures data isolation in a multi-user environment while improving performance.

[0154] In an exemplary embodiment, after constructing the index data structure based on the specification document, the method further includes: detecting whether the specification document has changed; if the specification document has changed, triggering the index construction mechanism to obtain a new index data structure; and switching the user's original language to the new index data structure for querying, so that the index data structure is updated without interruption.

[0155] In this embodiment, the index data structure (inverted index table, domain dictionary mapping table, etc.) is not static, but changes according to the API, as detailed below:

[0156] Automated monitoring and rebuilding: The system includes a background daemon that periodically polls the publication address of the Swagger / OpenAPI specification document for the cloud platform API. Once a change in the specification document content is detected (such as adding APIs, modifying parameters, deleting interfaces, etc.), the index building mechanism is triggered, and the index building module will be automatically triggered to re-parse the Swagger / OpenAPI specification document.

[0157] Uninterrupted Updates: The index rebuilding process is completed in the background. After the new index data structure is built, the system can use a "blue-green deployment" or similar strategy to smoothly switch natural language queries to the new index, while the old index is safely destroyed once there is no more traffic. The switching process has no impact on online services, ensuring business continuity. At the same time, to ensure compatibility with hot updates, the system supports cross-version compatibility strategies, allowing older versions of tool cards to still be parsed and mapped to new API operations within a preset window period.

[0158] Through the above implementation methods, this application embodiment, by constructing an index data structure for specification documents, can quickly respond to the user's input of raw language, filter multiple candidate language tools from a large-scale language tool pool, and reduce the cost of processing full information by a large language model; the tool card generation mechanism further narrows down the key information of candidate language tools. Finally, based on the target language tool and target parameters corresponding to the target tool card, an execution request is generated, and the target language tool is invoked through the execution request. This method solves the technical problems of low API call efficiency and high operation and maintenance costs caused by integrating large-scale APIs into large language models in related technologies, achieving accurate mapping from natural language queries to API calls, and achieving the technical effect of improving the efficiency and accuracy of language tool invocation.

[0159] To adapt to the continuous iteration of cloud platforms and the ongoing evolution of APIs, this application also provides a lifecycle management and hot update mechanism to ensure that the system can operate stably and efficiently for a long period of time. Figure 5 This is a schematic diagram illustrating lifecycle management and hot update based on language tool invocation according to an embodiment of this application, as shown below. Figure 5 As shown, this lifecycle management and hot update includes the following processes:

[0160] (1) Hot updates of indexes and domain dictionaries

[0161] The index data structure (inverted index table, domain dictionary mapping table, etc.) is not static, but changes according to the API, as detailed below:

[0162] Automated monitoring and rebuilding: The system includes a background daemon that periodically polls the publication address of the Swagger / OpenAPI specification document for the cloud platform API. Once a change in the document content is detected (such as adding an API, modifying parameters, deleting an interface, etc.), the index building module will be automatically triggered to re-parse the Swagger / OpenAPI specification document.

[0163] Uninterrupted Updates: The index rebuilding process is completed in the background. After the new index data structure is built, the system can use a "blue-green deployment" or similar strategy to smoothly switch natural language queries to the new index, while the old index is safely destroyed once there is no more traffic. The switching process has no impact on online services, ensuring business continuity. At the same time, to ensure compatibility with hot updates, the system supports cross-version compatibility strategies, allowing older versions of tool cards to still be parsed and mapped to new API operations within a preset window period.

[0164] Log-based dictionary expansion: The monitoring and learning module analyzes key user logs, especially those that cause routing failures or require multiple clarifications. By analyzing these logs, the system can discover new business terms and synonyms, and add them to the domain dictionary mapping table, thereby improving the accuracy of tool routing.

[0165] (2) Evolution of higher-order aggregation tools

[0166] Higher-order aggregation tools are used to handle complex tasks, and the embodiments of this application provide flexible definition and update methods.

[0167] Configuration and Version Management: New high-level aggregation tools can be defined using simple configuration files (such as YAML or JSON). These configuration files declare the tool ID, functionality description, parameter list, and the atomic API call chains and data aggregation logic it contains. To ensure the stability of changes, this configuration or its corresponding DSL (Domain Specific Language) is versioned.

[0168] Plug-in loading and hot reloading: The system automatically scans and loads the configuration files of advanced aggregation tools at startup. Meanwhile, operations and maintenance personnel can add, modify, or delete configuration files while the system is running. The system monitors changes to the configuration files and dynamically hot-reloads or unloads the corresponding aggregation tools without restarting the entire service.

[0169] Canary release and one-click rollback: For critical changes, canary release and one-click rollback strategies are supported to prevent runtime changes from causing widespread service inconsistencies.

[0170] (3) Permissions and User Management

[0171] In a multi-user cloud environment, this application provides a data isolation and access control system.

[0172] Access control during the routing phase: When a user request is received, the system first parses the user's authentication information (such as a token) to determine the user's permissions. The access control model is sourced from the securitySchemes definition in the Swagger documentation, the backend's "capability interfaces," or a centralized policy center. In tool routing, the system uses this access control information to pre-select candidate APIs. Even if a user's query semantically matches a certain API, if the user does not have the execution permissions for that API, that API will not appear in the available tool cards returned to the caller.

[0173] Credential delivery during execution: When the tool execution module initiates an HTTP request, it securely transmits the user's original authentication credentials or an authorized proxy credential to the downstream API gateway, ensuring that the underlying API calls comply with the cloud platform's security policies and auditing standards.

[0174] Full-chain audit logs: Tool calls, including internal atomic API calls, can be recorded in detail and associated with the caller's identity, user ID, and trace ID, providing a complete traceability chain for subsequent security audits and troubleshooting.

[0175] (4) Dynamic optimization of logs

[0176] The system forms a closed loop of continuous optimization through comprehensive log recording and data analysis.

[0177] End-to-end metrics monitoring: The system records detailed metrics throughout the entire process from receiving a user's natural language query to returning the execution result, including but not limited to the time consumed at each stage, routing accuracy, parameter completion success rate, cache hit rate, API call success / failure rate, and latency.

[0178] Log Compliance: To comply with data protection regulations, the system anonymizes personally identifiable information when recording logs, such as hashing or partially masking sensitive fields like resource names and usernames. It also sets a clear and configurable log retention period (e.g., 180 days by default) and provides a strict access auditing mechanism to record all log viewing operations.

[0179] Dynamic strategy adjustment: Through long-term analysis of the above indicators, the system can automatically adjust its internal strategies. For example, if an API call has extremely high latency, its relevance score in the reordering algorithm can be dynamically reduced; if the cache hit rate of a certain query pattern remains low, its caching strategy or TTL can be adjusted.

[0180] In order to ensure data consistency and handle various abnormal situations, this application provides an exception handling and rollback strategy. Figure 6This is a schematic diagram of an exception handling and rollback strategy based on language tool invocation according to an embodiment of this application, as shown below. Figure 6 As shown, the exception handling and rollback strategy includes the following process:

[0181] (1) Low confidence level processing

[0182] In the tool routing module, if the system cannot identify any high-confidence candidate APIs for a user's natural language query, the following approach is taken:

[0183] Threshold judgment: If the highest score calculated by the rearrangement algorithm is lower than the confidence threshold, the routing result is determined to be unreliable.

[0184] Interactive clarification: Generate a guiding clarification question. For example, if a user queries "Please check the status of a certain resource," and the information is too vague, generate a guiding clarification question: "What type of resource do you want to query? For example: cloud server, cloud disk, network, etc." This interactive communication strategy avoids resource waste caused by incorrect tool execution.

[0185] (2) Parameter error and conflict handling

[0186] Before parameter completion and tool execution, parameter validation is performed.

[0187] Missing parameters and incorrect formatting: For missing required parameters, interactive clarification will complete the field. For user-provided parameters with incorrect formatting (such as using an invalid string as an IP address), the system will explicitly point out the error and suggest the correct format.

[0188] Parameter conflict: In cases where APIs have mutually exclusive parameters (e.g., instance_id and instance_name cannot be specified simultaneously in a query), if a user's query triggers the filling of both mutually exclusive parameters, the system will select an API based on a preset priority rule (e.g., ID takes precedence over name), or directly initiate an interactive clarification question to the user.

[0189] (3) Server-side error mapping

[0190] To achieve end-to-end observability, the execution request passes the "trace_id" (e.g., via the "x-trace-id" header) from the client to the downstream service and connects with the trace chain of the cloud platform gateway and microservices. Execution of the underlying API may fail, returning HTTP error codes. This application embodiment performs a unified translation of HTTP error codes and forms an MCP-level error enumeration.

[0191] Error code translation: The system returns a structured error containing a "trace_id" to support link tracing. For example, a "403 Forbidden" response is mapped to {"error_code": "PERMISSION_DENIED","message":"Sorry, you do not have permission to perform this operation...", "trace_id":"..."}. A "404 Not Found" response is mapped to {"error_code": "NOT_FOUND",...}. Other error enumerations include "INVALID_PARAM", "RATE_LIMIT", "BACKEND_5XX", and "UNKNOWN".

[0192] Provide operational suggestions: When returning translated error messages, the system provides constructive suggestions for the next steps to help users resolve the problem.

[0193] (4) Rollback of intermediate state

[0194] If the task is interrupted during any of the above stages, the following procedure will be followed:

[0195] Transactional execution and state recording: For multi-step change operations (such as creating, modifying, and deleting resources), the system can wrap them in a transactional context. After each step is successfully executed, the intermediate results and state changes are recorded.

[0196] Compensation and Rollback: If an error occurs in a subsequent step, the system can roll back the completed steps according to predefined compensation logic. For example, in a task of "creating a cloud host and binding it to a public IP address," if the second step, "binding the public IP address," fails, the system automatically performs the reverse operation of the first step, namely, "deleting the created cloud host," thereby avoiding the creation of incomplete "zombie resources." After the rollback is performed, the system reports the failed steps and reasons to the agent, which then decides whether to perform partial corrections and retry, or to abandon the entire task.

[0197] The language tool invocation method and system provided in this application embodiment have high feasibility and incremental deployment capability, and can smoothly evolve from a minimal core system to a feature-rich, continuously automatically optimized intelligent system. Specifically, the system deployment evolution process is as follows:

[0198] (1) Initial deployment

[0199] The initial deployment of the system aims for "zero external dependencies" and "rapid results." Only the Swagger / OpenAPI specification document for the target cloud platform needs to be provided to start the core system, which specifically includes:

[0200] Core functionalities include index building, rule-based and inverted index-based tool routing, regular expression and dictionary-based parameter completion, single tool execution, and a basic FIFO (First In First Out) caching module.

[0201] Deployment constraints: The initial deployment does not rely on external vector databases or complex machine learning models, and meets the requirements for rapid deployment in resource-constrained or technology-stack-controlled private cloud environments.

[0202] (2) Gradually increase

[0203] After the initial version has been running stably and accumulated sufficient data, the system can be enhanced in stages through the following paths.

[0204] First enhancement: Introducing lightweight learning. During the tool routing reordering phase, a lightweight machine learning model (such as logistic regression or gradient boosting tree) is introduced. Using "query-correct API" logs collected by the monitoring module as training data, candidate APIs are ranked more accurately. The features of the machine learning model can include keyword matching scores, API call frequency, user historical preferences, etc.

[0205] The second enhancement: semi-automatic aggregation tool generation. The analysis engine mines frequently co-occurring API call sequences in logs. For example, if it discovers a large number of users querying cloud servers and then immediately querying their associated hard drive information, the analysis engine will recommend to the administrator to create a higher-level aggregation tool called "get_instance_with_volumes" and automatically generate its basic configuration framework.

[0206] The third enhancement: Implementing an intelligent caching strategy. The caching module is upgraded from a simple FIFO strategy to intelligent TTL management based on data change rate. By analyzing the frequency of changes in the results returned by specific APIs, the cache validity period is dynamically set, thereby achieving a balance between data freshness and cache hit rate.

[0207] (3) Long-term evolution

[0208] To be applicable to more complex application scenarios, the long-term evolution path of the system includes:

[0209] Introducing a complete learning sorter: Upgrading the reordering algorithm to a fully functional LTR model, capable of handling richer features and implementing more complex sorting logic, further improving routing accuracy.

[0210] Context-aware retrieval integration: Deeply integrates user context information during the tool routing stage, considering not only the query itself, but also the user's role permissions, historical operation habits, current workflow node, etc., to achieve highly personalized tool recommendations.

[0211] Exploring Hybrid Retrieval Techniques: For query scenarios where unstructured text (such as user manuals and technical documents) is the primary information source, vector retrieval and RAG techniques are explored as supplementary methods. In this hybrid mode, attempts are made to solve problems using keywords and structured indexes. When the matching degree is low, vector retrieval is then invoked to understand more ambiguous or broader query intents.

[0212] To quantitatively evaluate the effectiveness and performance of the system in this application, the embodiments of this application provide the following series of metrics and acceptance indicators, specifically including:

[0213] (1) Accuracy index

[0214] Top-K Candidate Coverage: Measures the performance of the tool routing module. The statistical caliber of this metric is based on a manually labeled set, that is, what proportion of user queries result in the correct API appearing in the Top-K (e.g., K=3) candidate tool cards returned by the tool routing module.

[0215] Parameter completion success rate: After selecting an API, the system can automatically and successfully fill in the required parameters from the user's query. This parameter completion success rate directly reflects the depth of the system's understanding of the user's intent.

[0216] End-to-end success rate: The percentage of user-initiated queries that are successfully executed and return the expected results. When obtaining the end-to-end success rate, cases of "expected failures" due to insufficient permissions should be excluded to more accurately reflect the system's core processing capabilities.

[0217] (2) Performance indicators

[0218] Token consumption: The average number of context tokens that need to be submitted to the upper-level large language model for decision-making for each user query.

[0219] Average response time: The end-to-end latency from when the system receives a user query to when it returns a result (or the first clarification of the question). It is usually measured using the P95 (95th percentile) or P99 value to evaluate the system's performance under high load.

[0220] Cache hit rate: The ratio of cache hits between query results cache and tool card cache. A high cache hit rate indicates a lower frequency of underlying API calls, faster response times, and lower system load.

[0221] (3) User interaction metrics

[0222] Clarification question occurrence rate: The percentage of queries generated by the system to guide the user's search. The system should be as "understandable on the first try" as possible, therefore a lower clarification question occurrence rate is better.

[0223] Secondary clarification rate: The percentage of users who provide insufficient information after the system raises the first clarification question, requiring the system to conduct a second or more clarifications.

[0224] (4) Maintainability indicators

[0225] Automatic index rebuild time: The total time required for the system to complete index rebuild and switch to the new index after detecting an update to the Swagger / OpenAPI specification document.

[0226] Uninterrupted service verification: During index rebuilding and hot updates of advanced tools, verify whether the system can continuously provide services without any interruption.

[0227] Log diagnostics: Assess whether system logs provide sufficiently clear and comprehensive information to support operations personnel in quickly diagnosing problems, analyzing performance bottlenecks, and making optimization decisions.

[0228] To better understand the above method, the following examples illustrate the process, but are not intended to limit the technical solutions of this application. Specifically:

[0229] For example, a user of a cloud platform might use natural language to query: "Please find all running cloud hosts with CentOS 7.9 version belonging to 'Project A', and their CPU core count must be greater than or equal to 4 cores."

[0230] This application embodiment converts the natural language query into a precise API call and returns the execution result, including the following steps:

[0231] Step 1: Construct the index data structure;

[0232] When the system is deployed for the first time or when an update to the API specification document is detected, the index building module is automatically executed to pull and parse the Swagger / OpenAPI document from the cloud platform, as follows:

[0233] Generate an atomic interface registry: parse the API path (e.g., " / v2.1 / servers / detail"), method ("GET"), operationId ("listServers"), summary ("List servers"), description, parameters, and detailed attributes (type, required, description, etc.), and store them in a structured database. For example, the listServers interface is recorded in detail.

[0234] Build an inverted index table: Extract keywords and establish mappings. For example, the keywords "cloud host" and "server" are mapped to compute.servers.listServers; the keywords "status" and "status" are mapped to API.

[0235] Populate the domain dictionary mapping table: map terms such as "cloud host" and "virtual machine" to "computing domain".

[0236] Load the regular expression library: Ensure that your system contains regular expressions for identifying operating system versions (such as "CentOS 7.9"), IP addresses, UUIDs, etc.

[0237] Step 2: Query processing flow and multi-stage tool routing;

[0238] After the user inputs a natural language query, the tool's routing module uses a multi-stage, step-by-step filtering "funnel-style" routing strategy.

[0239] Phase 1: Service Domain and Tag Filtering. The system preprocesses the query text, identifies the keyword "cloud host," and limits the search scope to the "computing domain" by querying the domain dictionary mapping table. Simultaneously, it identifies the entity "CentOS 7.9" using a regular expression library. This first phase narrows down the candidate APIs from hundreds to dozens relevant to the computing domain.

[0240] Phase Two: Inverted Index Recall and Lightweight Reordering. Within the compute domain's API set, the system uses segmented keywords ("running", "project A", "CentOS 7.9", "cloud host", "CPU", "4 cores") to recall results in the inverted index. The compute.servers.listServers interface was recalled due to matching multiple keywords such as "cloud host".

[0241] A reordering algorithm is used to score the retrieved APIs. For the listServers interface:

[0242] If the name or summary matches "cloud host", it will receive a higher β1 or β2 weight.

[0243] The parameter list includes status, project_id, imageRef, etc., which are related to concepts such as "running", "project A", and "CentOS" in the query, and obtain a β3 weight.

[0244] The query "CentOS 7.9" was recognized by the regular expression library, and the parameters of listServers can accept image-related information, thus earning a β5 regular expression reward.

[0245] Ultimately, the listServers interface achieved a high relevance score, exceeding the confidence threshold τ.

[0246] Phase 3: Generating Candidate Tool Cards. The system selects the Top-K (e.g., K=3) APIs with the highest relevance scores and generates corresponding tool cards. For example, the tool card for the listServers interface is shown below and placed first:

[0247] json

[0248] {

[0249] "tool_id":"compute.servers.listServers",

[0250] "title":"Query Cloud Server List",

[0251] "one_line":"Query and list detailed cloud server information based on multiple filtering criteria",

[0252] "required_params": [],

[0253] "score":0.95,

[0254] "why":["Matched the keyword 'cloud server'", "The query intent is highly relevant to the filtered query"],

[0255] "notes":{"regex": {"os_version": "CentOS 7.9"}},

[0256] "low_confidence": false,

[0257] "clarify_suggestion": null

[0258] }

[0259] ```

[0260] The aforementioned tool cards, along with the remaining tool cards, are submitted to the agent or LLM for final decision-making. Due to the highly condensed information, the token consumption for this submission is extremely low. The compute.servers.listServers tool can be selected based on the highest "score" and "why" fields.

[0261] Step 3: Parameter completion and clarification interaction;

[0262] After receiving the selection results from the compute.servers.listServers tool, the parameter completion module automatically fills in the parameters required by the tool from the original user language query.

[0263] Automatic parameter completion: The parameter completion module re-analyzes the original language query "Please find all running cloud hosts of CentOS 7.9 version belonging to 'Project A', and their CPU core count must be greater than or equal to 4 cores".

[0264] "Running" is mapped through the domain dictionary mapping table and populated into the "status" parameter, with a value of active.

[0265] "Project A" is populated into the "project_id" parameter by querying the project list API or by direct matching.

[0266] "CentOS 7.9" is first mapped to a specific image ID. The internal images.list interface is called, and the imageRef is determined by filtering with name=CentOS 7.9, and the value is populated.

[0267] The query "CPU core count must be greater than or equal to 4 cores" is a numerical range query. The parameter completion module parses the condition as flavor.vcpus≥4. It calls the flavors.list interface to obtain a list of all flavors with vcpus≥4, and passes the set of IDs of the above flavors as the filter condition to the listServers interface.

[0268] Clarifying Interaction: In this embodiment, if the required parameters ("project_id" is mandatory) and most optional parameters have been successfully filled, and the user enters "Create a cloud host with 4 CPU cores," then "image_id" and "network_id" are mandatory parameters. The parameter completion module generates a clarifying question: "Okay, creating a cloud host with 4 CPU cores. Which operating system image do you want to use, and which network do you want to connect it to?" This clarifying question maximizes information gain and can quickly fill in missing key information.

[0269] Step 4: Tool execution and return of execution results;

[0270] After the parameters are populated, the tool execution module receives an execution request containing the tool ID and the populated parameters, and executes the following process:

[0271] Parameter validation and permission check: The tool execution module validates the type and format of parameters based on the listServers Swagger Schema. It also parses the user's identity token to confirm that the user has permission to execute listServers operations in "Project A".

[0272] REST Call and Pagination Processing: The tool execution module constructs an HTTP request (execution request), which is sent to the cloud platform API gateway. If the returned execution result is paginated, the tool execution module automatically processes the "next" link, retrieving data for all pages according to the preset page_limit and max_items, until all data is retrieved or the upper limit is reached.

[0273] Execution result formatting: The tool execution module aggregates multiple paginated results into a list and wraps it in a unified response structure before returning it to the user or agent.

[0274] json

[0275] {

[0276] "ok":true,

[0277] "data":{

[0278] "items":[

[0279] {"id":"uuid-1", "name": "vm-01", ...} ,

[0280] {"id":"uuid-2", "name": "vm-02", ...}

[0281] ],

[0282] "stats":{"count": 2}

[0283] },

[0284] "meta":{

[0285] "source":["compute.servers.listServers"],

[0286] "page":1, "pageSize": 2, "total": 2,

[0287] "elapsed_ms":150,

[0288] "trace_id":"trace-xyz-123"

[0289] }

[0290] }

[0291] ```

[0292] Step 5: Monitoring, caching, and continuous learning;

[0293] Key information in the above process, including but not limited to user queries, route scores, final selections, parameter filling status, execution time, and results, is recorded by the monitoring and learning module.

[0294] Caching Application: If other users with the same permissions initiate the exact same natural language query within a short period of time, the caching and strategy module will directly return the above results from the cache, greatly reducing the response time.

[0295] Continuous learning: If subsequent analysis reveals that a large number of users are using "virtual machine" instead of "cloud server" for queries, resulting in low routing scores, the system will automatically add "virtual machine" to the synonym list for "cloud server," thereby improving its ability to understand this type of query in the future.

[0296] Through the above implementation methods, this application embodiment accurately transforms a complex natural language query into a precise call to the underlying cloud platform API in an automated, phased, and low-cost manner, and achieves continuous self-optimization.

[0297] Compared with related technologies, the embodiments of this application bring at least the following beneficial effects in integrating large-scale APIs into LLM applications.

[0298] (1) Reduced token consumption and operational costs: Related technologies typically input all API documentation into the LLM context, resulting in a huge amount of token consumption per query. This application's embodiment uses a "funnel-style" multi-stage routing strategy to filter out most irrelevant APIs in the early stages using a lightweight model and deterministic rules, providing the LLM with a "tool card" containing K (K is usually 3 to 5) highly relevant candidate APIs. This reduces the size of the context submitted to the LLM from tens of thousands or even hundreds of thousands of tokens to a few hundred tokens, achieving an order-of-magnitude reduction. This reduction in token consumption directly translates into a significant saving in API call costs.

[0299] (2) Improved accuracy of tool selection and task success rate: This embodiment of the application achieves more accurate tool routing than general semantic retrieval by deeply analyzing Swagger / OpenAPI structured information and combining inverted indexes, domain dictionaries, and regular expression libraries. The reordering algorithm comprehensively considers matching information from multiple dimensions such as name, summary, and parameters, and rewards specific entity formats, effectively distinguishing APIs that are functionally similar but have different details. This accurate routing ensures the quality of candidate tools provided to LLM, thereby greatly reducing the probability of the model generating illusions and improving the accuracy of tool selection. Combined with the clarifying interaction mechanism for intelligent completion of missing parameters, the success rate of the entire end-to-end task from natural language to API execution is significantly guaranteed.

[0300] (3) It achieves automated adaptation to the API lifecycle, greatly reducing maintenance costs: The index building module in this embodiment can automatically monitor changes to the Swagger / OpenAPI documentation and trigger uninterrupted hot updates of the index. Whether it is adding an API, modifying parameters, or deprecated interfaces, the system can automatically detect and reflect it in the tool routing and execution logic without manual intervention. It achieves the capability of "one-time access, continuous synchronization", getting rid of the huge burden of manually writing and maintaining adapters, making the system highly scalable and with extremely low maintainability costs.

[0301] (4) Enhanced system security, stability, and observability: This application embodiment performs pre-access permission pruning during the routing phase to ensure users can see and select the tools they are authorized to access. User credentials are transmitted during the execution phase, adhering to the platform's security policy. A multi-level caching mechanism and refined design of cache keys improve performance while ensuring data isolation in a multi-user environment. For API calls, the system incorporates stability safeguards such as pagination, exponential backoff retries, timeout control, and concurrency budgeting. Comprehensive logging and end-to-end trace ID transmission provide complete observability for troubleshooting and performance optimization, ensuring long-term stable and reliable operation of the system in complex production environments.

[0302] (5) Improved user interaction experience: The interactive clarification mechanism proposed in this application greatly enhances the user experience. When information is insufficient, the system does not simply return failure, but intelligently generates a single, focused clarification question with options to proactively help the user provide the necessary information. This approach avoids inefficient and tedious multi-round "toothpaste-squeezing" dialogues with users, making the interaction process smoother, more natural, and more efficient. Even if the user's initial question is vague, the system can help them complete the task with a high probability.

[0303] In summary, the embodiments of this application not only solve the technical problem of LLM accessing large-scale APIs, but also provide a reliable and easy-to-maintain solution.

[0304] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0305] According to another aspect of the embodiments of this application, a language tool invocation device is also provided, the structural schematic diagram of which is shown below. Figure 7 As shown, it includes the following modules:

[0306] Module 702 is used to construct the index data structure based on the specification document;

[0307] The query module 704 is used to query the index data structure based on the user's input in the original language to determine multiple candidate language tools;

[0308] The generation module 706 is used to generate multiple tool cards for multiple candidate language tools, and to determine the target language tool corresponding to the target tool card from the multiple tool cards, wherein one tool card is generated for each candidate language tool;

[0309] Module 708 is used to generate an execution request based on the target language tool and target parameters, and to invoke the target language tool based on the execution request.

[0310] The specific execution steps involved in the various calculation processes and dynamic optimization of storage space in the above modules can be referred to the description in the above embodiments, and will not be repeated here.

[0311] Obviously, the language tool invocation device described above can be used to implement the language tool invocation method provided in the above embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0312] It should be noted that the construction module 702 in this embodiment can be used to execute the above step S302, the query module 704 in this embodiment can be used to execute the above step S304, the generation module 706 in this embodiment can be used to execute the above step S306, and the calling module 708 in this embodiment can be used to execute the above step S308.

[0313] In one exemplary embodiment, the above-described apparatus further includes an index data structure comprising an atomic interface data table, an inverted index table, a domain dictionary mapping table, and a regular expression library.

[0314] In an exemplary embodiment, the above-mentioned construction module 702 includes: a parsing submodule, used to parse the specification document of the cloud platform to obtain the specification structure and specification format of the corresponding language tool; and a construction submodule, used to construct an index data structure based on the specification structure and specification format.

[0315] In an exemplary embodiment, the query module 704 includes: a first query submodule, configured to perform a first query on the domain dictionary mapping table and the regular expression library based on the original language input by the user to obtain multiple initial language tools; and a second query submodule, configured to perform a second query on the inverted index table based on the multiple initial language tools to determine multiple candidate language tools.

[0316] In an exemplary embodiment, the first query submodule includes: a first processing unit, configured to preprocess the original language to obtain multiple keywords, wherein the preprocessing includes at least one of the following: letter conversion, mixed word segmentation, named splitting, and word count sliding window; a first determining unit, configured to determine the service domain of the original language query based on the multiple keywords and the domain dictionary mapping table; and a first query unit, configured to perform a format signal query based on the service domain and the regular expression library to obtain multiple initial language tools.

[0317] In an exemplary embodiment, the second query submodule includes: a second query unit, configured to perform keyword queries based on the plurality of initial language tools through the inverted index table to generate a language tool list, wherein the language tool list includes a plurality of language tools; and a second determination unit, configured to score and sort the plurality of language tools to determine a plurality of candidate language tools.

[0318] In an exemplary embodiment, the second determining unit includes: a scoring subunit, configured to perform weighted scoring on the plurality of language tools according to a rearrangement formula to obtain a relevance score corresponding to each language tool; and a determining subunit, configured to sort the language tools corresponding to the relevance scores when the relevance scores meet a first threshold, to determine a plurality of candidate language tools.

[0319] In one exemplary embodiment, the apparatus further includes: the rearrangement formula includes weighted scoring of at least one of the following: word frequency, original language length, keyword hit, summary hit, parameter hit, word reward, and regularity reward.

[0320] In an exemplary embodiment, the generation module 706 includes: a filling submodule, configured to fill in the target parameters required by the target language tool based on the original language input by the user; and a first generation submodule, configured to encapsulate the identifier of the target language tool and the target parameters to generate an execution request.

[0321] In an exemplary embodiment, the above-mentioned filling submodule includes: a first filling unit, configured to determine a target value based on the original language input by the user, through the domain dictionary mapping table and the regular expression library, and fill the target value into the target parameters required by the target language tool.

[0322] In one exemplary embodiment, the above-mentioned filling submodule further includes: an acquisition unit, configured to initiate a clarification mechanism when the target parameter is missing a target value, wherein the clarification mechanism is to generate a single question and answer for the user to obtain the target value; and a second filling unit, configured to determine the target value according to the clarification mechanism and fill the target value into the target parameter required by the target language tool.

[0323] In an exemplary embodiment, the above-mentioned filling submodule further includes: a third determining unit, configured to determine a target value based on the original language input by the user through the domain dictionary mapping table and the regular expression library; and a third filling unit, configured to, when there are multiple enumerated values ​​corresponding to the target parameter, match the target value according to synonym mapping or near-synonym mapping, and fill the target value into the target parameter required by the target language tool.

[0324] In an exemplary embodiment, the first generation submodule includes: a verification unit, configured to verify the target parameters and authenticate the user identity token; and an encapsulation unit, configured to encapsulate the identifier of the target language tool and the target parameters to generate an execution request when the target parameters are successfully verified and the user identity token is successfully authenticated.

[0325] In an exemplary embodiment, the aforementioned invocation module 708 includes: a second generation submodule, configured to generate an execution result for the original language based on the target language tool and in conjunction with a large language model; and a processing submodule, configured to standardize the execution result to obtain a result list and return the result list to the user.

[0326] In an exemplary embodiment, the above-mentioned processing submodule includes: a second processing unit, configured to perform pagination processing based on the execution result and the control mechanism to obtain multiple pagination results, wherein the control mechanism is a pre-set number of pages, number of entries per page, and execution stop condition; and an aggregation unit, configured to aggregate the multiple pagination results to obtain a result list.

[0327] In one exemplary embodiment, the above-described apparatus further includes: long-term caching of multiple tool cards generated for the multiple candidate language tools, and short-term caching of the target language tool invoked based on the execution request.

[0328] In one exemplary embodiment, the above-described apparatus further includes: a detection submodule for detecting whether the specification document has been changed; a triggering submodule for triggering an index building mechanism to obtain a new index data structure when the specification document has been changed; and a third query submodule for switching the user-inputted original language to the new index data structure for querying, so that the index data structure is updated without interruption.

[0329] For a description of the features in the embodiment corresponding to the language tool invocation device, please refer to the relevant description of the embodiment corresponding to the language tool invocation method, which will not be repeated here.

[0330] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in the invocation method embodiments of any of the above-described language tools.

[0331] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in the above-described language tool invocation method embodiments at runtime.

[0332] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0333] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the invocation method embodiments of any of the above-described language tools.

[0334] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the invocation method embodiments of any of the above-described language tools.

[0335] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0336] The above provides a detailed description of a method for invoking a language tool provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for invoking a language tool, characterized in that, include: Construct the index data structure based on the specification document; The index data structure is queried based on the user's input of the original language to determine multiple candidate language tools. The index data structure includes an inverted index table, a domain dictionary mapping table, and a regular expression library. The language tools include application programming interfaces. Multiple tool cards are generated for the multiple candidate language tools, and the target language tool corresponding to the target tool card is determined from the multiple tool cards, wherein one tool card is generated for each candidate language tool; An execution request is generated based on the target language tool and target parameters, and the target language tool is invoked based on the execution request; The process of querying the index data structure based on the user's input in the original language to determine multiple candidate language tools includes: The original language is preprocessed to obtain multiple keywords, wherein the preprocessing includes at least one of the following: letter conversion, mixed word segmentation, named splitting, and word count sliding window; Based on the multiple keywords and the domain dictionary mapping table, the service domain of the original language query is determined; Based on the service domain and the regular expression library, a format signal query is performed to obtain multiple initial language tools; Based on the multiple initial language tools, a list of language tools is generated by performing keyword queries through the inverted index table, wherein the list of language tools includes multiple language tools; The multiple language tools are scored and ranked to identify multiple candidate language tools.

2. The method according to claim 1, characterized in that, in, The index data structure also includes atomic interface data tables.

3. The method according to claim 1, characterized in that, The construction of the index data structure based on the specification document includes: Parse the specification document of the cloud platform to obtain the specification structure and specification format of the corresponding language tools; Based on the aforementioned specification structure and format, an index data structure is constructed.

4. The method according to claim 1, characterized in that, The process of scoring and ranking the multiple language tools to determine multiple candidate language tools includes: The multiple language tools are weighted and scored according to the rearrangement formula to obtain the relevance score for each language tool; If the relevance score meets the first threshold, the language tools corresponding to the relevance score are sorted to determine multiple candidate language tools.

5. The method according to claim 4, characterized in that, in, The rearrangement formula includes weighted scoring of at least one of the following: word frequency, keyword hit, abstract hit, parameter hit, word bonus, and regularization bonus.

6. The method according to claim 1, characterized in that, The step of generating an execution request based on the target language tool and target parameters includes: The target parameters required by the target language tool are filled in based on the original language input by the user; The identifier of the target language tool and the target parameters are encapsulated to generate an execution request.

7. The method according to claim 6, characterized in that, The target parameters required by the tool to fill in the target language based on the original language input by the user include: Based on the original language input by the user, the target value is determined through a domain dictionary mapping table and a regular expression library, and the target value is filled into the target parameters required by the target language tool.

8. The method according to claim 7, characterized in that, Also includes: If the target parameter is missing a target value, a clarification mechanism is initiated, wherein the clarification mechanism generates a single question and answer for the user to obtain the target value; The target value is determined according to the clarification mechanism, and the target value is filled into the target parameters required by the target language tool.

9. The method according to claim 7, characterized in that, Also includes: Based on the original language input by the user, the target value is determined through the domain dictionary mapping table and the regular expression library; When there are multiple enumerated values ​​corresponding to the target parameter, the target value is matched according to the synonym mapping or near-synonym mapping, and the target value is filled into the target parameter required by the target language tool.

10. The method according to claim 6, characterized in that, The step of encapsulating the identifier of the target language tool and the target parameters to generate an execution request includes: The target parameters are verified, and the user identity token is authenticated. If the target parameter verification is successful and the user identity token authentication is passed, the identifier of the target language tool and the target parameter are encapsulated to generate an execution request.

11. The method according to claim 1, characterized in that, After invoking the target language tool based on the execution request, the method further includes: Based on the target language tool and combined with the large language model, the execution result for the original language is generated; The execution results are standardized to obtain a result list, and the result list is returned to the user.

12. The method according to claim 11, characterized in that, The result list obtained by standardizing the execution result includes: Based on the execution results and control mechanism, pagination is performed to obtain multiple pagination results. The control mechanism includes a pre-set number of pages, the number of items per page, and an execution stop condition. The multiple paginated results are aggregated to obtain a result list.

13. The method according to claim 1, characterized in that, in, The tool cards generated for the multiple candidate language tools will be cached for a long time, while the target language tool invoked based on the execution request will be cached for a short time.

14. The method according to claim 1, characterized in that, After constructing the index data structure based on the specification document, the method further includes: Check whether the specification document has been changed; When the specification document is changed, the index building mechanism is triggered to obtain a new index data structure; The query is performed by switching the user's original language to the new index data structure, so that the index data structure is updated without interruption.

15. A device for invoking a language tool, characterized in that, include: The building module is used to construct the index data structure based on the specification document; The query module is used to query the index data structure based on the user-input raw language to determine multiple candidate language tools. The index data structure includes an inverted index table, a domain dictionary mapping table, and a regular expression library. The language tools include application programming interfaces. The generation module is used to generate multiple tool cards for the multiple candidate language tools, and to determine the target language tool corresponding to the target tool card from the multiple tool cards, wherein one tool card is generated for each candidate language tool; The invocation module is used to generate an execution request based on the target language tool and target parameters, and to invoke the target language tool based on the execution request; The language tool invocation device is further configured to preprocess the original language to obtain multiple keywords, wherein the preprocessing includes at least one of the following: letter conversion, mixed word segmentation, naming splitting, and character count sliding window; determine the service domain of the original language query based on the multiple keywords and the domain dictionary mapping table; perform format signal query based on the service domain and the regular expression library to obtain multiple initial language tools; perform keyword query based on the multiple initial language tools through the inverted index table to generate a language tool list, wherein the language tool list includes multiple language tools; and score and sort the multiple language tools to determine multiple candidate language tools.

16. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as described in any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 14.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Man-machine interaction type query optimization method and system for incomplete query

    CN118643060A

  • Method and system for retrieving DOCX document content based on keywords

    CN121029978A