Centralized multi-agent platform for accurate and low-latency processing of user requests

CN122554539APending Publication Date: 2026-08-11GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,并且尤其是对于涉及许多组件的操作、特征等的复杂平台,这样的系统通常缺乏用于提供对广泛的问题的支持的灵活性/能力

Benefits of technology

[0008]而且,通过使用专用的或技能特定的AI智能体,而不是依赖于单个通用智能体或模型,所公开的技术可以在特定领域中提供更详细的专业知识和/或上下文特异性,例如,不需要单个大模型,该单个大模型将是处理更密集的、缓慢的和/或难以训练的、或者由于上下文窗口和/或其他限制而仅仅是不切实际的。此外,通过将查询路由到专用AI智能体,所公开的技术避免了对可能产生幻觉的单个更大模型的依赖,并且因此降低了产生不正确、不相关或虚构信息的风险。因为每个专用AI智能体在其相应领域中被特别训练/配置,所以通常由通用模型引起的不准确性和错误信息可以被减轻。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554539A_ABST
    Figure CN122554539A_ABST
Patent Text Reader

Abstract

A method for accurately and low-latency processing user requests includes receiving a user request and generating, by an orchestrator, a context-specific response to the user request. Generating the context-specific response includes generating, based on the user request and one or more context signals, a response development plan specifying one or more specialized AI agents. Generating the response development plan includes determining that the specified AI agents are relevant to the user request. Each specialized AI agent is configured to analyze queries specific to a respective domain. Generating the context-specific response also includes querying the specified AI agents according to the response development plan and, in response, receiving respective feedback from the AI agents. The context-specific response is generated based on the respective feedback.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 803,443, filed May 9, 2025, entitled “Centralized Multi-Agent Platform for Accurate and Low-Latency Handing of User Requests,” the entire disclosure of which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to query processing techniques, and more specifically, to techniques for providing accurate user request responses with low latency using a centralized architecture of artificial intelligence agents. Background Technology

[0004] The background description provided herein is for the purpose of presenting the general context of this disclosure. The work of the currently named inventors (to the extent described in this background section) and aspects of the specification that might not have been considered prior art at the time of filing are neither expressly nor impliedly acknowledged as prior art to this disclosure.

[0005] In many domains or use cases, users of computer / software-based platforms may have needs (problems, concerns, interests, etc.) that can vary greatly in focus, scope, and generality. For example, in the digital advertising space, users / entities with digital advertising accounts might want to understand how their ad performance has been trending, how to improve the performance of their campaigns (or specific aspects of a particular campaign), and why platform providers don't approve specific ads. However, users may need to take different and cumbersome steps to meet those needs. For instance, an advertising entity might query analytics tools and / or set various parameters to view past ad performance, consult with advisors / experts about techniques for optimizing campaign performance, and communicate with platform providers to inquire why specific ads were not approved, etc.

[0006] Automated query systems (e.g., user "help" tools) can provide a relatively fast, efficient, and low-cost way to resolve some user issues arising in certain applications, tools, or platforms. However, and especially for complex platforms involving the operation, features, etc., of many components, such systems often lack the flexibility / capability to provide support for a wide range of questions. While recent improvements to Large Language Models (LLMs) have made automated query systems more flexible, they are still generally limited in their scope of application by factors such as a tendency to produce illusions, limitations in context window size, and the inability to train reasonably sized models to adequately handle many different domains / aspects. While very large models and / or highly complex systems can solve some of these problems, this usually comes at the cost of greater processing resource usage and increased latency / delay in providing responses to user queries. Especially for chatbot platforms, even a moderate amount of "per-turn" latency is unacceptable. Summary of the Invention

[0007] The technology disclosed herein remedies the aforementioned deficiencies and other technical problems by providing a systematic, efficient, and automated query processing framework that uses an orchestrator to coordinate the actions of multiple specialized AI agents (e.g., reporting agents, diagnostic agents, optimization agents, billing agents, policy agents, etc.). Depending on the implementation, the orchestrator provides planning and response capabilities by using separate planner and response models or by using a single model. In response to receiving a user request (e.g., a query or instruction), the orchestrator generates a response development plan, at least in part, by determining which of the specialized AI agents are relevant to the user request, and generates a synthetic query (e.g., a prompt for the relevant AI agent) against the response development plan. As used herein, and unless the context clearly indicates otherwise, the terms “request” or “query” can refer, for example, to a single request or query transmission / message, or to a set of request or query transmissions / messages (e.g., in the course of a multi-round “dialogue”). Based on the feedback received in response to the synthetic query, the orchestrator generates a response to present to the user. Separating the responsibilities between the planner and the response phase or components in this way enables faster updates or optimizations and helps ensure that updates or optimizations have more predictable results.

[0008] Moreover, by using specialized or skill-specific AI agents, rather than relying on a single general-purpose agent or model, the disclosed techniques can provide more detailed expertise and / or context-specificity in a particular domain. For example, a single large model is unnecessary, as it would be more intensive, slow, and / or difficult to train, or simply impractical due to context windows and / or other limitations. Furthermore, by routing queries to specialized AI agents, the disclosed techniques avoid reliance on a single, larger model that may produce illusions, and thus reduce the risk of generating incorrect, irrelevant, or fabricated information. Because each specialized AI agent is specifically trained / configured in its respective domain, inaccuracies and misinformation typically caused by general-purpose models can be mitigated.

[0009] Therefore, the disclosed technology achieves a balance by providing the ability to accurately and flexibly handle a wide range of complex user requests without requiring computationally expensive, slow, and / or overly complex general-purpose models. Furthermore, the ability to add or remove specific dedicated AI agents as needed makes the disclosed architecture highly flexible and scalable, and the centralized architecture of the orchestrator and dedicated AI agents allows for more efficient conflict resolution, implementation of desired strategies, and so on. In addition, the disclosed technology can reduce latency and processing resource usage by resolving inconsistencies, conflicts, and lack of conversational fluency in a more efficient and centralized manner.

[0010] The disclosed techniques further help ensure that the final response is more relevant to the user who entered the request and / or another entity associated with the request (such as an advertiser who entered the request on their behalf) by leveraging contextual signals to generate response development plans and / or synthesize queries for dedicated AI agents. For example, the orchestrator may develop a response development plan in part based on information about the advertiser associated with the user's request, or information related to the advertiser's account or business. In some implementations, the disclosed techniques provide even greater contextual relevance by using one or more signals associated with the user interface through which the user's request was entered. For example, the orchestrator may develop a response development plan more accurately by analyzing information about the current state of the user interface (e.g., which page or section of a page the user is currently viewing, as captured in a screenshot).

[0011] In some implementations, further technical advantages are gained by adapting the operation of the dedicated AI agent, the responsive feedback provided by the dedicated AI agent to the orchestrator, and / or the way the orchestrator uses the responsive feedback from the dedicated AI agent. For example, when the dedicated AI agent receives a query from the orchestrator (according to the response development plan), the dedicated AI agent can respond by sending intermediate inference to the orchestrator that supports the response to the query (e.g., supports responses typically generated by the dedicated AI agent's response layer model) and the underlying data used by the dedicated AI agent to generate the intermediate inference, without having to generate and / or send the higher-level query response itself (or, in some implementations, the orchestrator receives but ignores the higher-level query response). In this way, the task of response formulation by the dedicated AI agent is effectively delegated to the orchestrator, which uses the intermediate inference and underlying data (and possibly other information from the dedicated AI agent, such as instructions on how to read the underlying data) to generate a response to the user request. By receiving the underlying inference and underlying data, the orchestrator obtains any necessary information sufficient (at least in some cases) to prevent re-querying the dedicated AI agent. By removing loops / rounds (or multiple loops / rounds) from orchestrator / agent communication, this implementation reduces the total latency of responses delivered to the user, while also reducing processing power usage. Furthermore, and contrary to intuition (given the greater domain expertise of the dedicated AI agent), removing the layer where the dedicated AI agent formulates its response to the orchestrator query reduces illusions, as conclusions or inferences drawn from the dedicated AI agent's response model can be a significant source of such illusions.

[0012] In some implementations, further technical advantages are gained by storing / maintaining the context of at least some of the dedicated AI agents. For example, for the purpose of generating a response development plan, the orchestrator may maintain a first agent context / history that indicates previous communication (e.g., queries / feedback) between the orchestrator and a first dedicated AI agent, and separately maintain a second agent context / history that indicates previous communication between the orchestrator and a second dedicated AI agent. Centralized maintenance of isolated agent-specific context / history effectively provides a much wider context window (or equivalently, eliminates the need for the dedicated AI agents themselves to maintain a larger context), while also reducing the likelihood of illusions by allowing each dedicated AI agent to focus more specifically on its own domain without cross-context contamination (e.g., no dedicated AI agent must analyze / consider unwanted data / history in synthetic queries).

[0013] Experiments have shown that a system where the orchestrator generates responses to user queries based on base data and intermediate inference (rather than higher-level responses) from dedicated AI agents, and where the orchestrator maintains separate / isolated contexts for different dedicated AI agents, can improve realism (no fact-based illusions) from approximately 87% to approximately 97%, and improve response latency to the user from approximately 8 seconds to approximately 3 seconds. Such improvements are possible because, for example, the system avoids illusions that are prone to occur at the response layer of dedicated AI agents, and typically requires fewer rounds / turns of communication between the orchestrator and the dedicated AI agents.

[0014] In some implementations and / or scenarios, the orchestrator responds to an initial user request at least in part by performing or triggering an action. For example, the orchestrator may identify one or more actions that should occur via the user interface, after which the orchestrator (or another component or device / system) may remotely assume control of the input functions of the client device to implement the identified actions (e.g., when the client device user observes a corresponding movement of the mouse pointer in interaction with the user interface).

[0015] Other advantages will become apparent to those skilled in the art after reading this disclosure and viewing the corresponding figures. Attached Figure Description

[0016] Figure 1 This is a block diagram of an example system that can implement the user request processing technology disclosed herein.

[0017] Figure 2 It describes what can be made by Figure 1 An example architecture for implementing a computing system.

[0018] Figure 3 It describes what can be made by Figure 2 within the architecture Figure 1 A collection of examples of dedicated artificial intelligence (AI) agents and tools used in computing systems.

[0019] Figure 4 This indicates that in the example scenarios and implementation methods, it is possible to... Figure 1 Computing systems and / or Figure 2 A message passing graph showing the sequence of messages exchanged within the architecture.

[0020] Figure 5 This indicates that alternative implementations are possible. Figure 1 Computing systems and / or Figure 2 Another sequence of messages exchanged within the architecture is represented by a message passing graph.

[0021] Figure 6An example process that can be implemented by one or more dedicated AI agents is described.

[0022] Figures 7A to 7D It describes various example scenarios and implementation methods. Figure 1 UI components can cause the user interface (UI) that they present to the user.

[0023] Figure 8 This is a flowchart of an example method for processing user requests accurately and with low latency. Detailed Implementation

[0024] Figure 1 This is a block diagram of an example system 100 in which the techniques of this disclosure for processing user requests can be implemented. Example system 100 includes a computing system 102, a client device 104, one or more servers 106, and a network 110. The computing system 102 may be located remotely from the client device 104 and / or the server 106, and is communicatively coupled to the client device 104 and the server 106 via the network 110.

[0025] Network 110 can be a single communication network (e.g., the Internet), or in some implementations may include one or more additional networks. As an example only, network 110 may include a cellular network, the Internet, and a server-side local area network (LAN). Although Figure 1 Only a single client device 104 is shown, but it should be understood that the computing system 102 may also communicate with multiple (e.g., thousands) other client devices that are typically similar to client device 104.

[0026] Generally, computing system 102 can perform user request processing services for entities (e.g., as a service / feature in a more comprehensive set of services), and client device 104 is typically configured to receive and provide information via one or more user interfaces (e.g., user interfaces of web pages and / or mobile applications or other applications) that implement interaction with computing system 102. For example, the entity associated with computing system 102 may be a provider of digital advertising services, and the entity associated with client device 104 may be a specific advertiser with an account with that provider. In such a case, a user of client device 104 can create and / or monitor one or more advertising campaigns associated with an advertising account (e.g., by creating new campaigns, placing ad groups, creating keywords, monitoring ad performance, setting bids, etc.) via the user interface of a website or application hosted or provided by computing system 102. As discussed further below, the user interface supported by the computing system 102 and displayed at the client device 104 may include functionalities and / or tools, such as areas of the user interface dedicated to chat interaction, that enable the user of the client device 104 to enter requests for help (e.g., queries or instructions) and view responses generated and sent by the computing system 102.

[0027] The computing system 102 includes a network interface 120, a processor 122, and a memory 124. The network interface 120 includes hardware, firmware, and / or software configured to enable the computing system 102 to exchange electronic data with client device 104 (and possibly other similar client devices) and server 106 via network 110. For example, the network interface 120 may include a wired or wireless router and a modem. The processor 122 may be a single processor (e.g., a central processing unit (CPU)) or may include multiple processors (e.g., multiple CPUs, or one or more CPUs and one or more graphics processing units (GPUs)). The computing system 102 may be a single computing device at a single location (e.g., a server) or may include multiple coordinated computing devices (e.g., multiple servers) distributed co-located or remotely.

[0028] Memory 124 is a computer-readable, non-transitory storage medium, cell, or device, or a collection of such media / cells / devices, and may include persistent and / or non-persistent memory components. Memory 124 stores instructions executable by processor 122 to perform various operations, including instructions for various software applications and data generated and / or used by such applications. Figure 1In example system 100, memory 124 stores user interface (UI) components 130, an orchestrator 132, multiple dedicated AI agents 134, a model 136, and tools 138. These components may all reside on a single server of computing system 102, or they may be distributed across two or more servers (e.g., servers geographically separated from each other). The operation of these components, depending on various implementations and / or scenarios, will be referenced below. Figures 2 to 8 This will be discussed in more detail. However, in general, orchestrator 132 is configured to centrally process user requests received from devices such as client device 104 by selectively identifying and using / invoking specific AI agents in dedicated AI agents 134, which are then specifically configured (e.g., trained and / or fine-tuned, and / or associated with different document databases, etc.) to analyze queries in different relevant domains. Tool 138 may include one or more dedicated components, modules, models, etc., that can be invoked as needed (e.g., by dedicated AI agents 134 and / or directly by planner component 140) to respond to specific requests / queries. For example, each of tools 138 may expose a corresponding application programming interface (API) for accessing that tool. See below for example. Figure 3 Specific examples of dedicated AI agents 134 and tools 138 are discussed.

[0029] UI component 130 communicates via network 110 with client device applications (e.g., with application 170 when executing on client device 104) to populate user interfaces and receive information (including requests) entered by the user through those user interfaces. To process user requests entered via the user interfaces, orchestrator 132 includes planner component 140 and response component 142 that respond to requests using "speculation" and "formulation" methods, respectively. For example, planner component 140 and response component 142 may represent separate models, different functions of a single model, or a model and other associated components. Planner component 140 generates a response development plan for a given user request by at least partially determining which of the dedicated AI agents 134 are relevant to the user request, and then queries the relevant dedicated AI agents among the dedicated AI agents 134 based on the response development plan. To this end, planner component 140 may issue synthetic queries to different dedicated AI agents 134 in parallel and / or serially (e.g., generating a synthetic query for a second dedicated AI agent using feedback from a first dedicated AI agent). In some implementations and / or scenarios (e.g., for certain response development plans), the planner component 140 can participate in multi-turn dialogues with each of one or more dedicated AI agents 134 (e.g., as follows regarding...). Figure 4One or more of the following (discussed) and / or dedicated AI agents 134 can be adapted to provide additional and / or different information to the orchestrator 132 as feedback (e.g., as discussed below regarding Figure 5 (As discussed). The response component 142 generates a response to the user directly or indirectly based on feedback from the queried dedicated AI agent in the dedicated AI agent 134 (e.g., after receiving feedback directly, or after another component or function of the planner component 140 or orchestrator 132 has aggregated, merged, filtered and / or otherwise manipulated the feedback).

[0030] In some implementations, planner component 140 and / or response component 142 include corresponding AI agents, and / or orchestrator 132 includes AI agents communicating with planner component 140 and / or response component 142. For example, the top-level AI agent of orchestrator 132 can analyze user requests and responsively generate synthetic queries for the AI ​​agents of planner component 140, wherein the synthetic queries prompt the AI ​​agents of planner component 140 to formulate a response development plan and generate appropriate synthetic queries. In addition to (or as part of) AI agents, planner component 140 may also include an execution engine that uses the synthetic queries generated by planner component 140 to invoke / prompt the relevant dedicated AI agent 134.

[0031] Each of the AI ​​agents in System 100 (e.g., dedicated AI agent 134, the AI ​​agent of planner component 140, the AI ​​agent of response component 142, the top-level AI agent of orchestrator 132, etc.) may use one or more types of machine learning models included in Model 136, such as large language models (LLM) or multimodal LLM (MLLM); or possibly, for some of the dedicated AI agents in Dedicated AI agents 134, classification models such as neural networks, support vector machines, clustering models, regression models, decision trees, etc. may be used. Each of some or all of the models in Model 136, and at least some of the models used by the dedicated AI agents 134, may be specifically trained (e.g., initial training and / or fine-tuning) to provide responses in their specific domain (e.g., using supervised learning techniques), and / or may be otherwise configured or arranged to provide responses in their specific domain (e.g., for a Retrieval Augmentation (RAG)-based AI agent as discussed below, by accessing a database of documents particularly relevant to a specific domain). In some implementations, the planner component 140 includes an AI agent using a model (in model 136) trained with supervised learning techniques, where manually created (or manually checked) queries / responses are used as ground truth (e.g., where the queries / responses indicate that a particular skill / topic / specialized AI agent is relevant to a particular user request, etc.).

[0032] In some implementations, when using RAG or other suitable techniques to resolve synthetic queries from planner component 140, one or more of the dedicated AI agents 134 may utilize internal and / or external information sources. In example system 100, account database 180 may store information about the current state of accounts of advertisers or other entities (e.g., activity entries / settings such as bid amounts, keywords, etc., and / or creatives / ads associated with advertisers' activities), and knowledge corpus 182 may include other information that may be useful to a particular dedicated AI agent in dedicated AI agent 134 (e.g., analytics data, policy guidelines, general business information, etc.). Although shown as two separate local data stores 180 and 182, it should be understood that account database 180 and / or knowledge corpus 182 may be stored locally (e.g., in memory 124), remotely, and / or in a distributed manner. In some implementations, each of one or more dedicated AI agents 134 is configured to access only a subset of the knowledge corpus 182 (and / or only a subset of the account database 180), such that the AI ​​agent can only retrieve (e.g., using RAG) documents that are more likely to be relevant to the AI ​​agent's domain. This specialization of the available document database can reduce illusions and / or improve efficiency by isolating AI agents from at least some information that may not be relevant to their domain.

[0033] As used herein, "AI agent" may, in some implementations, include one or more AI sub-agents (and / or peer AI agents) with finer-grained functionality, and "model" may, in some implementations, include one or more sub-models (and / or peer models). For example, the optimizing agent in dedicated AI agent 134 may include a first sub-agent specifically assuming an optimization strategy and a second sub-agent specifically collecting activities and / or other data related to those optimization strategies.

[0034] In some implementations, and such as Figure 1As shown in Example System 100, orchestrator 132 maintains configuration data 146 (e.g., configuration data stored in memory 124 or memory at one or more remote locations). Configuration data 146 may include data (parameter settings, etc.) entered by a user (e.g., an administrator using a client device as part of computing system 102) to configure orchestrator 132 for a specific use case or specific product / service. For example, configuration data 146 may indicate that a first subset of dedicated AI agents 134 and / or tools 138 will be available for selection / use by planner component 140 when a user on a client device (e.g., client device 104) enters a user request related to an advertising campaign management product, and may indicate that a second subset (possibly overlapping with the first subset) of dedicated AI agents 134 and / or tools 138 will be available for selection / use by planner component 140 when a user on a client device alternatively enters a user request related to an advertising analytics product.

[0035] In some implementations, and such as Figure 1As shown in Example System 100, Orchestrator 132 maintains User Context Data 150 and Agent Context Data 152 (e.g., stored in memory 124 or memory at one or more remote locations). User Context Data 150 typically provides planner component 140 (and possibly response component 142) with information that facilitates the generation of a response development plan (and thus the final user response) that is more relevant to and / or better suited to the context of the user request. For example, User Context Data 150 may include the chat history of the requesting user recorded in a single session or multiple sessions (e.g., earlier user requests from client device 104 and corresponding responses from response component 142), and / or data associated with a specific entity (e.g., characteristics of the advertiser or other entity associated with the user request, information associated with the business of such an entity, information associated with the account of such an entity, etc.). In some implementations, User Context Data 150 may also, or alternatively, include information about the current state of the user interface being presented to the user on client device 104, for example, as discussed in more detail below. As an example only, if user context data 150 indicates that a user on client device 104 has entered chat history stating that the advertiser does not want to increase the budget (or the user repeatedly ignores suggestions to increase the budget from response component 142, etc.), then planner component 140 can generate a synthetic query instructing one of the dedicated AI agents 134 (e.g., the optimization agent) to avoid recommendations involving increasing the advertiser's budget. As another example, if user context data 150 indicates that a user on client device 104 has entered chat history expressing frustration with frequent ad rejections, then planner component 140 can generate a response development plan that invokes the strategy agent of dedicated AI agent 134, even if the content of the corresponding user request would not normally trigger an invocation of the strategy agent.

[0036] Agent context data 152 includes data maintained (e.g., created and updated) by orchestrator 132 (e.g., planner component 140) for one, some, or all of the dedicated AI agents 134. Specifically, agent context data 152 may include a history reflecting previous communications with the dedicated AI agents (e.g., synthesized queries and corresponding responses or other feedback from the dedicated AI agents), and planner component 140 may analyze the relevant history as needed to facilitate the development of a response development plan. As an example only, if planner component 140 determines that a policy agent is relevant to a specific user request, and if the context / history maintained by planner component 140 indicates that the policy agent has frequently returned, and most recently returned, a synthesized query response indicating that a policy criterion involving sensitive content is relevant, then planner component 140 may decide to include a call to the sensitive content agent in the response development plan the next time planner component 140 receives a user request involving any policy issue. In some implementations, the planner component 140 maintains separate / isolated agent contexts / histories for different dedicated AI agents among some or all of the dedicated AI agents 134, and uses that AI agent's context when generating synthetic queries for a specific AI agent, but not when generating synthetic queries for any other AI agent. This centralized maintenance of isolated agent-specific contexts / histories effectively provides a much wider context window (or equivalently, eliminates the need for dedicated AI agents 134 to maintain larger contexts themselves), while also reducing the likelihood of illusions by allowing each of the dedicated AI agents 134 to focus more specifically on its own domain without cross-context contamination (e.g., not having to analyze / consider unwanted data / history, or being more relevant to other dedicated AI agents 134, in their synthetic queries).

[0037] In some implementations, orchestrator 132 (e.g., planner component 140) is configured (e.g., trained) to resolve conflicts between different dedicated AI agents in dedicated AI agent 134 to avoid inefficiency and further reduce illusions. As an example only, if planner component 140 identifies a conflict between feedback from two dedicated AI agents, planner component 140 may ignore feedback from the AI ​​agents that planner component 140 determines is less relevant to / less important to the response (or relevant portions of the feedback).

[0038] Client device 104 may be or include any fixed, mobile, or portable computing device with wired and / or wireless communication capabilities (e.g., smartphone, tablet, laptop, desktop computer, smart wearable device (such as smart glasses or smartwatch), vehicle main unit computer, etc.). Figure 1 In an example implementation, client device 104 includes a network interface 160, a processor 162, a memory 164, and a display 166. Processor 162 may be a single processor or may include multiple processors. Memory 164 includes one or more computer-readable non-transitory storage media, cells, or devices, which may include persistent and / or non-persistent memory components. Memory 164 stores instructions executable by processor 162 to perform various operations, including instructions for various software applications and data generated and / or used by such applications.

[0039] exist Figure 1 In example system 100, memory 164 stores at least application 170. For example, application 170 may be a web browser application for accessing a website hosted by computing system 102, or a dedicated application (e.g., a mobile application) downloaded from computing system 102 or otherwise associated with an entity related to computing system 102. Generally, application 170 is executed by processor 162 to present one or more user interfaces via display 166, wherein the user interface enables a user to enter user requests to be analyzed by planner component 140, view responses to user requests generated by response component 142, and possibly also facilitate various other tasks or operations (e.g., user changes to activity settings, user monitoring of activity performance, etc.). Example user interfaces that may be supported by UI component 130 (and more generally, by orchestrator 132) are described below. Figures 7A to 7D Let's have a discussion.

[0040] Display 166 includes hardware, firmware, and / or software configured to enable a user to view visual output from client device 104, and can use any suitable display technology (e.g., LED, OLED, LCD, etc.). In some implementations, display 166 is integrated into a touchscreen that has both display and manual input capabilities. Although Figure 1 Not shown, but client device 104 may also include one or more audio output devices or components, such as one or more speakers (e.g., for presenting a response to a user request in an audio format).

[0041] Network interface 160 includes hardware, firmware, and / or software configured to enable client device 104 to exchange electronic data with computing system 102 via network 110. For example, network interface 160 may include a cellular transceiver, a WiFi transceiver, and / or a transceiver for one or more other wired and / or wireless communication technologies.

[0042] Server 106 may include any external (e.g., third-party) server with which computing system 102 communicates to obtain data useful in generating responses to user requests. For example, in an implementation where dedicated AI agent 134 includes an analytical agent, the analytical agent may request relevant analytical data from the analytical server in server 106, or invoke tools (in tool 138) that obtain relevant analytical data from the analytical server. In other implementations, system 100 omits server 106.

[0043] Figure 2 It describes what can be made by Figure 1 The computing system 102 is implemented as an example architecture 200 to handle user requests. Figure 2 In the diagram, the arrows indicate the directionality of communication between layers in architecture 200.

[0044] Example architecture 200 includes a UI layer 210 implemented by UI component 130 above an orchestrator layer 220 implemented by orchestrator 132. Orchestrator layer 220 includes a planner layer 222 implemented by planner component 140 above a responsive layer 224 implemented by responsive component 142. Alternatively, planner layer 222 and responsive layer 224 may reside at the same level in the stack of architecture 200. Orchestrator layer 220 resides above a dedicated AI agent layer 230 implemented by dedicated AI agent 134, which in turn resides above a tool layer 240 implemented by tool 138. Alternatively, dedicated AI agent layer 230 and tool layer 240 may reside at the same level in the stack of architecture 200 (or wherein dedicated AI agent layer 230 is partially above tool layer 240 and partially at the same level as tool layer 240, etc.).

[0045] In the example architecture, UI layer 210 transmits user requests (and possibly other information, such as screenshots or other current UI states) to planner layer 222, which then generates a response development plan and queries the relevant dedicated AI agents in dedicated AI agent 134 based on that plan. Figure 2As indicated by the leftmost double-headed arrow, the dedicated AI agent layer 230 (i.e., the relevant dedicated AI agent in dedicated AI agent 134) can provide feedback (e.g., responses) to the planner layer 222, where the synthesized query and response potentially have multiple loops / rounds. The planner layer 222 can then provide feedback to the response layer 224 (or a synthesized query generated by the planner component 140 based on the feedback). Alternatively or additionally (and as...) Figure 2 As indicated by the dotted line in the diagram, the dedicated AI agent layer 230 can provide feedback directly to the response layer 224. Similarly, the dedicated AI agent layer 230 can invoke one or more of the tools 138 by communicating with the tool layer 240, and tool invocation and response may potentially have multiple loops / rounds.

[0046] After the orchestrator layer 220 (planner layer 222 or response layer 224) has received all responses anticipated under the response development plan from the dedicated AI agent layer 230 (excluding any responses that may have timed out due to non-responsive AI agents, etc.), the orchestrator 132 or planner component 140 can generate a synthesized query to be sent to the response layer 224, whereby the response component 142 then generates a response to the user request and sends the response to the UI layer 210. The UI layer 210 can then cause the user interface of the application 170 at the client device 104 to present the response to the user via the display 166 (and / or as audio output, etc.).

[0047] It should be understood that other implementation methods can support more than Figure 2 Inter-layer communication between more or fewer layers is shown in the diagram, and / or more than Figure 2 Inter-layer communication in more or fewer directions is shown. Moreover, in some implementations, architecture 200 may include one or more additional layers (e.g., a model layer implemented by model 136 between layers 230 and 240) and / or omit one or more layers (e.g., tool layer 240).

[0048] Figure 3 It describes what can be made by Figure 1 The computing system 102 and / or in Figure 2 An example set 300 of dedicated AI agents 302 and tools 304 used within the architecture 200. For example, the dedicated AI agent 302 and tool 304 can be respectively Figure 1 Dedicated AI agent 134 and tools 138.

[0049] exist Figure 3In the example set 300, the dedicated AI agent 302 includes reporting agents, billing agents, optimization agents, diagnostic agents, policy agents, sensitive content agents, creative agents, search engine agents, and web crawling agents. For example, the depicted dedicated AI agent 302 can be specifically configured (trained, fine-tuned, provided with access to certain databases for document retrieval, etc.) to perform the following operations:

[0050] The reporting agent responds to synthetic queries from planner component 140 by: (1) planning locally what data / information to collect (e.g., advertising performance / analysis data from knowledge corpus 182), (2) using one or more of tools 304 to collect the desired data / information, (3) using LLM or MLLM of model 136 to reason based on the collected data / information and synthetic queries, and / or (4) using the reporting agent's response layer model to formulate a response to planner component 140 (e.g., a concise presentation of relevant parts, summaries, etc. of the collected information), and / or providing other feedback as discussed below (e.g., basic data).

[0051] The billing agent responds to synthetic queries from planner component 140 by: (1) determining which data / information to collect locally (e.g., past budget / cost data from knowledge corpus 182), (2) using one or more of tools 304 to collect the desired data / information, (3) using LLM or MLLM of model 136 to reason based on the collected data / information and synthetic queries, and / or (4) using the billing agent's response layer model to formulate a response to planner component 140 (e.g., a concise presentation of relevant portions, summaries, etc. of the collected information), and / or providing other feedback as discussed below (e.g., basic data).

[0052] The optimizing agent responds to synthetic queries from planner component 140 by: (1) planning locally what data / information to collect (e.g., activity setup information, advertising performance / analysis data, etc. in account database 180), (2) using one or more of tools 304 to collect the desired data / information, (3) using LLM or MLLM of model 136 to reason based on the collected data / information and synthetic queries, and / or (4) using the response layer model of the optimizing agent to formulate a response to planner component 140 (e.g., proposed ways to improve past performance as indicated by the collected information), and / or providing other feedback as discussed below (e.g., basic data).

[0053] The diagnostic agent responds to synthetic queries from planner component 140 by: (1) planning locally what data / information to collect (e.g., notification of errors / problems with accounts in account database 180), (2) using one or more of tools 304 to collect the desired data / information, (3) using LLM or MLLM of model 136 to reason based on the collected data / information and synthetic queries, and / or (4) using the response layer model of the diagnostic agent to formulate a response to planner component 140 (e.g., a proposed way to fix errors / problems indicated by the collected information), and / or providing other feedback as discussed below (e.g., basic data).

[0054] The strategy agent responds to synthetic queries from planner component 140 by: (1) planning locally what data / information to collect (e.g., notification of strategy rules from knowledge corpus 182, policy-based rejection of ads in the activities of advertisers indicated in account database 180 or knowledge corpus 182, etc.), (2) collecting the desired data / information using one or more of tools 304, (3) reasoning based on the collected data / information and synthetic queries using LLM or MLLM of model 136, and / or (4) formulating a response to planner component 140 using the strategy agent's response layer model (e.g., a concise explanation of why an ad is not approved), and / or providing other feedback as discussed below (e.g., basic data).

[0055] Sensitive Content Agent: Responds to synthetic queries from Planner Component 140 by: (1) planning locally what data / information to collect (e.g., notification of specific policy rules related to sensitive content from Knowledge Corpus 182, policy-based rejection of ads in the activities of advertisers indicated in Account Database 180 or Knowledge Corpus 182, etc.), (2) collecting the desired data / information using one or more of Tools 304, (3) reasoning based on the collected data / information and synthetic queries using LLM or MLLM of Model 136, and / or (4) formulating a response to Planner Component 140 using the Sensitive Content Agent's response layer model (e.g., a concise explanation of why an ad is not approved due to sensitive content, or why a question or instruction in a user request raises a potential sensitive content issue, etc.), and / or providing other feedback as discussed below (e.g., basic data). For example, the Sensitive Content Agent may resemble a Policy Agent but be more specialized.

[0056] The creative agent responds to synthetic queries from planner component 140 by: (1) planning locally what data / information to collect (e.g., specific creatives for advertisers in account database 180), (2) using one or more of tools 304 to collect the desired data / information, (3) using LLM or MLLM of model 136 to reason based on the collected data / information and synthetic queries, and / or (4) using the creative agent’s response layer model to formulate a response to planner component 140 (e.g., a concise explanation of why an advertisement was not approved), and / or providing other feedback as discussed below (e.g., basic data).

[0057] The search engine agent responds to synthetic queries from planner component 140 by: (1) formulating one or more search queries, (2) running one or more queries on the search engine (e.g., directly or using one of tools 138), (3) using LLM or MLLM of model 136 to reason based on search results and synthetic queries, and / or (4) using the response layer model of the search engine agent to formulate a response to planner component 140 (or simply providing the top N search results, etc.), and / or providing other feedback as discussed below (e.g., underlying data).

[0058] The web crawling agent responds to synthetic queries from planner component 140 by: (1) specifying web crawling parameters (e.g., a list of URLs to crawl, and / or keywords indicating the type of information expected), (2) crawling the web based on the web crawling parameters (e.g., directly or using one of tools 138), (3) using the LLM or MLLM of model 136 to reason based on the web crawling results and synthetic queries, and / or (4) using the response layer model of the web crawling agent to specify a response to planner component 140 (or simply providing the results from the web crawling, etc.), and / or providing other feedback as discussed in the question (e.g., basic data).

[0059] As described above, in some implementations and / or scenarios, the planner component 140 can formulate a response development plan, wherein the planner component 140 queries all relevant dedicated AI agents in the dedicated AI agent 134 in parallel, while in other implementations and / or scenarios, the planner component 140 specifies a sequential order or some other more complex sorting (e.g., fanning out parallel synthetic queries to the reporting agent and the diagnostic agent, and then using additional synthetic queries based on the responses from the reporting agent and the diagnostic agent to query the optimization agent, etc.). Furthermore, in some implementations and / or scenarios, a first dedicated AI agent in the dedicated AI agent 134 can invoke a sub-agent, which can be, for example, a different second dedicated AI agent in the dedicated AI agent 134. For example, the diagnostic agent can invoke the policy agent, and / or the policy agent can invoke the sensitive content agent, etc.

[0060] Also in Figure 3 In the example set 300, tool 304 includes event information tools, notification tools, and website information tools. For example, the depicted tool 304 can be configured to perform the following actions (e.g., using the corresponding API) when invoked:

[0061] Event information tool: Retrieve information about a specific event (e.g., from the account database 180), such as event entries / settings (keywords, bid amount, etc.), costs, etc.

[0062] Notification tools: Retrieve information about current notifications associated with a specific account and / or activity (e.g., from the account database 180), such as ad disapproval notifications, budget notifications, etc.

[0063] Website information tools: Retrieve information about a specific website (e.g., URL or set of URLs), such as by crawling information (product descriptions, etc.) from landing pages associated with a specific advertisement.

[0064] Screen capture tool: retrieves screenshots of the user interface and / or other user interface state information to indicate which screen or part of the screen is currently being viewed (e.g., being viewed by a user on client device 104).

[0065] One or more of the dedicated AI agents 302 (e.g., the planning layer / model of the AI ​​agent) can invoke / use one or more of the tools 304 as appropriate for a specific query. In some implementations, the planner component 140 itself may also, or alternatively, directly invoke / use some or all of the tools 304. It should be understood that the dedicated AI agent 302 may include more than Figure 3 The document shows more, fewer, and / or different AI agents, and tool 304 may include more than Figure 3 More, fewer and / or different tools are shown in the document.

[0066] Figure 4 This indicates that in an example scenario and implementation, it is possible Figure 1 The computing system 102 and / or Figure 2 The message passing graph within architecture 200 shows the sequence of messages exchanged within architecture 400. Figure 4 For example, orchestrator 402 may represent orchestrator 132 or orchestrator layer 220, and dedicated AI agent 404 may represent one of dedicated AI agent 134 or 302 or dedicated AI agent layer 230.

[0067] exist Figure 4 In the example implementation and scenario, orchestrator 402 (e.g., planner component 140) generates a first synthetic query (based on a user request) and sends this first synthetic query 410 to a dedicated AI agent 404. The first synthetic query can be designed to directly elicit the desired final response from the dedicated AI agent 404 (e.g., if the response development plan anticipates a usable response within at least one dialogue round / turn), or it can be designed to gather information useful for generating a next second synthetic query for the same dedicated AI agent 404 (e.g., if the response development plan anticipates requiring multiple dialogue rounds / turns). In either case, Figure 1 In the scenario depicted, orchestrator 402 (e.g., planner component 140) determines that an additional query is needed after receiving a first response 412 from dedicated AI agent 404, and thus generates a second synthesized query. Depending on the implementation, the first response 412 may be received by planner component 140 or response component 142. For example, planner component 140 may receive the first response and determine that the first response lacks some necessary information, or the response development plan may specify subsequent queries regardless of the content of the first response, etc.

[0068] Therefore, orchestrator 402 (e.g., planner component 140) generates a second synthetic query and sends it 414 to dedicated AI agent 404, and receives a corresponding second response 416 from dedicated AI agent 404 (e.g., via planner component 140 or response component 142). Depending on the scenario (or implementation, e.g., if orchestrator 132 allows a maximum of two rounds / rounds of communication with dedicated AI agent 404 to limit latency), orchestrator 402 may or may not have one or more additional rounds / rounds of communication with dedicated AI agent 404. In some implementations, prompts to planner component 140 (e.g., generated by a higher-level AI agent of orchestrator 132) instruct planner component 140 to limit communication to no more than a certain number of rounds / rounds, or instruct planner component 140 to suppress the generation and / or sending of another query when planner component 140 deems the latest response to have a certain degree of usefulness, etc.

[0069] like Figure 4 As seen, the dedicated AI agent 404 may include its own response model 406 and one or more lower-level models 408. For example, the response model 406 may be a relatively small / weak LLM, and the lower-level models may be one or more larger, more powerful LLMs. The dedicated AI agent 404 may use the lower-level models 408 to perform intermediate inference while using the response model 406 to generate a "final" response based on that inference. For example, the lower-level model 408 may analyze a first synthetic query to generate first intermediate inference, and the response model 406 may generate a first response based on the first intermediate inference. Similarly, the lower-level model 408 may analyze a second synthetic query to generate second intermediate inference, and the response model 406 may generate a second response based on the second intermediate inference. The lower-level model 408 may generate intermediate inference based on underlying data, which includes information and / or other information from the synthetic query from the orchestrator 402 (e.g., one or more documents retrieved by the dedicated AI agent 404 using the RAG architecture).

[0070] Figure 5 This indicates that alternative implementations that can further reduce latency and / or illusions are possible. Figure 1 The computing system 102 and / or Figure 2 The message passing graph shows another sequence of 500 messages exchanged within the architecture of 200. Figure 5 For example, orchestrator 502 can represent orchestrator 132 or orchestrator layer 220, and dedicated AI agent 504 can represent one of dedicated AI agents 134 or 302 or dedicated AI agent layer 230. Furthermore, agent-side response model 506 and lower-level model 508 can be similar to... Figure 4The response model 406 and the lower-level model 408.

[0071] exist Figure 5 In the example implementation and scenario, orchestrator 502 (e.g., planner component 140) generates a first synthesized query (based on a user request) and sends this first synthesized query to dedicated AI agent 504 via 510. However, in Figure 5 In some implementations, the dedicated AI agent 504 is adapted to send 512 intermediate inferences generated by the lower-level model 508 based on the synthetic query, and the underlying data on which the lower-level model 508 performs intermediate inferences, to the orchestrator 502 (e.g., to the response component 142). In some implementations, the dedicated AI agent 504 also sends 512 instructions on how to read / interpret the underlying data, which may be in a format that the orchestrator 502 would not originally know.

[0072] exist Figure 5 In the implementation shown, the dedicated AI agent 504 does not send any response generated by the response model 506 based on intermediate inference. Therefore, higher-level response formulation is effectively delegated to the orchestrator 502 (e.g., response component 142), which generates a response based on intermediate inference and underlying data (and possibly instructions on how to read the underlying data) but not on any response from the response model 506. In some implementations, the dedicated AI agent 504 bypasses the response model 506 because a higher-level response is not required at the agent end. In other implementations, the dedicated AI agent 504 completely omits the response model 506. In other implementations, the dedicated AI agent 504 sends a higher-level response 512 with underlying data and intermediate inference (and possibly instructions), but the orchestrator 502 (e.g., response component 142) ignores the higher-level response for the purpose of formulating a response to a user request. In any of these implementations, the orchestrator 132 may be provided with sufficient low-level information to (at least in some cases) prevent re-queries such as... Figure 4 The dedicated AI agent 504 shown is shown to have any need for this. By removing one loop / round (or multiple loops / rounds) in the communication between the orchestrator 502 and the dedicated AI agent 504, such an implementation can reduce the total latency of the response provided to the user, while also reducing the use of processing power. Moreover, and contrary to intuition (in the context of the larger domain expertise of the given dedicated AI agent 504), removing (or ignoring, etc.) the layer in which the dedicated AI agent 504 formulates its response to the synthesized query from the orchestrator 502 can reduce the illusion, since conclusions or inferences drawn by the response model of the dedicated AI agent (e.g., response model 506) can be a significant source of illusion.

[0073] In some implementations, all dedicated AI agents 134 or 302 are used. Figure 5 The technology delegates higher-level response formulation to orchestrator 502 (e.g., to response component 142). In other implementations, only a subset (one or more dedicated AI agents) of dedicated AI agents 134 or 302 uses this technology. Figure 5 The technology used by [the AI ​​agent], while other specialized AI agents use other technologies (e.g., Figure 4 (The technology). Moreover, in some implementations, higher-level response delegation is determined on a case-by-case basis (e.g., by planner component 140, which may indicate where delegation is needed in the response development plan and / or in the synthetic query to the specialized AI agent; or by the specialized AI agent itself).

[0074] Figure 6 An example process 600 is described that can be implemented by each of one or more dedicated AI agents (e.g., dedicated AI agents 134 or 302). Figure 6 This corresponds to an implementation of an LLM (e.g., the LLM of Model 136) in which a dedicated AI agent uses an LLM specific to that AI agent (i.e., specifically trained and / or otherwise configured to handle queries in a specific domain).

[0075] exist Figure 6 In this process, synthetic query 602 (e.g., generated by planner component 140) is received as input by a dedicated AI agent, which, in response, constructs an LLM request (e.g., a prompt) at stage 604. At stage 606, the dedicated AI agent uses the LLM to generate an answer to the LLM request, and at stage 608, the dedicated AI agent determines whether a tool call exists in the answer and / or whether at least a threshold number of LLM calls have already been made for synthetic query 602. If a tool call exists and the threshold has not been reached, the process proceeds to stage 610. At stage 610, the dedicated AI agent uses the tool processor to invoke the appropriate tool in tool 138. If no tool call exists and / or the threshold has been reached, the process proceeds to stage 612, in which case the dedicated AI agent responds to the synthetic query using information provided by the tool. For example, stage 612 may include sending, for example, Figure 4 A higher-level response, or sending such as Figure 5 The underlying basic data and intermediate inference in the middle.

[0076] Figures 7A to 7D It describes various example scenarios and implementation methods. Figure 1 UI component 130 can cause various user interfaces to be presented to the user (e.g., to the user of client device 104 via application 177 and display 166).

[0077] First refer to Figure 7A The user interface 700 includes a work area 702 and a chat area 704. The work area 702 may support services or features of a digital advertising platform provider, while the chat area 704 allows the user to enter requests and view responses from the orchestrator 132. Specifically, the chat area 704 includes a user request field 710 where the user can enter their request and a response area 712 that displays the response. In some implementations, the chat area 704 also provides a further user request field to allow the user to continue their conversation with the orchestrator 132.

[0078] exist Figure 7A In the example, a user asks, “What is my top performing ad and how can I improve my ads further?”, and orchestrator 132 (at least partially) responds as shown. Specifically, the response displays an identifier for the advertiser’s top performing ad, along with performance metrics (click-through rate and number of clicks) used for that ad. The response in this example also provides at least one suggestion for improving the advertiser’s campaign, wherein interactive control 714, when activated by the user, causes additional details to be displayed (e.g., by activating a deep link to a screen / page associated with a product or service in work area 702).

[0079] As an example, refer to Figure 1 and 3 The response in response area 712 can be generated as follows: 1) Based on the user request in field 710, planner component 140 generates a response development plan specifying that the reporting agent and optimization agent should be invoked in parallel; 2) Planner component 140 generates and sends corresponding synthetic queries for the reporting agent and optimization agent; 3) The reporting agent invokes the activity information tool based on its query to obtain performance metrics for all ads for the advertiser; 4) Based on the performance metrics, the reporting agent identifies the best-performing ads and provides feedback to orchestrator 132 including at least the ad identifier and the corresponding performance metric; 5) In parallel with the operation of the reporting agent and based on its own query, the optimization agent invokes the activity information tool to obtain activity settings and performance metrics for the advertiser; 6) Based on the activity settings and performance metrics, the optimization agent formulates one or more optimization strategies and provides feedback to orchestrator 132 indicating the optimization strategies; and 7) Response component 142 formulates the text shown in response area 712 based on the feedback from the strategy agent and optimization agent. In some implementations, the reporting agent and / or optimizing agent provide basic data and intermediate inference, as described above. Figure 5 As stated above.

[0080] Next reference Figure 7B The user interface 720 includes a work area 722 and a chat area 724. The work area 722 may support services or features of a digital advertising platform provider, while the chat area 724 allows users to enter requests and view responses from the orchestrator 132. For example, Figure 7A Chat area 704 and chat area 724 (and / or Figure 7C and / or Figure 7D The chat area (724) can be the same chat area, that is, a persistent chat area that remains in place as the user navigates to different screens. Chat area 724 includes a user request field 730 in which the user can enter his or her request and a response area 732 that displays the response.

[0081] exist Figure 7B In the example, the user instruction is "Fix the policy violation for this disapprovedad," and orchestrator 132 (at least partially) responds as shown in the figure. Specifically, the response indicates the specific reason why the ad shown in work area 722 was not approved. This example response also provides an interactive control 734, which, when activated by the user, causes the system to conform to the policy guidelines, or at least attempts to do so.

[0082] As an example, refer to Figure 1 and Figure 3The response in response area 732 can be generated as follows: 1) Based on the user request in field 730, planner component 140 generates a response development plan specifying that the policy agent will be invoked; 2) The policy agent or planner component 140 invokes a screen capture tool to obtain a screenshot of the user interface as seen by the user of client device 104; 3) Planner component 140 generates and sends a synthetic query for the policy agent; 4) Based on its query and the screenshot (or information derived by planner component 140 from the screenshot), the policy agent invokes a notification tool to obtain the reason why the advertisement currently displayed in work area 722 was not approved; 5) Based on the reason for non-approval, the policy agent generates a description of the reason for non-approval and identifies remedial actions to resolve the reason; and 6) Response component 142 formulates the text shown in response area 732 based on feedback from the policy agent. User activation of the user interaction control 734 can cause the orchestrator 132 (or dedicated AI agent 134 or tool 138, etc.) to communicate with another system or platform (e.g., via an API), which in turn enables the specific changes described in response area 732 (possibly after the user provides further confirmation). In some implementations, the policy agent provides the underlying data and intermediate inference, as described above. Figure 5 As stated above.

[0083] Next reference Figure 7C The user interface 740 includes a work area 742 and a chat area 744. The work area 742 may support the services or features of a digital advertising platform provider, while the chat area 744 allows the user to enter requests and view responses from the orchestrator 132. The chat area 744 includes a user request field 750 where the user can enter his or her request and a response area 752 that displays the response.

[0084] exist Figure 7C In the example, a user asks, “Which campaigns need my attention?”, and orchestrator 132 (at least partially) responds as shown. Specifically, the response illustrates the specific reasons why a particular advertiser’s campaign could benefit from improvements, and possible ways to achieve at least some of those improvements. The response in this example also provides an interactive control 754, which, when activated by the user, causes the system to conform the advertisement to policy guidelines, or at least attempts to do so.

[0085] As an example, refer to Figure 1 and Figure 3The response in response area 752 can be generated as follows: 1) Based on the user request in field 750, planner component 140 generates a response development plan specifying that the reporting agent and optimization agent should be invoked sequentially; 2) planner component 140 generates a synthetic query and sends it to the reporting agent; 3) the reporting agent invokes the activity information tool based on its query to obtain performance metrics for all of the advertiser's activities (and / or for specific ads within them); 4) based on the performance metrics, the reporting agent generates a description of the advertiser's relatively poorly performing activities; 5) based on the description from the reporting agent, planner component 140 generates a synthetic query and sends it to the optimization agent; 6) the optimization agent identifies potential improvements for one or more of the poorly performing activities based on its query; and 7) response component 142 formulates the text shown in response area 752 based on feedback from the reporting agent and optimization agent. For example, user activation of user interaction control 754 can cause planner component 140 to generate another synthetic query for the optimization agent to obtain a more detailed optimization plan. Alternatively, user activation of the user interaction control 754 can activate deep links to screens / pages of products or services associated with the work area 742. In some implementations, the reporting agent and / or the optimizing agent provide basic data and intermediate inference, as described above. Figure 5 As stated above.

[0086] Next reference Figure 7D The user interface 760 includes a work area 762 and a chat area 764. The work area 762 may support another service or feature of the digital advertising platform provider, while the chat area 764 allows the user to enter requests and view responses from the orchestrator 132. The chat area 764 includes a user request field 770 where the user can enter his or her request and a response area 772 that displays the response.

[0087] exist Figure 7D In the example, a user asks, "Who else is selling this service near me?", and orchestrator 132 (at least in part) responds as shown. Specifically, the response lists other businesses geographically close to the advertiser's site based on one or more web sources.

[0088] As an example, refer to Figure 1 and Figure 3The response in response area 772 can be generated as follows: 1) Based on the user request in field 770, planner component 140 generates a response development plan specifying that an activity information tool (or another suitable tool) will be invoked, which will then invoke the search engine agent; 2) Planner component 140 uses the activity information tool (or other suitable tool) to obtain the advertiser's physical address or other location; 3) Planner component 140 generates a synthetic query including the location and sends the synthetic query to the search engine agent; 4) The search engine agent invokes the search engine's API (e.g., via tool 138 or 304) based on its query using another query indicating the location and receives the search results; 5) The search engine agent provides feedback indicating at least a subset of the relevant search results; and 6) Response component 142 formulates the text shown in response area 772 based on the feedback from the search engine agent. In some implementations, the search engine agent provides the underlying data and intermediate inference, as described above. Figure 5 As stated above.

[0089] As described above, in some implementations, orchestrator 132 (e.g., planner component 140) maintains one or more contexts, such as those provided by... Figure 1 The user context data 150 and / or agent context data 152 reflect this. As an example of user context data (besides...) Figure 7B (In addition to the screenshot example) Figure 7A In the scenarios and implementations, the planner component 140 can generate a response development plan not only based on the user request in field 710 but also based on chat history indicating the user's previous interest in performance over the last 30 days (e.g., causing the planner component 140 to generate a synthetic query for the reporting agent, which specifically inquires about the performance of the best-performing advertisement over the last 30 days). As an example of agent context data, in Figure 7B In the scenarios and implementations, the planner component 140 can generate a response development plan not only based on the user request in field 730, but also based on the history of queries / responses between the planner component 140 and the policy agent indicating that the user may refer to a specific advertisement (e.g., in an implementation where UI context / screen capture technology is not used).

[0090] As also mentioned above, in some implementations, the response component 142 responds to the initial user request at least in part by performing or triggering an action. For example, in Figure 7BIn the scenario and implementation, the response component 142 can cause the response area 752 to include interactive controls. User activation of the interactive controls can then cause the computing system 102 (or another device / system) to remotely control the input functions of the client device 104, thereby automatically selecting various controls in the work area 742, entering information in various fields, etc.

[0091] In some implementations, orchestrator 132 outputs text instructing its inference phase based on the response development plan, and computing system 102 can then send this text to client device 104 for real-time presentation to the user via application 170 and display 166. In this way, the planned steps can be more transparent to the user, and the user can find it easier to remain engaged during delays caused by processing latency. (Reference) Figure 7A For example, after a query is entered in the user request field 710 (when the reporting agent invokes the activity information tool), the text "Collecting performance metrics for all your ads" can initially be presented to the user. This text can then be replaced with the text "Looking for your best-performing ad" (when the reporting agent identifies the best-performing ad), etc.

[0092] Figure 8 This is a flowchart of an example method 800 for processing user requests accurately and with low latency. For example, method 800 could be... Figure 1 The computing system 102 (e.g., an orchestrator 132 executed by processor 122, and possibly a dedicated AI agent 134 and / or other components) is used to implement this.

[0093] At box 802, a user request (e.g., by planner component 140) is received via a user interface and from a client device (e.g., client device 104). At box 804, an orchestrator (e.g., orchestrator 132) generates a context-specific response to the user request. Box 804 includes boxes 806, 808, 810, and 812.

[0094] At box 806, an orchestrator (e.g., planner component 140 of orchestrator 132) generates a response development plan based on the user request and one or more contextual signals (e.g., from user context data 150 and / or agent context data 152). The response development plan specifies one or more dedicated AI agents (e.g., dedicated AI agents 134 or 302) from a plurality of dedicated AI agents by identifying one or more dedicated AI agents as relevant to the user request. Each of the dedicated AI agents is configured to analyze queries specific to its respective domain.

[0095] At box 808, the orchestrator queries one or more dedicated AI agents based on the response development plan. At box 810, in response to the query, it receives corresponding feedback from one or more dedicated AI agents. The feedback can be a higher-level response (e.g., similar to...). Figure 4 ) or may include intermediate inference and underlying data (e.g., similar to Figure 5 ).

[0096] At box 812, the orchestrator (e.g., response component 142 of orchestrator 132) generates a context-specific response based on the corresponding feedback received at box 810.

[0097] At box 814, a context-specific response is rendered on the client device via the user interface.

[0098] In other implementations, method 800 may include more or fewer boxes, and / or some boxes may be in a different manner than... Figure 8 The order in which the contents are shown. For example, box 812 may include an identifier of an action to be taken in response to a user request, and method 800 may include an additional box in which an action is caused (e.g., by remotely assuming control of the user interface at the client device, causing one or more operations to occur via the user interface).

[0099] As is apparent from the above description, some of the techniques disclosed herein utilize artificial intelligence to generate high-quality videos for specific products and audiences. Artificial intelligence (AI) is a branch of computer science focused on creating models that can perform tasks with minimal human intervention. For example, AI systems can leverage machine learning, natural language processing, and computer vision. Machine learning and subsets thereof (such as deep learning) focus on developing models that can infer outputs from data. Outputs may include, for example, predictions and / or classifications. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. AI systems may include generative models that generate new content, such as images, videos, text, audio, and / or other content, in response to input prompts and / or based on other information.

[0100] Example machine learning models include neural networks or other multi-layered nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some machine learning models may include multi-head self-attention models (e.g., transformer models).

[0101] The model can be trained using various training or learning techniques. Training can be supervised, unsupervised, reinforcement learning, etc. Training can utilize techniques such as backpropagation of errors. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent can be used to iteratively update parameters across multiple training iterations. Various generalization techniques (e.g., weight decay, dropout) can be used to improve the generalization ability of the model being trained.

[0102] The model can be pre-trained before domain-specific alignment. For example, the model can be pre-trained on a general training data corpus and then fine-tuned on a more targeted training data corpus. The model can be aligned using cues designed to elicit domain-specific outputs. The cues can be designed to include learned cue values ​​(e.g., soft cues). The trained model can be validated with input data other than the training data before its use and can be further updated or refined during its use based on additional feedback / input.

[0103] In some implementations, computing system 102 may use one or more of the machine learning models or techniques described above to combine machine learning to perform any or more of the operations discussed herein. For example, any model in model 136 may use one or more such machine learning techniques to help formulate a response (and / or intermediate inference, etc.).

[0104] The following list of examples reflects the various implementations explicitly envisioned by this disclosure:

[0105] Example 1. A computer-implemented method for accurately and with low latency processing a user request, the computer-implemented method comprising: receiving a user request from a client device via a user interface; generating a context-specific response to the user request by an orchestrator communicatively coupled to a plurality of dedicated artificial intelligence (AI) agents at least in part by: generating a response development plan specifying one or more of the plurality of dedicated AI agents based on (i) the user request and (ii) one or more context signals, wherein generating the response development plan includes determining that the one or more dedicated AI agents are relevant to the user request, and wherein each of the plurality of dedicated AI agents is configured to analyze a query specific to a corresponding domain; querying the one or more dedicated AI agents according to the response development plan; receiving corresponding feedback from the one or more dedicated AI agents in response to the query; generating the context-specific response based on the corresponding feedback from the one or more dedicated AI agents; and causing the context-specific response to be presented at the client device via the user interface.

[0106] Example 2. A computer-implemented method as described in Example 1, wherein: the one or more dedicated AI agents include a first dedicated AI agent and a second dedicated AI agent; the computer-implemented method further includes maintaining (i) a first agent context based on past communications between the orchestrator and the first dedicated AI agent and (ii) a second agent context based on past communications between the orchestrator and the second dedicated AI agent, the orchestrator maintaining the first agent context in isolation from the second agent context; and querying the one or more dedicated AI agents includes (i) generating a first synthetic query for the first dedicated AI agent based in part on the first agent context, and (ii) generating a second synthetic query for the second dedicated AI agent based in part on the second agent context.

[0107] Example 3. A computer-implemented method as described in Example 1 or 2, wherein: querying the one or more dedicated AI agents includes sending a first synthetic query to a first dedicated AI agent; receiving the corresponding feedback includes receiving (i) intermediate inference supporting a response to the first synthetic query and (ii) basic data used by the first dedicated AI agent to generate the intermediate inference from the first dedicated AI agent; and generating the context-specific response includes generating the context-specific response based on the basic data and the intermediate inference.

[0108] Example 4. A computer-implemented method as described in Example 3, wherein: the first dedicated AI agent includes a first AI agent response model configured to generate a response to a synthetic query based on the intermediate inference; and generating the context-specific response includes generating the context-specific response without using any response generated by the first AI agent response model.

[0109] Example 5. A computer-implemented method as described in Example 3 or 4, wherein receiving the corresponding feedback includes receiving instructions from the first dedicated AI agent regarding how to read the underlying data.

[0110] Example 6. A computer-implemented method as described in any one of Examples 1 to 5, wherein generating the response development plan and generating the context-specific response are performed by a single generative AI model.

[0111] Example 7. A computer-implemented method as described in any one of Examples 1 to 5, wherein: the orchestrator includes a planner model for generating the response development plan and a response model for generating the context-specific response; the planner model includes a first AI agent; and the response model includes a second AI agent.

[0112] Example 8. A computer-implemented method as described in Example 7, wherein: the orchestrator includes a third AI agent that receives the user request, generates a prompt based on the user request, and applies the prompt to a first AI agent such that the first AI agent generates the response development plan.

[0113] Example 9. A computer-implemented method as described in any one of Examples 1 to 8, wherein the one or more context signals include one or more of the following: one or more characteristics of an entity; information associated with the business of the entity; or information associated with the account of the entity.

[0114] Example 10. A computer-implemented method as described in any one of Examples 1 to 9, wherein the one or more context signals include data indicating the state of the user interface.

[0115] Example 11. A computer-implemented method as described in any one of Examples 1 to 10, wherein: generating the context-specific response includes identifying an action to be taken in response to the user request; and the computer-implemented method further includes causing the action to occur.

[0116] Example 12. A computer-implemented method as described in Example 11, wherein: identifying the action includes determining a set of one or more desired operations via the user interface; and causing the action to occur includes causing the client device to perform the set of one or more desired operations via the user interface by remotely assuming control over the input functions of the client device.

[0117] Example 13. A computer-implemented method as described in any one of Examples 1 to 10, wherein: the one or more dedicated AI agents include a first dedicated AI agent and a second dedicated AI agent; and generating the context-specific response includes resolving a conflict between (i) a first feedback provided by the first dedicated AI agent in response to the query and (ii) a second feedback provided by the second dedicated AI agent in response to the query.

[0118] Example 14. A computer-implemented method as described in any one of Examples 1 to 13, wherein the response development plan further specifies the order in which the planner model queries the one or more dedicated AI agents.

[0119] Example 15. A computer-implemented method as described in Example 14, wherein the sequence instructs the planner model to query at least two dedicated AI agents in parallel.

[0120] Example 16. A computer-implemented method as described in Example 15, wherein the sequence instructs the planner model to use the corresponding feedback from the first dedicated AI agent to generate a synthetic query for querying the second dedicated AI agent.

[0121] Example 17. A computer-implemented method as described in any one of Examples 1 to 16, further comprising: maintaining historical data by the planner model based on past interactions associated with (i) a user of the client device or (ii) an entity associated with the user or the client device, wherein (i) generating the response development plan and (ii) querying one or both of the one or more dedicated AI agents are based in part on the historical data.

[0122] Example 18. A computer-implemented method as described in Example 17, wherein the historical data includes chat history that is at least partially displayed on the user interface.

[0123] Example 19. A computer-implemented method as described in any one of Examples 1 to 18, further comprising: generating the corresponding feedback by the one or more dedicated AI agents.

[0124] Example 20. A computer-implemented method as described in Example 19, wherein generating the corresponding feedback includes: querying another AI agent by a first dedicated AI agent among the one or more dedicated AI agents.

[0125] Example 21. A computer-implemented method as described in Example 19, wherein generating the corresponding feedback includes: a tool configured to provide one or more types of information being invoked by a first dedicated AI agent of the one or more dedicated AI agents via an application programming interface (API).

[0126] Example 22. A computer-implemented method as described in any one of Examples 1 to 21, wherein the one or more dedicated AI agents include one or both of (i) a web crawling agent and (ii) a search engine agent.

[0127] Example 23. A computer-implemented method as described in any one of Examples 1 to 22, wherein the one or more dedicated AI agents include one or more of the following: a reporting agent; an optimization agent; a billing agent; a policy agent; or a diagnostic agent.

[0128] Example 24. A system comprising: one or more processors; and one or more memories storing instructions which, when executed by said one or more processors, cause said one or more processors to perform a computer-implemented method as described in any one of Examples 1 to 23.

[0129] Example 24. A non-transitory computer-readable medium storing one or more instructions that, when executed by one or more processors, cause the one or more processors to perform a computer-implemented method as described in any one of Examples 1 to 23.

[0130] Although the foregoing text describes many different aspects and implementations of the invention in detail, it should be understood that the scope of this patent is defined by the text of the claims set forth at the end of this patent. The detailed description should be understood as exemplary only and does not describe every possible implementation, as describing every possible implementation would be impractical or even impossible. Many alternative implementations may be implemented using prior art or technology developed after the date of this patent application, which will still fall within the scope of the claims.

[0131] The following additional considerations apply to the foregoing discussion and the appended claims. Throughout the specification, multiple instances may implement components, operations, or structures described as single instances. Although individual operations of one or more methods are shown and described as separate operations, one or more of the separate operations may be performed simultaneously, and the operations are not required to be performed in the order shown. Structures and functions presented as separate components in the example configuration may be implemented as combined structures or components. Similarly, structures and functions presented as single components may be implemented as single components. These, and other variations, modifications, additions, and improvements, all fall within the scope of the subject matter of this disclosure.

[0132] Unless otherwise apparent from the context of use, references in this disclosure to the same group of “one or more processors” (or the same “multiple processors”, etc.) performing multiple operations can cover implementations in which the execution of operations is divided among processors in any suitable manner. For example, “X is generated by one or more processors; and Y is generated by one or more processors” can cover: (1) an implementation in which a first group of one or more processors (e.g., in a first computing device) generates X and a different second group of one or more processors (e.g., in different second computing devices) independently generates Y; (2) an implementation in which all processors in the group of one or more processors (e.g., all in the same device or distributed among multiple devices) contribute to the generation of both X and Y; and (3) other variations.

[0133] Unless otherwise expressly stated, discussions using words such as “processing,” “calculating,” “deriving,” “determining,” “presenting,” “displaying,” etc., in this disclosure may refer to the actions or processing of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or combinations thereof), registers, or other machine components that receive, store, transmit, or display information.

[0134] As used in this disclosure, any reference to "one implementation" or "an implementation" means that a particular element, feature, structure, or characteristic described in connection with that implementation is included in at least one implementation or implementation. The appearance of the phrase "in one implementation" in various places in the specification does not necessarily refer to the same implementation.

[0135] As used in this disclosure, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, article of manufacture, or apparatus that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article of manufacture, or apparatus. Furthermore, unless expressly stated otherwise to the contrary, “or” means inclusive or rather than exclusive or. For example, condition A or B is satisfied by either: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).

[0136] Upon reading this disclosure, those skilled in the art will understand the additional alternative structures and functional designs implemented through the principles described herein. Therefore, while specific implementations and applications have been shown and described, it should be understood that the disclosed implementations are not limited to the precise constructions and components disclosed herein. Various modifications, alterations, and variations can be made to the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims, as will be apparent to those skilled in the art.

Claims

1. A computer-implemented method for accurately and with low latency processing user requests, the computer-implemented method comprising: Receive user requests from the client device via the user interface; An orchestrator, communicatively coupled to multiple dedicated artificial intelligence (AI) agents, generates a context-specific response to the user request at least in part by: Based on (i) the user request and (ii) one or more context signals, generate a response development plan specifying one or more of the plurality of dedicated AI agents, wherein generating the response development plan includes determining that the one or more dedicated AI agents are relevant to the user request, and wherein each of the plurality of dedicated AI agents is configured to analyze a query specific to a corresponding domain; Query the one or more dedicated AI agents according to the response development plan; In response to the query, receive corresponding feedback from the one or more dedicated AI agents; and Generate the context-specific response based on the corresponding feedback from the one or more dedicated AI agents; and This causes the context-specific response to be presented on the client device via the user interface.

2. The computer-implemented method as described in claim 1, wherein: The one or more dedicated AI agents include a first dedicated AI agent and a second dedicated AI agent; The computer-implemented method further includes the orchestrator maintaining (i) a first agent context based on past communications between the orchestrator and the first dedicated AI agent and (ii) a second agent context based on past communications between the orchestrator and the second dedicated AI agent, wherein the orchestrator maintains the first agent context in isolation from the second agent context; and Querying the one or more dedicated AI agents includes (i) generating a first synthetic query for the first dedicated AI agent in part based on the context of the first agent, and (ii) generating a second synthetic query for the second dedicated AI agent in part based on the context of the second agent.

3. The computer-implemented method as described in claim 1, wherein: Querying the one or more dedicated AI agents includes sending a first synthetic query to a first dedicated AI agent; Receiving the corresponding feedback includes receiving (i) intermediate inference supporting the response to the first synthetic query and (ii) the basic data used by the first dedicated AI agent to generate the intermediate inference; and Generating the context-specific response includes generating the context-specific response based on the base data and the intermediate inference.

4. The computer-implemented method as described in claim 3, wherein: The first dedicated AI agent includes a first AI agent response model, which is configured to generate a response to the synthetic query based on the intermediate inference; and Generating the context-specific response includes generating the context-specific response without using any response generated by the first AI agent response model.

5. The computer-implemented method as described in claim 3, wherein: Receiving the corresponding feedback includes receiving instructions from the first dedicated AI agent on how to read the basic data.

6. The computer-implemented method of claim 1, wherein, The generation of the response development plan and the generation of the context-specific response are performed by a single generative AI model.

7. The computer-implemented method of claim 1, wherein, The one or more context signals include one or more of the following: One or more properties of an entity; Information related to the entity's business; or Information associated with the entity’s account.

8. The computer-implemented method of claim 1, wherein, The one or more context signals include data indicating the state of the user interface.

9. The computer-implemented method as described in claim 1, wherein: Generating the context-specific response includes identifying the action to be taken in response to the user request; and The computer-implemented method further includes causing the action to occur.

10. The computer-implemented method as described in claim 9, wherein: Identifying the action includes determining one or more desired actions via the user interface; and Causing the action to occur includes causing the client device to perform one or more desired operations via the user interface by remotely assuming control over the input functions of the client device.

11. The computer-implemented method as described in claim 1, wherein: The one or more dedicated AI agents include a first dedicated AI agent and a second dedicated AI agent; and Generating the context-specific response includes resolving the conflict between (i) a first response provided by the first dedicated AI agent in response to the query and (ii) a second response provided by the second dedicated AI agent in response to the query.

12. The computer-implemented method of claim 1, wherein, The response development plan further specifies the order in which the orchestrator queries the one or more dedicated AI agents.

13. The computer-implemented method of claim 12, wherein, The sequence instructs the orchestrator to use the corresponding feedback from the first dedicated AI agent to generate a synthetic query for querying the second dedicated AI agent.

14. The computer-implemented method of claim 1, further comprising: The orchestrator maintains historical data based on past interactions associated with (i) the user of the client device or (ii) entities associated with the user or the client device. Wherein, (i) generating the response development plan and (ii) querying one or both of the one or more dedicated AI agents are based in part on the historical data.

15. The computer-implemented method of claim 14, wherein, The historical data includes, at least in part, the chat history displayed on the user interface.

16. The computer-implemented method of claim 1, wherein, The one or more dedicated AI agents include one or both of (i) web crawling agents and (ii) search engine agents.

17. The computer-implemented method of claim 1, wherein, The one or more dedicated AI agents include one or more of the following: Reporting agent; Optimize the intelligent agent; Billing agent; Policy agent; or Diagnostic intelligent agents.

18. A system comprising: One or more processors; as well as One or more memories storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: Receive user requests from the client device via the user interface; An orchestrator, communicatively coupled to multiple dedicated artificial intelligence (AI) agents, generates a context-specific response to the user request at least in part by: Based on (i) the user request and (ii) one or more context signals, generate a response development plan specifying one or more of the plurality of dedicated AI agents, wherein generating the response development plan includes determining that the one or more dedicated AI agents are relevant to the user request, and wherein each of the plurality of dedicated AI agents is configured to analyze a query specific to a corresponding domain; Query the one or more dedicated AI agents according to the response development plan; In response to the query, receive corresponding feedback from the one or more dedicated AI agents; and Generate the context-specific response based on the corresponding feedback from the one or more dedicated AI agents; and This causes the context-specific response to be presented on the client device via the user interface.

19. The system of claim 18, wherein: The one or more dedicated AI agents include a first dedicated AI agent and a second dedicated AI agent; The operation further includes the orchestrator maintaining (i) a first agent context based on past communications between the orchestrator and the first dedicated AI agent and (ii) a second agent context based on past communications between the orchestrator and the second dedicated AI agent, wherein the orchestrator maintains the first agent context in isolation from the second agent context; Furthermore, querying the one or more dedicated AI agents includes (i) generating a first synthetic query for the first dedicated AI agent based in part on the context of the first agent, and (ii) generating a second synthetic query for the second dedicated AI agent based in part on the context of the second agent.

20. The system of claim 18, wherein: Querying the one or more dedicated AI agents includes sending a first synthetic query to a first dedicated AI agent; receiving the corresponding feedback includes receiving (i) intermediate inference supporting the response to the first synthetic query and (ii) basic data used by the first dedicated AI agent to generate the intermediate inference from the first dedicated AI agent; and Generating the context-specific response includes generating the context-specific response based on the base data and the intermediate inference.