Storage for automating context grounding and other results using agent-based memory.

Agent-based extraction and search methods enhance traditional software automation by processing complex document structures and industry terminology, improving AI model responses through context grounding and memory storage.

JP2026122457APending Publication Date: 2026-07-28UIPATH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
UIPATH INC
Filing Date
2025-12-15
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Traditional software automation struggles with processing complex document structures and industry jargon, leading to inefficient and inaccurate responses due to the inability to decipher relevance, especially when using Large-Scale Language Models (LLMs) for repetitive human-computer tasks.

Method used

Implementing agent-based extraction and search methods using AI agents to perform multi-stage processing of text and images, generate context grounding, and store context embeddings in agent-based memory to enhance the accuracy of AI model responses.

Benefits of technology

Improves the efficiency and accuracy of AI model responses by providing context grounding and storing relevant information, reducing the number of actions required and enhancing processor efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026122457000007
    Figure 2026122457000007
  • Figure 2026122457000008
    Figure 2026122457000008
  • Figure 2026122457000009
    Figure 2026122457000009
Patent Text Reader

Abstract

This invention provides an agent-type memory and method for storing context grounding and other results for automation purposes. [Solution] The method is implemented by an artificial intelligence (AI) agent. The agent-type memory method includes storing a vector containing context embedding generated by advanced agent-type extraction in response to a query in agent-type memory; receiving and matching a complex query to the vector in agent-type memory by advanced agent-type search in order to determine context grounding; providing the context grounding along with the complex query to the AI ​​model in order to improve the accuracy of the AI ​​model's response; and receiving a response from the AI ​​model to the complex query.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates in general to automation, and more specifically to agent-type memory for storing context grounding and other results for automation. [Background technology]

[0002] Generally, traditional software automation performs basic, repetitive human-computer tasks. When performing basic, repetitive human-computer tasks, traditional software automation statically executes hundreds (100 or more) operations, consuming large amounts of processing and memory resources.

[0003] For example, traditional software automation may require parsing data from multiple databases containing specific industry jargon and complex document structures to retrieve relevant information. In this respect, traditional software automation requires processing power to gain control of database access and analyze data to determine relevance. In most cases, industry jargon and complex document structures present challenges to data access and analysis because traditional software automation cannot decipher the relevance from them. Furthermore, the retrieved data is passed to a model that performs final processing to provide a response to a human-computer task. However, because this model does not inherently recognize the relevance, it may take an extremely long time to generate a response, or it may even be completely off the mark.

[0004] Furthermore, traditional software automation utilizes Large-Scale Language Models (LLMs) when performing basic, repetitive human-computer tasks. The problem with LLMs is that while they appear to work well for general purposes (when using well-known keywords), they become significantly inaccurate when basic, repetitive human-computer tasks transform into complex business processes with specific goals. In other words, the real world is not general-purpose; business applications reflect this concept, and complex business processes are unique or have certain unique characteristics. Returning to the aforementioned industry jargon and complex document structures, different terminology and structures present challenges to LLMs and their functionality.

[0005] What is needed is a mechanism or method to improve traditional software automation by linking automation to relevance. In other words, improved and / or alternative approaches for sophisticated chunking of complex document structures and advanced extraction and retrieval techniques for industry terminology may be beneficial in ensuring that relevant information is passed to the model without noise and that responses are tailored to diverse industries and applications. Furthermore, improved and / or alternative storage approaches for semantically mapping what queries, actions, or tasks actually mean may also be beneficial in facilitating correlation within such advanced extraction and retrieval techniques. [Overview of the Initiative]

[0006] Certain embodiments described herein may offer alternatives or solutions to, or useful alternatives to, challenges and needs in the technical field that have not yet been fully identified, evaluated, or addressed by current conventional software automation technologies. For example, one or more embodiments relate to agent-type memory for storing context grounding and other results for automation.

[0007] According to one or more embodiments, an agent-based extraction method is provided. The agent-based extraction method is implemented by one or more artificial intelligence (AI) agents and generates and provides context grounding for an AI model. The agent-based extraction method includes extracting text data from one or more documents containing complex document structures and obtaining one or more images from the one or more documents. The agent-based extraction method includes performing multi-stage processing of the text data and the one or more images using two or more large-scale language models to generate an output set containing contextualization and one or more keywords. The agent-based extraction method includes converting the output set into one or more vectors containing contextual embeddings to provide the context grounding.

[0008] According to one or more embodiments, an agent-based search method is provided. The agent-based search method is implemented by one or more artificial intelligence (AI) agents. The agent-based search method includes performing an advanced agent-based search of semantic storage in response to a complex query. The agent-based search method includes outputting the context grounding along with the complex query to the AI ​​model in order to improve the accuracy of the AI ​​model's response. The agent-based search method includes receiving the response to the complex query from the AI ​​model.

[0009] According to one or more embodiments, an agent-based memory method is provided. The agent-based memory method is implemented by one or more artificial intelligence (AI) agents. The agent-based memory method includes storing one or more vectors, including context embeddings generated by advanced agent-based extraction, in agent-based memory in response to one or more queries. The agent-based memory method includes determining context grounding by matching the complex queries to the one or more vectors in the agent-based memory using advanced agent-based search. The agent-based memory method includes providing the context grounding along with the complex queries to the AI ​​model to improve the accuracy of the AI ​​model's response, and receiving a response from the AI ​​model to the complex queries.

[0010] Any of the methods described herein may be implemented as a computer program product, system, apparatus, and / or device. [Brief explanation of the drawing]

[0011] To facilitate understanding of the advantages of specific embodiments, a more detailed description of the above overview will be given with reference to specific embodiments shown in the accompanying drawings. It should be understood that these drawings illustrate only representative embodiments and do not limit the scope of this disclosure. This disclosure will be described in more specific and detail with reference to the accompanying drawings.

[0012] [Figure 1] This is an architectural diagram showing a hyperautomation system configured to perform agent-based automation and orchestration in one or more embodiments.

[0013] [Figure 2] This illustrates some of the capabilities for combining artificial intelligence (AI) agents and robotic process automation (RPA) robots in one or more embodiments.

[0014] [Figure 3] Shows an AI agent, RPA robot, agent-based orchestration process (AOP), and a pool of applications according to one or more embodiments.

[0015] [Figure 4] Shows an example of an AI agent service interface according to one or more embodiments.

[0016] [Figure 5] Shows an example of an AOP development interface according to one or more embodiments.

[0017] [Figure 6] Shows an example of an RPA development interface according to one or more embodiments.

[0018] [Figure 7] Shows an end-to-end AI agent, RPA robot, and AOP development and deployment system according to one or more embodiments.

[0019] [Figure 8] An architecture diagram showing an agent-based automation and RPA system according to one or more embodiments.

[0020] [Figure 9] An architecture diagram showing a deployed RPA system according to one or more embodiments.

[0021] [Figure 10] An architecture diagram showing the relationship between a designer, an activity, and a driver according to one or more embodiments.

[0022] [Figure 11]This is an architectural diagram showing a computing system configured to perform agent-based automation and orchestration in one or more embodiments.

[0023] [Figure 12A] Examples of neural networks trained to perform agent-based automation and orchestration in one or more embodiments are shown.

[0024] [Figure 12B] Examples of neurons according to one or more embodiments are shown.

[0025] [Figure 13] This is an architecture diagram showing a reference architecture for a generative AI model according to one or more embodiments.

[0026] [Figure 14] This flowchart shows the learning process of an AI / ML model according to one or more embodiments.

[0027] [Figure 15] The process is shown according to one or more embodiments.

[0028] [Figure 16] The process is shown according to one or more embodiments.

[0029] [Figure 17] This is an architectural diagram showing a hyperautomation system configured to perform agent-based automation and orchestration in one or more embodiments.

[0030] [Figure 18] The process is shown according to one or more embodiments.

[0031] Unless otherwise stated, the same reference numerals consistently indicate the corresponding features throughout the attached drawings. [Modes for carrying out the invention]

[0032] (Detailed description of the embodiment) One or more embodiments described herein relate to agent-type memory. More specifically, agent-type memory may be used for automation to store context grounding and to cache other results.

[0033] According to one or more embodiments, enhanced extractive and retrieval techniques tailored to diverse industries and applications (e.g., tailored to specific industry terminology and complex document structures) are provided, thereby improving model responses. More specifically, advanced agent-based extractive and retrieval techniques provide contextual grounding within automation to improve models such as Large-Scale Language Models (LLMs) by integrating company-specific information with pre-trained knowledge, enabling accurate responses to specialized or modern queries.

[0034] For example, in context grounding, a query is sent via a prompt to be executed on any LLM in the cloud (e.g., one owned by the user or a third party). However, the LLM has no context because it can only access data prior to a certain date (i.e., the LLM is missing the current calendar year, and the query is sent in August). How does the LLM then obtain the context necessary to respond to the query? How is additional context provided to the LLM? The advanced agent-based extraction and search embodiments described herein infuse the necessary context into the LLM by performing a search alongside the query and incorporating the search results into the prompt as if they were sent with the query. Thus, the query, the LLM, and the response are grounded based on context that would otherwise not have been available in the data.

[0035] According to one or more embodiments, agent-based memory may be provided as long-term memory for automation. In this regard, agent-based memory may be used to store context grounding generated by extended extract and retrieve techniques, as well as to cache other results generated by automation across the entire system. In this way, agent-based memory can overcome the shortcomings of LLM by providing an improved alternative storage approach that semantically maps context grounding and other results for use by system-wide automation.

[0036] Figure 1 is an architectural diagram showing a hyperautomation system 100 configured to perform agent-based automation and orchestration in one or more embodiments. In this specification, “hyperautomation” refers to an automation system that integrates process automation, agent-based automation, integration tools, and technological components to amplify the automation capabilities of work. Examples of components include, but are not limited to, artificial intelligence (AI) agents, agent-based orchestration processes (AOPs), and robotic process automation (RPA) robots.

[0037] Generally, as used herein, “AI agent” refers to an AI-enhanced probabilistic automation that operates independently, dynamically, makes decisions, performs actions, and acts adaptively. In some cases, an AI agent operates in this manner by utilizing a Large-Scale Language Model (LLM) or other AI model, which itself is typically probabilistic. According to one or more embodiments and as described herein, one or more AI agents may implement context-grounding techniques (e.g., Search Augmentation Generation (RAG), extraction, and semantic storage), advanced agent-based search (semantic search and search, hybrid search, and wide-area search with re-ranking), tethering (including automatic tethering), and query decomposition.

[0038] Generally, as used herein, "AOP" refers to automation that combines probabilistic and deterministic methods in a way that is both dynamic and predictable. In some cases, AOP is automation that allows a user to describe the overall business process. A user includes, but is not limited to, any person who has access to the system performing the automation (e.g., developers, engineers, customers, etc.). AOP may be created using an interface that allows the creation of a business flowchart of the business process, which may be described in Business Process Model and Notation (BPMN). BPMN is an Extended Markup Language (XML) description of a business process (see, for example, Figure 5).

[0039] Generally, as used herein, "RPA robot" refers to rule-based deterministic automation that operates predictably and makes deterministic decisions.

[0040] For example, in some embodiments, one or more RPAs may be used at the core of a hyperautomation system, and in certain embodiments, automation capabilities may be extended by AI / machine learning (ML), process mining, analytics, agent-based automation, and / or other advanced tools. As the hyperautomation system learns processes, learns AI / ML models, and uses analytics, more knowledge work can be automated, and computing systems within the organization, both those used by individuals and those operating autonomously, can be involved as participants in the hyperautomation process. In some embodiments, the hyperautomation system enables users and organizations to discover, understand, and scale automation efficiently and effectively.

[0041] In such embodiments, the AI ​​agent "coexists" with the RPA robots and AOPs that perform the RPA. As mentioned above, the AI ​​agent is AI-skilled automation that can operate independently, make dynamic decisions, perform actions, and adapt its performance. AI agents can dynamically utilize the tools provided by these RPA robots to perform tasks such as document processing (see, for example, U.S. Patent Application Publication No. 2021 / 0097274), user interface (UI) automation (see, for example, U.S. Patents No. 10,654,166, 10,990,876, 11,080,548, 11,507,259, 11,733,668, and 11,748,069), and semantic copy and paste between source and target (see, for example, U.S. Patent No. 12,124,806 and U.S. Patent Application Publications No. 2023 / 0107316, 2023 / 0415338, and 2024 / 0220581). AI agents can dynamically select these tools and execute them in a pipeline.

[0042] Generally, agent-based automation is probabilistic automation performed by one or more AI agents. Agent-based automation expands an organization's automation potential by focusing not only on individual tasks but also on the entire end-to-end process. Teams of RPA robots and / or AOP led by AI agents can enable one employee to accomplish the work of many people. Agent-based automation with AI agents and / or AOP gives managers room for mentoring, physicians more time for patient care, developers the ability to fine-tune their work, engineers the freedom to innovate, and customers a seamless, personalized experience.

[0043] Agent-based automation can achieve a variety of technical effects, benefits, and advantages. It improves memory usage by reducing the amount of data stored and enhances processor efficiency by reducing the number of calls and actions. Agent-based automation can also provide the ability to process gigabytes, terabytes, petabytes, or more of data that would be impossible to implement mentally or manually. Furthermore, agent-based automation can use fewer triggers and models through dynamic decision-making. While traditional software automation alone might require 100 actions in an example scenario, agent-based automation in the same example scenario can significantly reduce the required actions (e.g., to 15 actions). Agent-based automation can also improve the efficiency of LLMs or AI models by using context grounding to tether AI agents to the desired context, "constraining" LLMs or AI models to their relevant contexts.

[0044] An AI agent may have agent-type memory that evolutionarily stores user interactions, feedback, corrections, and solutions (e.g., dynamic and / or direct user input in human-in-the-loop operations). As used herein, “human-in-the-loop operations” or “human-in-the-loop” may include AI agents and RPA robots working in conjunction with users to receive dynamic and / or direct user input. As agent-type memory grows, the AI ​​agent may become more autonomous, reducing the need for dynamic and / or direct user input and improving efficiency. The AI ​​agent may also learn to become more efficient based on agent-type memory if more efficient solutions are contained within or derived from it. For example, the AI ​​agent may periodically process agent-type memory to analyze patterns in order to achieve greater autonomy.

[0045] As used herein, “agent-based memory” refers to a dynamic caching (i.e., storage) system for managing escalations and tool calls. For example, if an AI agent encounters a problem during execution, it may prompt for user interaction or feedback to overcome the problem, store / cached that interaction or feedback, and learn from it to reduce the need for repeated user input. Through one or more technical effects, benefits, and advantages, agent-based memory improves efficiency by storing solutions to common problems and minimizing potentially costly tool calls. The collaborative operation between the AI ​​agent and agent-based memory can “bend the curve” so that as the AI ​​agent continuously learns through agent-based memory, the frequency of user involvement decreases.

[0046] Generally, agent-based orchestration is implemented by a conductor application that implements one or more AOP and / or AI agents to orchestrate AI agents (e.g., UiPath Agents®), third-party agents, RPA robots (e.g., UiPath Robots®), AOP, and users (where user involvement is required or requested) that execute workflows. Agent-based orchestration enables the modeling and monitoring of complex business processes as agent-based automation from start to finish. Agent-based orchestration also offers a unique ability to orchestrate RPA robots, AI agents, AOP, third-party agents, and users across an end-to-end workflow. Agent-based orchestration is beneficial for the successful scaling of agent-based automation.

[0047] AI agents for agent-based automation are based on AI models as described herein, allowing them to operate independently of the user and implement these agent-based automations. AI agents are also goal-oriented and make probabilistic decisions using context. Furthermore, AI agents are well-suited for ad-hoc tasks requiring high adaptability. AI agents learn how tasks are performed and improve over time. AI agents can use and select various tools for task accomplishment, context gathering, and action execution (often through RPA robots used as tools by the AI ​​agent), and in some embodiments, AI agents may generate automations by building workflows to be executed by RPA robots and / or other AI agents, leveraging UiPath Autopilot® for developers or other applications that accelerate the creation and testing of agent-based automations. For example, an AI agent may utilize a designer application via an API to generate another AI agent or RPA robot to execute parts of a workflow, and may also trigger human intervention to escalate issues with the workflow. Where correct, the workflow can be deployed. AI agents may also exhibit varying degrees of autonomy controlled by agent-based orchestration.

[0048] The AI ​​agent, by executing an "agent loop," uses given tools and context to generate a dynamic plan to follow instructions and achieve its goals. Once the dynamic plan is generated, the AI ​​agent utilizes an efficient execution path for that plan. If the dynamic plan has two or more steps that can be executed in parallel, the AI ​​agent executes those two or more steps in parallel based on available resources. After each step is completed, the AI ​​agent retrieves the output of that step and regenerates the next step or multiple steps. Thus, the agent loop continues until the goal is achieved. Executing steps in a dynamic plan in parallel and utilizing ecosystem tools and context grounding are advanced capabilities of agent-based orchestration.

[0049] As mentioned above, RPA robots are rule-based automations that operate predictably and make deterministic decisions. RPA robots are highly reliable and efficient and are suitable for routine tasks. RPA robots, along with AI agents, can use human intervention for exception handling. According to one or more embodiments, AI agents are more flexible, abstract, and self-willed than RPA robots and AOP. Furthermore, RPA robots are more stable, concrete, and controllable than AI agents and AOP. Moreover, AOP processes lie between the flexibility / stability, abstraction / concreteness, and self-will / controllability of AI agents and RPA robots, respectively.

[0050] As will be discussed later with respect to Figure 3, AI agents, AOPs, and RPA robots can find and use each other as tools to accomplish tasks. AI agents, AOPs, and RPA robots can also access and utilize various applications (e.g., via Application Programming Interfaces (APIs)). Tools can be manually configured by developers for automation, or AI agents and RPA robots can discover and use tools at runtime.

[0051] In some embodiments, AI agents, AOPs, and RPA robots can work in collaboration with users (e.g., human participants), enabling them to make faster, more consistent, and more informed decisions. Furthermore, the use of AI agents, AOPs, and RPA robots allows users to accomplish more. Specifically, AI agents, AOPs, and RPA robots can take on repetitive, monotonous, and ad-hoc additional tasks at a scale that humans cannot handle. Users can make decisions when the AI ​​agent, AOP, or RPA robot encounters an exception. Thus, users are elevated to and can focus on roles as supervisors, decision-makers, and organizational leaders.

[0052] AI models provide AI agents with the ability to reason, plan, create, and make autonomous decisions. AI models can also be used by RPA robots for task-specific activities such as document processing and data analysis. AI models can be enhanced with business-specific content and context (e.g., from a company's context repositories) to improve their accuracy and results. AI models can be applied individually or concurrently, depending on the complexity of the task. AI model selection can be from the RPA vendor's model library, third-party models, and Bring Your Own Model (BYOM) options (see, for example, U.S. Patents 11,738,453 and 11,748,479).

[0053] The hyper-automation system 100 includes user computing systems such as a desktop computer 102, a tablet 104, and a smartphone 106. However, any user computing system, including smartwatches, laptop computers, servers, and Internet of Things (IoT) devices, may be used without departing from the scope of this disclosure. Also, although three user computing systems are shown in Figure 1, any number of user computing systems may be used without departing from the scope of this disclosure. For example, in some embodiments, tens, hundreds, thousands, or millions of user computing systems may be used. User computing systems may be used actively by a user, or they may be executed automatically by AI agents, AOPs, and / or RPA robots with little or no user input.

[0054] As disclosed herein, some embodiments include three types of automation: (1) agent-based automation implemented by each AI agent, (2) RPA implemented by each RPA robot, and (3) composite automation achieved by a combination of AI agents and RPA robots to accomplish a more complex overall task. Automations 110, 112, and 114 may include, but are not limited to, those performed by RPA robots and / or AI agents (either individually or to achieve a larger composite automation). Other processes, such as listeners, may also be implemented. These processes may be implemented as standalone applications, subprocesses of other applications, parts of operating systems, any other suitable software and / or hardware, or any combination thereof. Without departing from the scope of this disclosure, the logic of a process may be partially or completely implemented by physical hardware.

[0055] Each user computing system 102, 104, and 106 executes automations 110, 112, and 114, respectively, which are implemented by RPA robots, AI agents, AOP, etc. In some embodiments, automations 110, 112, and 114 may be stored remotely (e.g., on server 130, or stored in database 140 and accessed via network 120) and loaded by RPA robots and / or AI agents to implement automations 110, 112, and 114. Database 140 may store structured data and / or unstructured data, although the former is typically required for RPA. RPA automations may exist as scripts (e.g., Extended Markup Language (XML), Extended Application Markup Language (XAML), etc.) or be compiled as machine-readable code (e.g., dynamic link libraries). In the case of AI agents, agent-type automations may be generated based on a plaintext description of the desired target.

[0056] The listener monitors and records data regarding user operations in each computing system and / or the operation of the unattended computing system, and transmits this data to the core hyperautomation system 120 via a network (e.g., a local area network (LAN), mobile communication network, satellite communication network, the Internet, or any combination thereof). The data may include, but is not limited to, which buttons were clicked, where the mouse moved, text entered into fields, when one window was minimized and another was opened, and the applications associated with the windows. In certain embodiments, data from the listener may be transmitted periodically as part of a heartbeat message. In some embodiments, data may be transmitted to the core hyperautomation system 120 when a predetermined amount of data has been collected, when a predetermined amount of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.

[0057] If automations 110, 112, and 114 are RPAs, they can execute the logic developed in the workflow at design time. A workflow may include a set of steps executed in a series or other logical flow, such steps are defined herein as “activities.” Each activity may include actions such as clicking a button, reading a file, or writing to a log panel. In some embodiments, workflows may be nested or embedded.

[0058] In some embodiments, long-running RPA workflows are master projects that support long-running transactions in service orchestration, human-participatory, and unattended environments. See, for example, U.S. Patent No. 10,860,905 (which is incorporated herein by reference in its entirety). Human participation involves situations where a particular process requires user input (dynamic and / or direct user input) to handle exceptions, approvals, or validations before proceeding to the next step. In this situation, process execution is interrupted and the RPA robot is released until the human-participatory portion of the task is completed.

[0059] Long-running workflows can support workflow fragmentation through persistent activities and, in combination with invocation processes and non-user interaction activities, can orchestrate human-participatory operations alongside RPA robot tasks. In some embodiments, multiple or numerous computing systems may participate in the logical execution of a long-running workflow. Long-running workflows may run within a session to facilitate rapid execution. In some embodiments, a long-running workflow can orchestrate background processes, including activities that perform API calls, etc., which are executed within the long-running workflow session. These activities may be invoked by invocation process activities in some embodiments. Processes with user interaction activities executed within a user session may be invoked by starting a job from a conductor activity (conductors are described later). In some embodiments, the user may interact with the conductor through tasks that require filling out forms. Activities may include causing the RPA robot to wait for the form task to complete and then resuming the long-running workflow.

[0060] One or more automations 110, 112, 114 can communicate with the core hyperautomation system 120. In some embodiments, the core hyperautomation system 120 may run conductor applications on one or more servers, such as server 130. Although one server 130 is shown for illustrative purposes, multiple servers in close proximity to each other, or a number of servers in a distributed architecture, may be used without departing from the scope of the invention. For example, one or more servers may be provided for conductor functions, AI / ML model distribution, authentication, governance, or any other appropriate functions. In some embodiments, the core hyperautomation system 120 may incorporate, or be part of, a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyperautomation system 120 may host multiple software-based servers on one or more computing systems, such as server 130. In some embodiments, one or more servers of the core hyperautomation system 120 (e.g., server 130) may be implemented by one or more virtual machines (VMs).

[0061] In some embodiments, one or more of the automations 110, 112, and 114 may invoke one or more AI / ML models 132 that are deployed on or accessible by the core hyperautomation system 120 and are trained to accomplish various tasks. For example, the AI / ML models 132 may include, but are not limited to, models trained to explore various application versions, models that perform computer vision (CV), models that perform optical character recognition (OCR), models that generate user interface (UI) descriptors, models that provide suggestions for the next activity or activity sequence in a workflow, models that perform semantic matching, models that perform natural language processing (NLP), models that generate or modify code and / or workflows, etc. The AI / ML models may be trained using labeled data, which may include, but are not limited to, elements from data sources (e.g., web pages, forms, scanned documents, application interfaces, screens, etc.), previously created workflows, screenshots of various versions of various application screens with corresponding UI elements, libraries of UI objects, etc. The AI / ML model 132 can be trained to achieve a desired confidence threshold while avoiding overfitting to a given set of training data. In general, UI elements, UI descriptors, applications, and application screens can be considered UI objects.

[0062] AI / ML models 132 can be trained for any suitable purpose without departing from the scope of the present invention, which will be described in more detail later in this specification. In some embodiments, two or more AI / ML models 132 can be chained together (e.g., in series, parallel, or a combination thereof) so that they work together to provide output. AI / ML models 132 can perform or assist in CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated workflow generation, sequence extraction, clustering detection, speech-to-text conversion, NLP, semantic matching, and any combination thereof. However, any number and / or types of AI / ML models can be used without departing from the scope of the present invention. By using multiple AI / ML models, the system may be able to form a whole picture of what is happening on a given computing system, for example. For example, one AI / ML model may perform OCR, another model may detect buttons, and yet another model may compare sequences, and so on. The patterns may be determined individually by a single AI / ML model, or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one of the computing systems 102, 104, and 106.

[0063] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that operates on data, and for example, the AI / ML model itself may be a deep learning neural network (DLNN) consisting of artificial "neurons" learned based on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions such as statistical modeling, such as Hidden Markov Models (HMMs), and may perform desired functions using deep learning techniques (e.g., Long Short-Term Memory (LSTM) deep learning, encoding of previous hidden states, etc.).

[0064] The hyper-automation system 100 may provide four main functional groups in some embodiments: (1) discovery, (2) automation construction, (3) management, and (4) engagement. Automation (e.g., executed on a user computing system, server, etc.) may, in some embodiments, be performed by, for example, an RPA robot, AOP, or AI agent, and may provide any of the functions described herein. For example, an RPA robot may include a manned robot, an unmanned robot, and / or a test robot. A manned robot collaborates with a user to assist with a task (e.g., via UiPath Assistant®). An unmanned robot may operate independently of a user and run in the background without user awareness. A test robot executes test cases against an application or workflow. In some embodiments, a test robot may run in parallel on multiple computing systems.

[0065] Discovery capabilities can discover business process automation opportunities and provide automated recommendations for different opportunities. Such capabilities may be implemented by one or more servers, such as server 130. In some embodiments, discovery capabilities may include an automation hub, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath Automation Hub®) may provide a mechanism for automation deployment management with visibility and control. For example, automation ideas may be crowdsourced from employees via a submission form. Feasibility and return on investment (ROI) calculations for automating these ideas may be provided, documentation for future automations may be collected, and collaboration may be provided to accelerate the automation discovery to build process.

[0066] Process mining (e.g., via UiPath Automation Cloud® and / or UiPath AI Center®) refers to collecting and analyzing data from applications (e.g., enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify end-to-end processes within an organization and how to effectively automate them, as well as the impact of automation. This data may be acquired, for example, by listeners from user computing systems 102, 104, and 106 and processed by servers such as server 130. In some embodiments, one or more AI / ML models 132 may be used for this purpose. This information can be exported to an automation hub to accelerate implementation and avoid manual information transfer. The objective of process mining may be to increase business value by automating processes within an organization. Examples of objectives for process mining may include, but are not limited to, increased profits, improved customer satisfaction, regulatory and / or contractual compliance, and improved employee efficiency.

[0067] Task mining (e.g., via UiPath Automation Cloud® and / or UiPath AI Center®) identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in daily tasks, and scores those tasks for ease of automation and potential savings (e.g., time and / or cost savings). In some embodiments, one or more AI / ML models 132 may be used to find recurring task patterns in the data. Recurring tasks suitable for automation can be identified. In some embodiments, this information is initially provided by a listener and can be analyzed on a server (e.g., server 130) of the core hyperautomation system 120. The results of task mining (e.g., XAML process data) can be exported to process documentation or to a designer application such as UiPath Studio® to more quickly create and deploy automations. In some embodiments, task mining may include taking screenshots with user actions (e.g., mouse click locations, keyboard inputs, application windows and graphic elements the user was interacting with, timestamps of actions, etc.), collecting statistical data (e.g., execution time, number of actions, number of text inputs, etc.), editing and annotating screenshots, specifying the types of actions to be recorded, and so on.

[0068] Task capture (e.g., via UiPath Automation Cloud® and / or UiPath AI Center®) automatically documents human processes as users work or provides a framework for unmanned processes. Such documentation may include desired tasks to be automated as process definition documents (PDDs), skeletal workflows, behavioral captures of each part of the process, automated generation of comprehensive workflow diagrams including user interaction records and step-by-step details, Microsoft Word® documents, XAML files, etc. The built-to-go workflows can be exported directly to designer applications such as UiPath Studio® in some embodiments. Task capture can simplify the requirements gathering process for both subject matter experts describing the process and Center of Excellence (CoE) members providing production-quality automation.

[0069] The construction of automation can be achieved through designer applications (e.g., UiPath Studio®, UiPath StudioX®, or UiPath Studio Web®). For example, developers in the RPA development facility 150 can use the designer application 154 on the computing system 152 to build and test agent-based automation, RPA, AOP, and / or hybrid automation for various applications and environments such as web, mobile, SAP®, and virtual desktops. Developers can also build AOP. For example, developers can create automations that are executed by RPA robots, AI agents, AOP, or a combination thereof. API integrations can be provided for various applications, technologies, and platforms. Predefined activities, drag-and-drop modeling, and workflow recording capabilities can facilitate automation with minimal coding. Document understanding capabilities can be provided as drag-and-drop AI skills for data extraction and interpretation, which call one or more AI / ML models 132. Such automations can handle substantially any document type and format, including tables, checkboxes, signatures, and handwritten text. When data is validated or exceptions are handled, this information can be used to retrain the corresponding AI / ML model, improving its accuracy over time.

[0070] The designer application 152 may be designed to invoke one or more trained AI / ML models 132 on the server 130 and / or one or more generated AI models 172 in a cloud environment via the network 120 (e.g., LAN, mobile communication network, satellite communication network, internet, or any combination thereof) to support the automated development process. In some embodiments, one or more AI / ML models may be packaged in the designer application 152 or stored locally on the computing system 150.

[0071] In some embodiments, one or more of the designer application 152 and AI / ML models 132 may be configured to use an object repository stored in the database 140. See, for example, U.S. Patent No. 11,748,069 (this document is incorporated herein by reference in its entirety). Generally, an object repository is a storage mechanism used by automation for images, text, semantic data, taxonomic associations, ontological associations, UI objects, etc. For example, an object repository may include a library of UI objects that can be used to develop workflows through the designer application 152. The object repository may be used for UI automation to add UI descriptors to activities in the workflow of the designer application 152. In some embodiments, one or more of the AI / ML models 132 may generate new UI descriptors and add them to the object repository in the database 140.

[0072] Once automation is complete in the designer application 152, it can be exposed on the server 130 and distributed to computing systems 102, 104, 106, etc. For example, as new UI descriptors are created or existing UI descriptors are modified, a global repository of a UI object library that can be shared and coordinated across all automations can be built. Taxonomies and ontologities can be used for the object repository. A taxonomy is a hierarchical structure of subcategories. An ontology is a formal representation of a knowledge domain that includes concepts, attributes, and their relationships. In ontologities, the relationships between categories do not necessarily have to be hierarchical, and ontological relationships can span multiple screens of an application.

[0073] Integrated services can enable developers to seamlessly combine, for example, UI automation and API automation. Any type of automation described herein can be built to require APIs or to span both API and non-API applications and systems. A repository of pre-built automation templates and solutions (e.g., UiPath Object Repository®) or a marketplace (e.g., UiPath Marketplace®) may be provided to enable developers to automate a wide variety of processes more quickly. Thus, when building automation, the hyperautomation system 100 can provide a user interface, development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. In some embodiments, the hyperautomation system 100 can enable the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, AOPs, and AI agents, which can provide automation for the hyperautomation system 100.

[0074] In some embodiments, components of the hyperautomation system 100 (e.g., designer applications and / or external rule engines) assist in managing and enforcing governance policies that control the various functions provided by the hyperautomation system 100. Governance is the ability of an organization to set policies to prevent users from developing automations (e.g., RPA robots, AOPs, and / or AI agents) that could perform actions that could harm the organization (e.g., violations of the EU General Data Protection Regulation (GDPR), the US Health Insurance Portability and Accountability Act (HIPAA), third-party application terms of service, etc.). Because developers may create automations that violate privacy laws, terms of service, etc., in some embodiments, access control and governance restrictions are implemented at the robot and / or robot design application level. This prevents developers from relying on unauthorized software libraries (which may introduce security risks or operate in a way that violates policies / regulations / privacy laws / privacy policies) and, in some embodiments, can provide an additional security and compliance layer to the automation development pipeline. See, for example, U.S. Patent No. 11,733,668 (this document is incorporated herein by reference in its entirety).

[0075] The management functions can provide management, deployment, and optimization of automation across the entire organization. In some embodiments, the management functions may include orchestration, test management, AI capabilities, and / or insights. The management functions of the hyperautomation system 100 may also function as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management capabilities of the hyperautomation system 100 may include, but are not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interoperability of RPA robots, AOP, and / or AI agents.

[0076] Conductor applications such as UiPath Orchestrator® (which in some embodiments may be offered as part of UiPath Automation Cloud®, or as a cloud-native single-container suite via on-premises, VMs, private or public clouds, Linux® VMs, or UiPath Automation Suite®) provide orchestration capabilities for deploying, monitoring, optimizing, scaling, and securing RPA robots, AOP, and / or AI agent deployments. Test suites such as UiPath Test Suite® may provide test management for monitoring the quality of deployed automations. Test suites can facilitate test planning and execution, requirements fulfillment, and defect traceability. Test suites may include comprehensive test reports.

[0077] Analytics software (e.g., UiPath Insights®) can track, measure, and manage the performance of deployed automations. Analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes within the organization. Analytics software can present results in a dashboard format for easier user understanding.

[0078] A data service (e.g., UiPath Data Service®) may store data in a database 140, for example, and a drag-and-drop storage interface may allow data to be aggregated into a single, scalable, and secure location. In some embodiments, low-code or no-code data modeling and storage for automation may be provided, enabling seamless access while ensuring enterprise-grade security and scalability. AI functionality may be provided by an AI Center (e.g., UiPath AI Center®) that facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may make such functionality available even to non-data scientists. Deployed automations (e.g., RPA robots, AOP, and AI agents) may call AI / ML models, such as AI / ML model 132, from the AI ​​Center. The performance of AI / ML models may be monitored and improved using user-validated data provided by a data review center 160, etc. Users (as reviewers) may provide labeled data to the core hyper-automation system 120 via a review application 152 on a computing system 154. For example, a reviewer may verify that the predictions made by AI / ML model 132 and / or generative AI model 172 are accurate, or provide corrections if they are not. A user (as a reviewer) may also provide dynamic and / or direct user input to the AI ​​agent (e.g., within the scope of human-participatory operation), and such dynamic and / or direct user input (e.g., responses and corrections provided by the reviewer) may be used to train the AI ​​agent to make its LLMs more accurate. In other words, dynamic and / or direct user input may be stored as training data for retraining AI / ML model 132 and / or generative AI model 172, for example, in a database 140. The AI ​​center may then schedule and execute training jobs to train a new version of the AI / ML model using the training data.Both positive and negative examples are saved and can be used to retrain the AI / ML model 132 and / or the generative AI model 172.

[0079] The engagement feature involves automation and users as a single team to seamlessly collaborate on the desired process. In some embodiments, low-code applications (e.g., using UiPath Apps®) can be built to connect browser tabs and legacy software, including those lacking APIs. For example, applications can be quickly created in a web browser using a rich library of drag-and-drop controls. The application can connect to a single automation or multiple automations.

[0080] An Action Center (e.g., UiPath Action Center®) provides a user-friendly and efficient mechanism for handing over processes from automation to the user, or vice versa. The user can approve or escalate, create exceptions, etc. Automation can then execute automated functions of a given workflow.

[0081] A local assistant (e.g., UiPath Autopilot®) may be provided as a launching pad for users to initiate automations. Such an assistant may also provide semantic cut-and-paste functionality (e.g., UiPath Clipboard AI®). See, for example, U.S. Patent No. 12,124,806 and U.S. Patent Application Publications 2023 / 0107316, 2023 / 0415338, and 2024 / 0220581. This functionality may be provided, for example, on a tray provided by the operating system, allowing users to interact with RPA robots, AOPs, AI agents, and automation-driven applications on their computing system. The interface may list automations approved by a given user and allow the user to execute them. These may include ready-to-use automations from an automation marketplace or automations from an internal automation store in an automation hub, etc. Once an automation is executed, it may run as a local instance in parallel with other processes on the computing system, allowing the user to use the computing system while the automation is performing its operations. In certain embodiments, the assistant is integrated with a task capture function, allowing the user to document processes to be automated from the assistant launcher.

[0082] In some embodiments, the hyperautomation system 100 can provide end-to-end measurement and governance of automation programs at any scale. As previously mentioned, analytics (e.g., via UiPath Insights®) can be used to understand the performance of the automation. Data modeling and analytics using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards allow for data visualization across desired metrics, discovery of new analytical insights, tracking of performance indicators, discovery of automation ROI, telemetry monitoring on user computing systems, detection of errors and anomalies, and debugging of automation. An automation management console (e.g., UiPath Automation Ops®) is provided to manage automation throughout its entire lifecycle. Organizations can control how automation is built, what users can do with automation, and which automations users can access.

[0083] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes are discovered, automation is built, tested, deployed, performance is measured, the use of automation is easily made available to users, feedback is obtained, AI / ML models are trained and retrained, and the process is repeated. This promotes a more robust and effective automation suite.

[0084] In some embodiments, the generative AI model 172 is used as described above. For example, an AI agent utilizes the generative AI model. The generative AI model 172 can generate various types of content, such as text, images, audio, and synthetic data. Types of generative AI models may include, but are not limited to, LLMs, generative opposite networks (GANs), diffusion models, flow-based models, variational autoencoders (VAEs), transformers, etc. For example, in the case of LLMs, NLP models such as word2vec, BERT, GPT-3, and ChatGPT may be used in some embodiments to facilitate semantic understanding and provide more accurate and human-like responses. These models may be part of the AI / ML model 132 hosted on server 130. For example, the generative AI model 172 may be trained on a large text information corpus to perform semantic understanding, grasp what is present on the screen from text, and automatically generate code. An AI agent may use such a generative AI model 172. In certain embodiments, generative AI models 172 provided by existing cloud ML service providers such as OpenAI®, Google®, Amazon®, Microsoft®, IBM®, Nvidia®, and Meta® may be used and trained to provide such functionality. In embodiments where the generative AI model 172 is remotely hosted, a server 130 may be configured to integrate with a third-party API, allowing the server 130 to send requests containing the necessary input information to the generative AI model 172 and receive responses (e.g., semantic matching of fields between application versions, classification of application types on a screen, responses to natural language queries from the user, etc.). Such embodiments may provide a more advanced and sophisticated user experience and access to state-of-the-art NLP and other ML capabilities offered by these companies.

[0085] One aspect of the generative AI model 172 in some embodiments is the use of transfer learning. In transfer learning, a pre-trained generative AI model, such as an LLM, is fine-tuned for a specific task or domain. This allows the LLM to leverage knowledge already learned during initial training and adapt to a specific application. In the case of an LLM, the pre-training stage typically involves training the LLM on a large text corpus consisting of billions of words. At this stage, the LLM learns the relationships between words and phrases, enabling it to generate consistent and human-like responses to text input. The output of this pre-training stage is an LLM with a high level of understanding of fundamental patterns in natural language.

[0086] In the fine-tuning phase, a pre-trained LLM is adapted to a specific task or domain by being trained on a small, task-specific dataset. For example, in some embodiments, the LLM may be trained to analyze specific or multiple types of data sources to improve its accuracy regarding their content. This data may include, but is not limited to, prompt tuning or instruction tuning (where the model is specifically trained to better understand and follow specific types of instructions or prompts, improving its ability to perform specific tasks when given appropriate instructions). Such information may be provided as part of the training data, allowing the LLM to focus on these areas and learn to more accurately identify data elements within those data sources. Through fine-tuning, the LLM can learn the nuances of the task or domain (e.g., specific vocabulary and syntax used in that domain) without requiring as much data as would be needed to train the LLM from scratch. By leveraging the knowledge gained in the pre-training phase, a fine-tuned LLM can achieve state-of-the-art performance on a specific task with relatively little training data.

[0087] LLM can utilize vector databases. Vector databases index, store, and provide access to structured or unstructured data (e.g., text, images, time-series data) along with their vector embeddings. Data such as text can be tokenized, with single characters, words, or sequences of words parsed as tokens. These tokens are "embedded" into vector embeddings, which are numerical representations of the data. Vector databases enable LLM in production environments to rapidly search for and retrieve similar objects on a scale impossible to achieve manually.

[0088] AI and ML allow unstructured data to be numerically represented as vector embeddings without losing their semantic meaning. A vector embedding is a long list of numbers, each describing a feature of the data object represented by the embedding. Similar objects group together closely in the vector space; that is, the more similar the objects, the closer their vector embeddings are to each other. Similar objects can be found using vector search, similarity search, or semantic search and search. The distance between vector embeddings can be calculated using various techniques, including, but not limited to, squared Euclidean distance (L2-squared distance), Manhattan distance (L1 distance), cosine similarity, dot product, and Hamming distance. It may be beneficial to select the same metrics used to train the AI / ML model.

[0089] Vector indexing can be used to organize vector embeddings and efficiently retrieve data. Calculating the distance between a vector embedding and all other vector embeddings in a vector database using the k-nearest neighbors (kNN) algorithm can be computationally intensive when the number of data points is large. This is because the required computations increase linearly (O(n)) with respect to the number of dimensions and data points. Finding similar objects using an approximate nearest neighbor (ANN) approach is more efficient. Similar objects can be found much faster because the distances between vector embeddings are pre-calculated, and similar vectors are organized and stored close to each other (e.g., in clusters or graphs). This process is called "vector indexing." ANN algorithms that can be used in several embodiments may include, but are not limited to, clustering-based indexing, proximity-graph-based indexing, tree-based indexing, hash-based indexing, and compression-based indexing.

[0090] Figure 2 shows some of the combined capabilities 200 of an AI agent 210 and an RPA robot 220 in one or more embodiments. The AI ​​agent 210 is configured to process natural language instructions and achieve expected goals 231 therefrom, executes through dynamic decision-making or dynamic flow control with self-correcting capabilities 233, stores information in long-term memory and evaluates its own execution performance 235, and learns from human participation and self-performance during execution 237. The RPA robot 220 may be used as a tool by the AI ​​agent 210 to respond to triggers 241 (e.g., from a conductor application such as UiPath Orchestrator®), respond based on context 243 (i.e., the RPA robot 220 may take information from the context to perform deterministic steps based on the acquired context, e.g., update a document based on the acquired context; alternatively, the agent 210 may use the acquired context to update a dynamic plan and follow instructions to complete the goal), leverage models 245 (e.g., CV models, document processing models, speech-to-text models, OCR models, AI models, etc.), leverage tools 247 (e.g., tools available in the RPA ecosystem, completed automations, in-automation workflows, integrated service connector calls for third-party and first-party services, RPA designer application activities, LLM calls, automations, etc.), and perform actions 249 (i.e., using the RPA robot 220 as a tool) that the RPA robot 220 may take based on input from the AI ​​agent 210. The AI ​​agent 210 may also perform actions 249 to update its memory, update its plan to follow instructions and achieve goals, self-evaluate and learn from its actions, self-correct when it encounters problems, and escalate to the user when assistance is needed.

[0091] As described above, agentic automation achieves various technical effects, benefits, and advantages. Agentic automation improves memory usage by reducing the storage space required for data and improves processor efficiency by reducing the number of calls and actions. Agentic automation provides the ability to process gigabytes, terabytes, petabytes, or more of data that would be impossible to implement mentally or manually by humans. Agentic automation also makes it possible to reduce the number of triggers and models used through dynamic decision-making. For example, as described herein, while conventional software automation alone may require 100 (100) actions in the exemplary scenario, agentic automation in the same exemplary scenario can significantly reduce the required actions (e.g., to 15 (15) actions). Agentic automation can also improve the efficiency of LLM by using context grounding to anchor the AI ​​agent 210 to the desired context and "constrain" the LLM to the relevant context.

[0092] In this specification, “context grounding” refers to a methodology that improves models such as LLMs by integrating company-specific information with pre-trained knowledge, enabling accurate responses to specialized or up-to-date queries. In one embodiment, context grounding uses external data to augment LLM responses, enabling LLMs to answer queries within a given context, even for matters they would not inherently know. For example, proprietary industry jargon and complex document structures can be challenges in ensuring effective retrieval and semantic matching, but context grounding solves this problem by precisely chunking documents, allowing relevant information derived from proprietary industry jargon and complex document structures to be passed to the LLM without noise. As another example, context grounding improves LLM responses by providing extract and retrieval techniques tailored to diverse industries and applications (e.g., proprietary industry jargon and complex document structures).

[0093] Figure 3 shows a diagram 300 of an AOP, AI agent, RPA robot, and application according to one or more embodiments.

[0094] AOP pool 310 includes AOP1, 2, ..., P, which implement business processes. As mentioned above, AOP can be implemented as BPMN and executed by an AOP execution engine such as Temporal®. AOP can utilize AI agents and / or RPA robots to execute parts of its business processes.

[0095] The AI ​​agent pool 320 includes AI agents 1, 2, ..., I that have been trained to perform various tasks such as investigating claims, exploring solutions with employees, and summarizing policies and technical specifications. The RPA robot pool 330 includes RPA robots 1, 2, ..., J that perform various automations such as UI automation, semantic matching automation, and form input automation.

[0096] The application pool 340 includes applications 1, 2, ..., K that AI agents and / or RPA robots can interact with. For example, applications may include CRM applications, invoicing applications, payroll applications, banking applications, web applications, legacy system applications, word processor applications, spreadsheet applications, email applications, etc. The AI ​​agents, RPA robots, and applications may reside on a single computer system, or they may be distributed across multiple or numerous computer systems. The AOP is typically on the cloud or other server side, and in some embodiments, it may reside on the same computer system as the conductor application 350.

[0097] AOP can trigger or invoke AI agents and RPA robots via the conductor application 350. AI agents and RPA robots can also trigger or invoke each other via the conductor application. For example, to invoke an RPA robot, an AI agent may make a "Start Job" call in the conductor application 350. Note that RPA robots are deployed as automation controlled by the conductor application 350. AI agents, AOP, and RPA robots can also trigger or invoke specific applications. For example, based on information obtained from human-in-the-loop operations, an AI agent may dynamically learn which RPA robots, other AI agents, and / or applications to trigger or invoke in order to accomplish a task. For example, an AI agent may learn to trigger an RPA robot via the conductor application 350 to fill out and submit a web form. The AI ​​agent may also learn to open Microsoft Excel® and enter form information into the appropriate tab, or open and update a payroll application, etc. The AI ​​agent may also learn to call or trigger an email resolution AI agent via the conductor application 350 to contact a bank customer service representative if a problem occurs. The technical effects, benefits, and advantages may, in some embodiments, be similar to those described above with respect to Figures 1 and 2.

[0098] AI agents may belong to a tenant so that AI agents, AOPs, and RPA robots can find each other. A designer application may call a conductor to retrieve a list of available RPAs. In one embodiment, there are three ways to obtain automation capabilities: (1) a user providing a description of what the automation does when creating a workflow in a designer application; (2) using AI agents and ML technologies to generate a summary of what a particular workflow does; or (3) a developer describing what the automation does in a designer application. A conductor application may also maintain a list of applications available to a particular AI agent and RPA robot. That is, descriptions of available AI agents, RPA robots, and / or applications are derived or assigned by the AI ​​agent, ML technology, or user.

[0099] Figure 4 shows an example of an agent service interface 400 according to one or more embodiments. As shown in Figure 4, the agent answers questions regarding policy documents provided within context grounding. The agent instruction pane 410 contains a natural language description entered by the user about what the AI ​​agent is intended to do. The user prompt 420 allows the developer to enter the content of the user prompt in the content field 422 as needed. The tool dropdown 430 allows the developer to select the tools the AI ​​agent will use, such as using an API for the application or calling an RPA robot to perform RPA.

[0100] The context dropdown 440 allows developers to configure context grounding for the AI ​​agent. The context configuration pane 442 allows developers to provide descriptions via the description field 444 and ECS indexes via the Elastic Common Schema (ECS) index field 446 for specific policy documents, such as contracts, conditions, and information on what to do. Developers can also add contexts 450 to further supplement the context grounding. User escalation options can be configured via the escalation dropdown 460. The query field 470 allows users to provide queries that the AI ​​agent will respond to. When the user clicks the run button 480, the AI ​​agent executes the query. The results of the AI ​​agent execution are then displayed in the run pane 490 as the AI ​​agent retrieves and outputs the results.

[0101] Figure 5 shows an example of an AOP development interface 500 according to one or more embodiments. The AOP development interface 500 includes an AOP 510, an AI agent 520, and an RPA robot 530 that can be selected by the user when developing a business process. These can be selected by the AI ​​agent, AOP, or developer and dragged onto the canvas 540 to develop the AOP. In this example, a credit check 541 is implemented by retrieving customer data 543 from the database 545 based on a credit check request 542. Next, the AI ​​agent 546 is called and determines the customer type (e.g., very likely to pay, likely to default on payments, frequently unemployed, etc.) by analyzing the customer data in the database 545. This type is then provided to the RPA robot 547, which takes this information into consideration when performing the credit check and generating the credit check result 549. Alternatively, the AI ​​agent, AOP, or developer can enter a description of the business process in field 550 and select the generate button 560. This description is provided to the LLM, which attempts to understand the business process and automatically generate the AOP. Subsequently, the AI ​​agent, AOP, or developer can edit the AOP.

[0102] Figure 6 shows an example of an RPA development interface 600 according to one or more embodiments. The RPA development interface 600 includes a component 610 that can be selected by an AI agent, AOP, or developer when developing a workflow for an RPA robot. The AI ​​agent, AOP, or developer can select this and drag it onto the canvas 620. Alternatively, the AI ​​agent, AOP, or developer can enter a description of the workflow in field 630 and select the generate button 640. This description is provided to the LLM, which attempts to understand the workflow and automatically generate an RPA robot. The AI ​​agent, AOP, or developer can then edit the workflow of the RPA robot. Note that the functions shown and described with respect to Figures 4, 5, and 6 may, in some embodiments, be provided by a single designer application.

[0103] Figure 7 shows an end-to-end AI agent, RPA robot, and AOP development and deployment system 700 according to one or more embodiments. A designer application 710 enables the AI ​​agent, AOP, and developer to design workflows for subsequent automation (e.g., AOP, AI agent, and / or RPA robot). After these AOP, AI agent, and / or RPA robots are tested and validated, the validated and tested AOP, AI agent, and / or RPA robots are packaged and published to the automation database 720.

[0104] The conductor application 730 manages the deployment of these packaged and exposed automations. When a software process 732 requests the execution of an automation, the conductor application 730 sends a job start command to the AOP engine 740, which then selects and starts the automation from AOP 742. During the execution of AOP 742, it may encounter steps implemented by an AI agent 750 and / or an RPA robot 760. In some cases, the AOP engine 740 pauses the running AOP 742 and sends a request to the conductor application 730 to send a job start request to the appropriate AI agent 750 or RPA robot 760 to perform the step in question.

[0105] When an AI agent is requested, the conductor application 730 sends a job start request to the appropriate AI agent 750. This request may include natural language text or other information provided to the conductor application 730 from the AOP engine 740. The AI ​​agent 750 then performs this step by executing LLM 752 to assist in task execution. The AI ​​agent 750 sends information related to the task (e.g., requested information, instructions that the step was completed, instructions that the step failed, etc.) to the conductor 730, which then provides this information to the AOP engine 740. The AOP engine 740 then resumes operation.

[0106] When an RPA robot is requested, the conductor application 730 sends a job start request to the appropriate RPA robot 760. The RPA robot 760 then executes the requested RPA 762. The RPA robot 760 sends information related to the task (e.g., requested information, instructions that a step was completed, instructions that a step failed, etc.) to the conductor application 730, which then provides this information to the AOP engine 740. The AOP engine 740 then resumes operation.

[0107] According to one or more embodiments, the AOP 742, AI agent 750, or RPA 762 may request user action. In this case, the AOP engine 740, AI agent 750, or RPA robot 760 contacts user 770 for human-in-the-loop operation that contributes to automation. After user 770 provides user action, the AOP engine 740, AI agent 750, or RPA robot 760 resumes automation.

[0108] Figure 8 shows an architectural diagram of an agent-based automation and RPA system 800 according to one or more embodiments. In one embodiment, the agent-based automation and RPA system 800 is part of the hyperautomation system 100 in Figure 1. The agent-based automation and RPA system 800 includes an AI agent, AOP, or a designer 810 that enables developers to design automations (e.g., workflows, natural language instructions for AI agents and AOPs, context grounding, tool configurations, RPA robots, AOPs, AI agents, etc.). The designer 810 provides solutions for application integration and can enable the automation of third-party applications, administrative information technology (IT) tasks, and business IT processes. The designer 810 facilitates the development of automation projects, which are graphical representations of business processes. The designer 810 facilitates the development and deployment of automations (indicated by arrow 811). The designer 810 may be an application running on a user's desktop, an application running remotely in a VM, a web application, etc.

[0109] An automation project enables the automation of rule-based processes by allowing an AI agent, AOP, or developer to control the execution order and relationships of a set of custom steps, or "activities," developed within a workflow, as described herein. One commercial example of an embodiment of Designer 810 is UiPath Studio®. Each activity may include actions such as clicking a button, reading a file, or writing to a log panel. In some embodiments, the workflow may take the form of a nested or embedded structure.

[0110] Workflow types may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. Sequences are particularly suitable for linear processes that flow from one activity to the next without complicating the workflow. Flowcharts are particularly suitable for more complex business logic, enabling the integration of decision-making and diverse ways of connecting activities through multiple branching logic operators. FSMs are particularly suitable for large-scale workflows. FSMs use a finite number of states in their execution, which are triggered by conditions (i.e., transitions) or activities. Global exception handlers are particularly suitable for debugging processes by determining how a workflow behaves when it encounters execution errors.

[0111] Once automation is developed in Designer 810, the execution of business processes is orchestrated by Conductor 820, which orchestrates one or more robots 830, one or more AI agents 850, and / or one or more AOPs 870 that execute the workflow developed in Designer 810. A commercial example of one embodiment of Conductor 820 is UiPath Orchestrator®. Conductor 820 facilitates the creation, monitoring, and deployment management of resources in an environment. Conductor 820 can function as an integration point with third-party solutions and applications. As described above, in some embodiments, Conductor 820 may be part of the Core Hyperautomation System 120 in Figure 1.

[0112] It should be noted that the RPA robot 830 can operate independently for deterministic processes. The AI ​​agent 850 and AOP 870 can also operate independently (for example, for non-deterministic processes), or they can utilize the RPA robot 830 or other AI agent 850 as tools to achieve part of their agent-based automation. The AI ​​agent 850 can drive composite automations that utilize both the RPA robot 830 and the AI ​​agent 850 (or vice versa), and the AOP 870 can include such composite automations.

[0113] The conductor 820 can manage a fleet of RPA robots 830 and AI agents 850, and can connect and run the RPA robots 830 and AI agents 850 from a central aggregation point (e.g., if requested by an AOP engine implementing AOP), as indicated by arrow 881. The types of RPA robots 830 that can be managed include, but are not limited to, attendant robots, unattended robots, development robots (similar to unattended robots but used for development and testing purposes), and non-production robots (similar to attendant robots but used for development and testing purposes). Attended robots are triggered by user events and operate in parallel with the user on the same computing system. Attended robots can be used with the conductor 820 for centralized process deployment and log recording. Attended robots can assist users in accomplishing various tasks and can be triggered by user events. In some embodiments, this type of robot cannot be started from the conductor 820 and / or run under a locked screen. In some embodiments, the attendant robot can only be started from a robot tray or command prompt. In some embodiments, the attendant robot should be run under user supervision.

[0114] Unattended robots run unattended in a virtual environment and can automate many processes. Unattended robots can provide remote execution, monitoring, scheduling, and work queue support. Debugging of all robot types can be performed in the designer 810 in some embodiments. Both attended and unattended robots can automate applications on mainframes, web applications, VMs, enterprise applications (e.g., those manufactured by SAP®, Salesforce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.), but are not limited to these.

[0115] Conductor 820 may have various functions, including but not limited to provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interoperability, as indicated by arrow 882. Provisioning may include creating and maintaining connections between RPA robots 830, AI agents 850, and / or AOP 870 and Conductor 820 (e.g., a web application). Deployment may include ensuring that package versions are correctly delivered to RPA robots 830, AI agents 850, and / or AOP 870 assigned for execution. Configuration may include maintaining and delivering the environment and process configuration for RPA robots and AI agents. Queuing may include providing management of queues and queue items. Monitoring may include tracking identification data for robots and AI agents and maintaining user privileges. Logging may include storing and indexing logs in a database (e.g., a Structured Query Language (SQL) database or a "Not Only" SQL (NoSQL) database) and / or other storage mechanisms (e.g., ElasticSearch® for storing and rapidly querying large datasets). Conductor 820 may provide interoperability by acting as a central communication point for third-party solutions and / or applications.

[0116] The RPA robot 830 is an execution agent that implements workflows built in the designer 810. A commercial example of several embodiments of the RPA robot 830 is UiPath Robots®. In some embodiments, the RPA robot 830 installs the Microsoft Windows® Service Control Manager (SCM) management service by default. As a result, such an RPA robot 830 can open an interactive Windows® session under the local system account and has Windows® service privileges.

[0117] In some embodiments, the RPA robot 830 can be installed in user mode. For such an RPA robot 830, this means that it has the same privileges as the user on which it is installed. This feature may also be available for high-density (HD) robots, ensuring that each machine is fully utilized to its maximum potential. In some embodiments, any type of RPA robot 830 may be configured in an HD environment.

[0118] In some embodiments, the RPA robot 830 is divided into several components, each dedicated to a specific automation task. The robot components include, but are not limited to, an SCM-managed robot service, a user-mode robot service, an executor, an agent, and a command line. The SCM-managed robot service manages and monitors Windows® sessions and acts as a proxy between the conductor 820 and the execution host (i.e., the computing system on which the RPA robot 830 runs). These services are trusted and manage the credentials of the RPA robot 830. A console application is launched by the SCM under the local system.

[0119] In some embodiments, the user-mode robot service manages and monitors Windows® sessions and acts as a proxy between the conductor 820 and the execution host. The user-mode robot service is trusted and can manage the credentials of the RPA robot 830. If the SCM management robot service is not installed, the Windows® application may be launched automatically.

[0120] The executor can execute a given job (i.e., execute a workflow) under a Windows® session. The executor can be aware of per-monitor dots per inch (DPI) settings. AI Agent 850 can be a Windows® Presentation Foundation (WPF) application that displays available jobs in a system tray window. Note that these agents are different from AI Agent 850. AI Agent 850 can be a service client and can request to start or stop jobs and change their settings. The command line is a service client. The command line is a console application that can request to start a job and await its output.

[0121] By dividing the RPA robot 830 into components as described above, developers, support users, and computing systems can more easily execute, identify, and track what each component is doing. This method allows for configuring specific behaviors for each component, such as setting different firewall rules for the executor and services. In some embodiments, the executor may always be aware of the DPI settings per monitor. As a result, workflows can run at any DPI, regardless of the configuration of the computing system on which they were created. Projects from the designer 810 may also be independent of the browser's zoom level in some embodiments. For applications that are not DPI-compatible or are intentionally marked as not DPI-compatible, DPI can be disabled in some embodiments.

[0122] The agent-based automation and RPA system 800 of this embodiment is part of a hyperautomation system such as the hyperautomation system 100 in Figure 1. Developers can use the designer 810 to build and test RPA, AOP, and AI agents that utilize AI / ML models deployed in the core hyperautomation system 840 (for example, as part of its AI center). Such an RPA robot can send inputs for the execution of an AI / ML model to the core hyperautomation system 840 and receive its output via the core hyperautomation system 840.

[0123] One or more of the RPA robots 830 may be listeners, as described above. These listeners may provide the core hyperautomation system 840 with information about what the user is doing while using the computing system. This information can then be used by the core hyperautomation system for process mining, task mining, task capture, etc.

[0124] An assistant / chatbot (of Conductor 820) may be provided on the user computing system to enable the user to launch a local RPA robot. The assistant / chatbot may be placed, for example, in the system tray. The chatbot may have a user interface so that the user can see the text within the chatbot. Alternatively, the chatbot may lack a user interface, operate in the background, and listen to the user's voice using the computing system's microphone.

[0125] In some embodiments, data labeling may be performed by a user of the computing system on which the RPA robot or AI agent is running, or by a user of another computing system from which the robot or AI agent provides information. For example, if a robot invokes an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify buttons on the screen, the user may draw rectangles around the misidentified or unidentified components and provide text with the correct identification. This information is provided to the core hyperautomation system 540, which can then be used to train a new version of the AI / ML model.

[0126] Figure 9 is an architectural diagram showing a deployed RPA system 900 in one or more embodiments. In some embodiments, the RPA system 900 may be part of the agent-based automation and RPA system 800 in Figure 8 and / or the hyperautomation system 100 in Figure 1. Note that the architecture of the deployed RPA system 900 may not be used in some embodiments. The deployed RPA system 900 may be a cloud-based system, an on-premise system, a desktop-based system (providing enterprise-level, user-level, or device-level automation solutions for automating different computing processes), etc.

[0127] It should be noted that the client side 901, the server side 902, or both may include any number of computing systems without departing from the scope of the present invention. On the client side 901, the robot application 910 includes an executor 912, an execution agent 914, and a designer 916. However, in some embodiments, the designer 916 may not run on the same computing system as the executor 912 and the execution agent 914. The executor 912 executes processes. Multiple business projects may be executed simultaneously. The execution agent 914 (e.g., Windows® service) is a single point of contact for all executors 912 in this embodiment. The execution agent 914 is also responsible for transmitting the robot's status (e.g., periodically sending "heartbeat" messages indicating that the robot is still functioning) and downloading the required versions of packages to be executed.

[0128] The listener 930 monitors and records data relating to user operations on the attended computing system and / or the operation of the unattended computing system in which the listener 930 resides. The listener 930 may be an RPA robot, an AOP, an AI agent, part of an operating system, a downloadable application for the computing system, or any other software and / or hardware that does not deviate from the scope of the present invention. In fact, in some embodiments, the logic of the listener is partially or completely implemented by physical hardware.

[0129] On the server side 902, the presentation layer 933, service layer 934, and persistence layer 935 are provided, and the conductor 940 is also provided. Furthermore, the presentation layer 933 includes a web application 942, an Open Data Protocol (oData) Representational State Transfer (REST) ​​Application Programming Interface (API) endpoint 944, and notification / monitoring 946, and the service layer 934 includes the API implementation / business logic 948. Thus, as shown in Figure 9, the conductor 940 includes the web application 942, the oData REST API endpoint 944, notification / monitoring 946, and the API implementation / business logic 948. The persistence layer 935 includes a database server 950, an AI / ML server 960, and an indexer server 970.

[0130] In some embodiments, communication between the execution agent 914 and the conductor 940 is always initiated by the execution agent 914. In notification scenarios, the execution agent 914 may open a WebSocket channel, which the conductor 940 later uses to send commands (e.g., start, stop, etc.) to the RPA robot. Note that, although not shown here to reduce the complexity of Figure 9, the AI ​​agent can also interact with the conductor 940, as described above with respect to, for example, Figures 1 and 8. The conductor 940 can orchestrate the actions of the AI ​​agent, AOP, and RPA robot.

[0131] In this embodiment, all messages are logged to the conductor 940, which then processes them via the database server 950, the AI / ML server 960, the indexer server 970, or any combination thereof. As described herein and with reference to Figure 8, the executor 912 may be a robotic component.

[0132] In some embodiments, an RPA robot represents an association between a machine name and a username. A robot may manage multiple executors simultaneously. On computing systems that support the simultaneous execution of multiple interactive sessions (e.g., Windows® Server 2012), multiple RPA robots may run concurrently, each running in a separate Windows® session with a unique username.

[0133] The conductor 940 can also facilitate interaction between the AI ​​agent and the AI / ML model via the AI / ML server 960. The AI / ML server 960 can store and / or facilitate access to the generated AI model (e.g., the generated AI model 172 in Figure 1).

[0134] In some embodiments, most operations performed by the user on the conductor 940 interface (e.g., via the browser 981) are carried out by calling various APIs. Such operations include, but are not limited to, starting jobs on robots, adding / deleting data in queues, scheduling unattended jobs, etc. The web application 942 is the visual layer of the server platform. In this embodiment, the web application 942 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any markup language, scripting language, or any other form may be used without departing from the scope of the invention. In this embodiment, the user interacts with web pages from the web application 942 via the browser 941 and performs various operations to control the conductor 940. For example, the user may create RPA robot groups, assign packages to RPA robots, perform log analysis per RPA robot and / or process, start and stop RPA robots, etc.

[0135] In addition to the web application 942, the conductor 940 also includes a service layer that exposes an oData REST API endpoint 944. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 942 and the execution agent 914. In this embodiment, the execution agent 914 is a supervisor of one or more robots on a client computer.

[0136] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functions. The configuration endpoint may be used in some embodiments to define and configure application users, permissions, robots, assets, releases, and environments. The logging REST endpoint may be used to log different information, such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used for robots to query the package version that should be executed when a start job command is used on conductor 940. The queuing REST endpoint may be responsible for queue and queue item management (e.g., adding data to the queue, retrieving transactions from the queue, setting the status of transactions, etc.).

[0137] A monitoring REST endpoint may monitor the web application 942 and the execution agent 914. The notification and monitoring API 946 may be a REST endpoint used for registering the execution agent 914, delivering configuration settings to the execution agent 914, and sending and receiving notifications between the server and the execution agent 914. In some embodiments, the notification and monitoring API 946 may also use WebSocket communication. As shown in Figure 9, one or more activities / operations described herein are represented by arrows 949.

[0138] The APIs of service layer 934 may, in some embodiments, be accessed through the configuration of appropriate API access paths, for example, based on whether the conductor 940 and the entire hyperautomation system are deployed on-premises or cloud-based. The APIs of conductor 940 may provide custom methods for querying statistics about various entities registered with conductor 940. Each logical resource may, in some embodiments, be an oData entity. In such entities, components such as robots, processes, and queues may have properties, relationships, and operations. The APIs of conductor 940 may be consumed by the web application 942 and / or execution agent 914 in two ways, in some embodiments: (1) obtaining API access information from conductor 940, or (2) registering an external application and using the oAuth flow.

[0139] The persistence layer 935 in this embodiment includes three servers: a database server 950 (e.g., an SQL server), an AI / ML server 960 (e.g., a server providing AI / ML model provisioning services such as AI Center functions), and an indexer server 970. In this embodiment, the database server 950 stores the configurations of robots and AI agents, robot and AI agent groups, AOP, associated processes, users, roles, schedules, etc. This information is managed through a web application 942 in some embodiments. The database server 950 may manage queues and queue entries. In some embodiments, the database server 950 may store messages logged by robots and AI agents (in addition to or instead of the indexer server 970). The database server 950 may also store process mining, task mining, and / or task capture-related data received, for example, from a listener 930 installed on the client side. Although no arrow is shown between listener 930 and database 950, it should be understood that in some embodiments, listener 930 can communicate with database 950 and vice versa. This data may be stored in the form of PDDs, images, XAML files, etc. Note that structured and / or unstructured data may be stored. Listener 930 may be configured to intercept user activity, processes, tasks, and performance metrics on the computing system in which listener 930 resides. For example, listener 930 may record user activity (e.g., clicks, typed characters, position, applications, active elements, time, etc.) on its computing system and then convert it into a format suitable for being provided to and stored in database server 950.

[0140] The AI / ML Server 960 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options make such functionality accessible even to non-data scientists. Deployed automations (e.g., RPA robots and / or AI agents) can invoke AI / ML models from the AI / ML Server 960. AI / ML model performance can be monitored, trained, and improved using user-validated data. The AI / ML Server 960 can schedule and execute training jobs to train new versions of AI / ML models. The AI / ML model server can also store and / or provide access to generated AI models.

[0141] The AI / ML server 960 may store data relating to AI / ML models and ML packages for users to configure various ML skills during development. In this specification, "ML skill" refers to a pre-built and trained ML model for a particular process, which can be used, for example, through automation. The AI / ML server 960 may also store data relating to document understanding techniques and frameworks, algorithms, and software packages for various AI / ML functions, including but not limited to intent analysis, NLP, speech analysis, and various types of AI / ML models.

[0142] An indexer server 970, which is optional in some embodiments, stores and indexes information logged by the robot. In some embodiments, the indexer server 970 can be disabled through configuration settings. In some embodiments, the indexer server 970 uses ElasticSearch®, an open-source full-text search engine. Messages logged by the robot (e.g., using log messages or activities such as light lines) are sent to the indexer server 970 via a logging REST endpoint, where they can be indexed for future use.

[0143] Figure 10 is an architectural diagram showing the relationships between Designer 1010, Activities 1020, 1030, 1040, 1050, Driver 1060, API 1070, and AI / ML Model 1080 in one or more embodiments. As described above, an AI agent, AOP, or developer uses Designer 1010 to develop workflows for automation (e.g., workflows executed by RPA robots, AI agents, and AOPs).

[0144] The AI ​​agent, AOP, or developer may design and configure workflow 1092, agent-based automation 1094 for the AI ​​agent (e.g., providing natural language descriptions, context grounding, tools, etc., to the AI ​​agent), and AOP 1096 (see also Figures 4B, 5, and 6). In some embodiments, various types of activities may be displayed to the developer. The designer 1010 may be local to the user's computing system or remote to it (e.g., accessed via a VM or via a local web browser interacting with a remote web server). The workflow for the RPA robot may include user-defined activities 1020, API-driven activities 1030, AI / ML activities 1040, and / or UI automation activities 1050. The user-defined activities 1020 and API-driven activities 1030 interact with the application via their APIs. In some embodiments, a user-defined activity 1020 and / or an AI / ML activity 1040 may call one or more AI / ML models 1080, which may be local to or remote from the computing system on which the robot operates.

[0145] In some embodiments, non-textual visual components within an image can be identified, which are referred to herein as CVs. However, it should be noted that in some embodiments, CVs include OCRs. CVs may be performed, at least in part, by an AI / ML model 1080. CV activities relating to such components may include, but are not limited to, text extraction from segmented label data using OCRs, fuzzy text matching, cropping of segmented label data using MLs, and comparison of text extracted from label data with ground truth data. In some embodiments, the number of activities that can be implemented in a user-defined activity 1020 may be in the hundreds or thousands. However, any number and / or types of activities may be used without departing from the scope of the invention.

[0146] The UI automation activity 1050 is a subset of special low-level activities written in lower-level code to facilitate interaction with the screen. The UI automation activity 1050 facilitates these interactions via a driver 1060 that enables the robot to interact with desired software. For example, the driver 1060 may include an operating system (OS) driver 1062, a browser driver 1064, a VM driver 1066, an enterprise application driver 1068, etc. In some embodiments, one or more AI / ML models 1080 may be used to perform the interaction with the computing system by the UI automation activity 1050. In some embodiments, the AI / ML model 1080 may extend or completely replace the driver 1060. In fact, in some embodiments, the driver 1060 is not included.

[0147] Driver 1060 can interact with the OS at a low level via OS driver 1062 for purposes such as hook discovery and key input monitoring. Driver 1060 can facilitate integration with applications such as Chrome®, IE®, Citrix®, and SAP®. For example, a "click" activity can perform the same role in these different applications via driver 1060.

[0148] Figure 11 is an architectural diagram showing a computing system 1100 configured to provide advanced agent-based extraction and retrieval for context grounding within automation, in one or more embodiments. In some embodiments, the computing system 1100 may be one or more of the computing systems described and / or mentioned herein. In some embodiments, the computing system 1100 may be part of a hyperautomation system, such as those shown in Figures 1 and 8. The computing system 1100 includes a bus 1105 or other communication mechanism for communicating information and a processor 1110 coupled to the bus 1105 for processing information. The processor 1110 may be any kind of general-purpose or application-specific processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The processor 1110 may also have multiple processing cores, at least some of which may be configured to perform specific functions. In some embodiments, multi-parallel processing may be used. In some embodiments, at least one of the processors 1110 may be a neuromorphic circuit that includes processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit does not require typical components of a von Neumann computing architecture.

[0149] The computing system 1100 further includes memory 1115 for storing information and instructions executed by the processor 1110. Memory 1115 may consist of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage such as magnetic or optical disks, or any other non-volatile computer-readable medium, or any combination thereof. The non-volatile computer-readable medium may be any available medium accessible by the processor 1110, and may include volatile media, non-volatile media, or both. The medium may be removable, non-removable, or both. The computing system 1100 also includes communication equipment 1120, such as transceivers, for providing access to a communication network via wireless and / or wired connections. In some embodiments, the communication equipment 1120 may include one or more antennas, including single, array, phased, switched, beamforming, beam-steering, a combination thereof, or any other antenna configuration.

[0150] The processor 1110 is further coupled to the display 1125 via the bus 1105. Any suitable display device and haptic I / O may be used without departing from the scope of the present invention. A keyboard 1130 and a cursor control device 1135 such as a computer mouse or touchpad are also coupled to the bus 1105 to allow the user to interface with the computing system 1100. However, in some embodiments, there is no physical keyboard and mouse, and the user may interact with the device only via the display 1125 and / or a touchpad (not shown). Any type and combination of input devices may be used as a design choice. In some embodiments, there is no physical input device and / or display. For example, the user may interact with the computing system 1100 remotely via another computing system that is communicably connected, or the computing system 1100 may operate autonomously.

[0151] Memory 1115 stores software modules that provide functionality when executed by processor 1110. These modules include an operating system 1140 for computing system 1100.

[0152] The module further includes an Extract and Search module 1145 configured to perform all or part of the processes described herein, or derivatives thereof, including, but not limited to, the performance of advanced agent-based extract and search for context grounding within automation. The Extract and Search module 1145 is configured to perform RAG, extract, and semantic storage. RAG is a technique that retrieves relevant data chunks from a document repository and integrates them into prompts for an AI model to improve response accuracy. Extract is the process of parsing various document formats (e.g., PDFs) to extract meaningful text and flattening complex data structures (e.g., tables and images) to make them available for storage and retrieval. Semantic storage involves storing the extracted information in a way that enables meaning-based, rather than exact-match-based, search and retrieve, using vector embeddings (numerical representations of text).

[0153] The module further includes an agent-type memory module 1147 configured to perform all or part of the processes described herein, or derivatives thereof, including, but not limited to, the execution of agent-type memory for storing context grounding and other results for automation. For example, the agent-type memory module 1147 is configured to dynamically cache (i.e., store) escalations, tool calls, user interactions or feedback, context grounding, etc., to provide improved efficiency and minimized calls.

[0154] The computing system 1100 may include one or more additional function modules 1150 that include additional functions.

[0155] Those skilled in the art will understand that “computing system” can be embodied as a server, embedded computing system, personal computer, console, personal digital assistant (PDA), mobile phone, tablet computing device, smartwatch, quantum computing system, or any other suitable computing device or combination thereof. Presenting the above functions as being performed by a “system” is not intended to limit the scope of the invention in any way, but rather to provide an example of a number of embodiments of the invention. Indeed, the methods, systems, and devices disclosed herein can be implemented in local and distributed forms that are compatible with computing technologies, including cloud computing systems. A computing system may be part of, or accessible by, a LAN, mobile communication network, satellite communication network, the Internet, a public or private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of the invention.

[0156] It should be noted that some of the system functions described herein are presented as modules to more clearly emphasize the independence of their implementations. For example, modules may be implemented as hardware circuits including custom very large-scale integrated (VLSI) circuits or gate arrays, or commercially available semiconductors (e.g., logic chips, transistors, or other discrete components). Modules may also be implemented in programmable hardware devices such as field-programmable gate arrays, programmable array logic, programmable logic devices, and graphics processing units.

[0157] Modules can also be implemented, at least partially, as software executed by various types of processors. A defined unit of executable code may include, for example, one or more physical or logical computer instruction blocks that can be organized as objects, procedures, or functions. However, the executable code of a defined module does not need to be physically located in one place; it may include distributed instructions stored in different locations, which, when logically combined, constitute a module and achieve its intended purpose. Furthermore, modules can be stored, for example, on hard disk drives, flash devices, RAM, tape, and / or any other non-volatile computer-readable medium.

[0158] In fact, a module of executable code may be a single instruction, a series of instructions, or distributed across multiple different code segments, different programs, and multiple memory devices. Similarly, operational data may be identified and illustrated within a module herein, embodied in any suitable form, and organized within any suitable type of data structure. Operational data may be collected as a single data set, or distributed across different locations in different storage devices, or at least partially exist as electronic signals on a system or network.

[0159] Without departing from the scope of the present invention, various types of AI / ML models can be trained and deployed. For example, Figure 12A shows an example of a neural network 1200 trained to receive inputs (indicated in column 1210) from input "neurons" 1 to I in an input layer (indicated in column 1220), according to one or more embodiments. The neural network 1200 includes several hidden layers (indicated in columns 1230 and 1240). Both DLNNs and shallow learning neural networks (SLNNs) typically have multiple layers, although SLNNs may have only one or two layers, and usually fewer than DLNNs. Typically, a neural network architecture includes an input layer, several intermediate layers (e.g., hidden layers), and an output layer (indicated in column 1250), and the neural network 1200 is similar.

[0160] DLNNs often have a large number of layers (e.g., 10, 50, 200, etc.), and subsequent layers typically reuse features from previous layers to compute more complex and general functions. SLNNs, on the other hand, usually have fewer layers and are trained relatively quickly because expert features are generated beforehand from raw data samples. However, feature extraction is labor-intensive. In contrast, DLNNs usually do not require expert features, but training takes longer and they tend to have more layers.

[0161] In both approaches, layers are typically trained concurrently on the training set and overfitting is usually checked on separate cross-validation sets. Both techniques can yield excellent results, and there is considerable interest in both approaches. The optimal size, shape, and number of individual layers vary depending on the problem the neural network addresses.

[0162] Returning to Figure 12A, the input layer provides company-specific information, including data from diverse industries and applications, unique industry terminology, complex document structures, and pre-trained knowledge, which is then fed as input to J neurons in Hidden Layer 1. Other possible inputs include, but are not limited to, computing system status information, published automations, business rules, information on related workflows and / or tasks, initial automation definitions, and process automation documents. In this example, all of these inputs are supplied to each neuron, but without departing from the scope of the present invention, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks may be used individually or in combination.

[0163] Hidden layer 2 (1240) receives input from hidden layer 1 (1230), hidden layer 3 receives input from hidden layer 2 (1240), and so on until the last hidden layer (indicated by the ellipsis 1255), which provides its output as input to the output layer. Although multiple proposals are shown as outputs here, in some embodiments only a single output proposal is provided. In some embodiments, proposals are ranked based on confidence scores. In this embodiment, the output is an accurate response to a professional or up-to-date query.

[0164] It should be noted that the number of neurons I, J, K, and L are not necessarily equal. Therefore, any number of layers can be used for a given layer of the neural network 1200 without departing from the scope of the present invention. In fact, in some embodiments, the types of neurons within a given layer do not all need to be the same.

[0165] The neural network 1200 is trained to assign confidence scores to appropriate outputs. To reduce inaccurate predictions, in some embodiments, only results whose confidence scores meet or exceed a predetermined confidence threshold are provided. For example, if the confidence threshold is 80%, outputs with confidence scores exceeding this value may be used, and the rest may be ignored. According to one or more embodiments, the output layer 1250 indicates that two text fields (shown in outputs 1261 and 1262), one text label (shown in output 1263), and a submit button (shown in output 1265) have been detected. Without departing from the scope of one or more embodiments herein, the neural network 1200 may provide the location, dimensions, images, and / or confidence scores of these elements, and such outputs may be subsequently used by an RPA robot or other automation for a predetermined purpose.

[0166] Neural networks are typically probabilistic configurations that have confidence scores. These can be scores trained by an AI / ML model based on how correctly it identified similar inputs during training. Common types of confidence scores include decimals between 0 and 1 (which can be interpreted as confidence percentages), numerical values ​​ranging from negative infinity to positive infinity, and a set of representations (e.g., "low," "medium," "high"). Various post-processing calibration techniques, such as temperature scaling, batch normalization, weight decay, and negative log-likelihood (NLL), can be used to obtain more accurate confidence scores.

[0167] In neural networks, "neurons" are algorithmically implemented as mathematical functions, often based on the functions of biological neurons. A neuron receives weighted inputs and possesses a summation function and an activation function, the activation function controlling whether or not the output is passed to the next layer. This activation function can be a nonlinear threshold function (i.e., modified linear unit (ReLU) nonlinearity) that does nothing when the value is below a threshold and responds linearly when it exceeds the threshold. The summation function and ReLU function are used in deep learning because real neurons can generally have similar activity functions. Information can be subtracted, added, etc., through linear transformations. Essentially, neurons act as gating functions controlled by their underlying mathematical functions, passing the output to the next layer. In some embodiments, different functions may be used for at least some of the neurons.

[0168] An example of neuron 1295 is shown in Figure 12B. Inputs from the preceding layer x1, x2, ..., x n Each of these has a weight w1, w2, ..., w n This is assigned. Therefore, the set input from preceding neuron 1 is w1x1. These weighted inputs are used in the sum function of the neuron in question, which is modified by a bias as shown in the following equation.

number

[0169] This sum is compared to the activation function f(x) to determine whether or not a neuron "fires". For example, f(x) can be given by the following equation:

number

[0170] Therefore, the output y of neuron 1295 can be given by the following equation.

number

[0171] In this case, neuron 1295 is a single-layer perceptron. However, any suitable neuron type or combination of neuron types may be used without departing from the scope of the present invention. It should also be noted that in some embodiments, the range of weight values ​​and / or the range of activation function output values ​​may differ without departing from the scope of the present invention.

[0172] Often, a goal, or "reward function," is used. The reward function guides the exploration of the state space, seeking intermediate transitions and steps with both short-term and long-term rewards in an attempt to achieve the goal (e.g., finding the most accurate answer to a user query based on relevant metrics). During training, various labeled data are fed into the neural network 1200. If the classification is successful, the weight of the input to the neuron is strengthened; if the classification fails, the weight is weakened. Cost functions such as mean squared error (MSE) or gradient descent may be used, and large errors may be punished more severely than small errors. If the performance of the AI / ML model does not improve after a predetermined number of training iterations, the data scientist may modify the reward function and provide corrections for incorrect predictions, etc.

[0173] Backpropagation is a technique for optimizing synaptic weights in feedforward neural networks. Backpropagation "interiorally examines" the hidden layers of the neural network to understand how much each node contributes to the loss, and then can minimize the loss by updating the weights so that nodes with high error rates have lower weights and nodes with low error rates have higher weights. In other words, backpropagation allows data scientists to iteratively adjust weights to minimize the difference between the actual output and the desired output.

[0174] The backpropagation algorithm has a mathematical basis in optimization theory. In supervised learning, training data with known outputs is passed through a neural network, and an error is calculated from the known target output by a cost function, which becomes the error for backpropagation. The error is calculated at the output, and this error is transformed into a correction of the network weights to minimize the error.

[0175] In the case of supervised learning, an example of backpropagation is shown below. The column vector input x is processed through a series of N non-linear activation functions f i between each layer i = 1, …, N of the network, and the output at a given layer is first multiplied by the synaptic matrix W i and then the bias vector b i is added. The network output o given by the following equation is,

Equation

[0176] In some embodiments, o is compared with the target output t, and as a result, an error

Equation

[0177] Optimization in the form of a gradient descent procedure can be used to minimize the error by modifying the synaptic weights W i of each layer. The gradient descent procedure requires calculating the output o when the input x corresponding to the known target output t is given and generating the error o - t. This global error is then propagated backward and gives a local error for weight update by calculations similar but not identical to those used in forward propagation. In particular, the backpropagation process typically requires an activation function in the form of p j (n j ) = f j ’(n j ), where n j is the network activity at layer j (i.e., nj =W j o j-1 +b j ) and o j =f j (n j ) and the apostrophe ' indicates the derivative of the activity function f.

[0178] Weight updates can be calculated using the following formula:

number

[0179] Here, o represents the Hadamard product (i.e., the element-wise product of two vectors), T This shows the matrix transpose, o j is f j (W j o j-1 +b j ) shows that o0 = x. Here, the learning rate η is selected based on machine learning considerations. Below, η is associated with the Hebb learning mechanism used in neural implementations. Note that synapses W and b can be joined to one large synaptic matrix, in which case it is assumed that 1 is added to the input vector and an additional column representing the b synapse is included in W.

[0180] The AI / ML model can be trained over multiple epochs until a good level of accuracy is achieved (e.g., 97% or higher using an F2 or F4 threshold for detection, and over approximately 2,000 epochs). In some embodiments, this level of accuracy can be determined using an F1 score, F2 score, F4 score, or any other suitable method that does not deviate from the scope of the invention. After being trained on training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This ensures that the AI / ML model is not "overfitted," i.e., it performs well on training data but not on other data.

[0181] In some embodiments, the level of accuracy that an AI / ML model can achieve may be unknown. Therefore, if the accuracy of the AI / ML model begins to decline when analyzing evaluation data (i.e., it performs well on training data but its performance begins to decline on evaluation data), the AI / ML model may undergo further training for more epochs on the training data (and / or new training data). In some embodiments, an AI / ML model is deployed only when its accuracy reaches a predetermined level, or when the accuracy of the trained AI / ML model is better than that of an existing deployed AI / ML model. In some embodiments, a set of trained AI / ML models may be used to accomplish a task. For example, one AI / ML model may be trained for image recognition, another for text recognition, and yet another for recognizing semantic and / or ontology associations.

[0182] It should be noted that, in addition to or instead of neural networks, transformer networks such as SentenceTransformers® may be used in some embodiments. SentenceTransformers® is a state-of-the-art Python® framework for embedding sentences, text, and images. Such transformer networks learn the relevance of words and phrases that have both high and low scores. This trains the AI / ML model to distinguish between those that are close to the input and those that are not. Rather than simply using word / phrase pairs, transformer networks may also use field length and field type.

[0183] As described above, NLP models such as word2vec, BERT, GPT-3, ChatGPT, and other LLMs can be used in several embodiments to facilitate semantic understanding and provide more accurate and human-like responses. Other techniques, such as clustering algorithms, can be used to find similarities between groups of elements. Clustering algorithms may include, but are not limited to, density-based, distribution-based, centroid-based, and hierarchy-based algorithms. Examples include K-means clustering algorithms, DBSCAN clustering algorithms, Gaussian mixture model (GMM) algorithms, and BIRCH (Balance Iterative Reducing and Clustering using Hierarchies) algorithms. Such techniques can also contribute to categorization.

[0184] Figure 13 is an architecture diagram showing a reference architecture 1300 for a generative AI model in one or more embodiments. This architecture consists of multiple layers: an API plugin, a prompt library, vector data source ingestion, access processing control, a model training pipeline, an evaluation layer for evaluating hallucinations / telemetry / evaluation, a BYOM embedding layer, and an LLM orchestration layer. There are also search plugins, access control plugins, and API plugins that are integrated into enterprise systems.

[0185] Reference Architecture 1300 has three main flows.

[0186] Flow 1: Ingestion may include data ingestion and training flows. For example, data is read from multiple data stores, preprocessed, chunked, and trained through an embedding model (e.g., RAG and training pipeline (i.e., fine-tuning)). A vector database stores the chunked document embeddings (e.g., vector embeddings), enabling better semantic and similarity-based data retrieval.

[0187] Flow 2: Search may include prompt extension using a data retrieval flow. For example, when a user query reaches the API layer, a prompt is selected, and then, before the prompt is passed to the LLM layer, a data retrieval is performed via a vector database or API plugin to retrieve the appropriate contextual data.

[0188] Flow 3: Inference may include the LLM inference flow. For example, here a choice is made whether to use a general-purpose foundational model or a self-hosted foundational model. Fine-tuned models may be used if they are tailored to a specific task or use case. The response is evaluated for accuracy and other metrics (including hallucinations).

[0189] It should be noted that in some embodiments, generative AI models having multiple "heads" may be used. A head refers to the output layer of the generative AI model. Generative AI models such as generative AI model 172 in Figure 1 typically have a sequence of multiple layers, and each head often shares the first few layers of the model before branching off into its own layers.

[0190] Figure 14 is a flowchart illustrating a process 1400 for training an AI / ML model according to one or more embodiments. In some embodiments, as described above, the AI / ML model may be a generative AI model. In the case of a neural network, the architecture typically includes multiple neuron layers, including an input layer, an output layer, and hidden layers (see, for example, Figures 12A and 12B). The hidden layers between these generate an intermediate representation of the input, which is used to process the input data and produce the output. These hidden layers may include various types of neurons, such as convolutional neurons, recurrent neurons, and / or transformer neurons. A generative AI model may also have various layers.

[0191] The process 1400 performed in Figure 14 may be executed by automation as described herein, implemented by a computer program according to one or more embodiments. The computer program may be implemented on a non-temporary computer-readable medium. The computer-readable medium may include, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other medium or combination of mediums used to store data. The computer program may include coded instructions for controlling a processor of a computing system (see, for example, Figure 11) to implement all or part of the process 1400 described in Figure 14, which may also be stored on the computer-readable medium.

[0192] In some embodiments, the training process begins in block 1410 by providing data (labeled or unlabeled) such as enterprise-specific information, along with data from diverse industries and applications, specific industry terminology, complex document structures, and pre-trained knowledge. For generative AI models, which are generally trained, the training process may be omitted unless a fine-tuned model is desired, as will be discussed in more detail below. The AI / ML model is then trained over multiple epochs in block 1420, and the results are reviewed in block 1430. While various types of AI / ML models can be used, LLMs and other generative AI models are typically trained (fine-tuned) using a process called "supervised learning," which was also discussed above. Supervised learning involves providing the model with a large dataset that the model will use to learn the relationship between inputs and outputs. During the training process, the model adjusts the weights and biases of neurons in the neural network to minimize the difference between the predicted output and the actual output in the training dataset.

[0193] One aspect of the model in some embodiments is the use of transfer learning. For example, transfer learning can utilize a pre-trained model such as ChatGPT, which is fine-tuned for a specific task or domain in block 1420. This allows the model to leverage knowledge already trained in the pre-training phase and adapt to a specific application through the training phase in block 1420.

[0194] The pre-training phase involves training the model on an initial training dataset that may be more general. In this phase, the model trains on relationships within the data. In the fine-tuning phase (which in some embodiments is performed in addition to or instead of the initial training phase in block 1420, for example, when the pre-trained model is used as the initial foundation for the final model), the pre-trained model is adapted to a specific task or domain by training the model on a smaller, task-specific dataset. For example, in some embodiments, the model may focus on a particular type of data source. This can help the model identify data elements within that data source more accurately than if the generative AI model were only pre-trained. Fine-tuning allows for training on the nuances of the source, such as specific vocabulary and syntax, specific graphic properties, or specific data formats, without requiring as much data as would be needed to train the model from scratch. By leveraging the knowledge trained in the pre-training phase, the fine-tuned model can achieve state-of-the-art performance on a particular task with relatively little additional training data.

[0195] In some embodiments, if the AI / ML model does not meet the desired confidence threshold in decision block 1440, process 1400 proceeds to block 1450 (indicated by the NO arrow). The training data is augmented and / or the reward function is modified to help the AI / ML model achieve the objective better (block 1450), and process 1400 returns to block 1420.

[0196] If the AI / ML model meets the confidence threshold in decision block 1440, process 1400 proceeds to block 1460 (indicated by a YES arrow). In block 1460, the AI / ML model is tested on evaluation data to ensure that the AI / ML model generalizes well and that it is not overtrained with respect to the training data. The evaluation data includes information that the AI / ML model has not previously processed.

[0197] If the confidence threshold is met in decision block 1470 for the evaluation data, process 1400 proceeds to block 1480 (indicated by the YES arrow). The AI / ML model is deployed in 1480. If the threshold is not met, process 1400 returns to block 1450 (indicated by the NO arrow), and the AI / ML model is further trained.

[0198] Figure 15 is a flowchart of process 1500 relating to advanced agent-based extraction and retrieval for context grounding within automation, according to one or more embodiments. Process 1500 performed in Figure 15 may be executed by automation described herein (e.g., AI agent, AOP, or RPA robot) implemented by a computer program according to one or more embodiments. The computer program may be embodied on a non-temporary computer-readable medium. The computer-readable medium may include, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other medium or combination of media used to store data. The computer program may also be stored on a computer-readable medium and may include coded instructions for controlling a processor of a computing system (see, for example, Figure 11) to implement all or part of process 1500 as described in Figure 15.

[0199] Process 1500 generally provides a context-grounding technique that improves AI models by integrating enterprise-specific information with pre-trained knowledge, enabling accurate responses to specialized or recent queries. Process 1500 includes extraction of block 1520, semantic storage of block 1540, and search / calculation of block 1530, which are described, for example, as implemented as agent-based automation performed by one or more AI agents. Some or more technical effects, benefits, and advantages of Process 1500 include, but are not limited to, one or more AI agents acting independently, adaptively, and dynamically to make decisions and take action to solve the key challenge of ensuring effective search and semantic matching for specialized or recent queries. Thus, one or more AI agents can decipher specific industry terminology and complex document structures, provide precise chunking of documents, and extract and retrieve relevant information within them while reducing processing and memory resources.

[0200] Starting from block 1520, one or more AI agents perform extraction operations. In a general sense, extraction operations involve taking data from one or more sources, transferring it, and converting it to text.

[0201] According to one or more embodiments, one or more AI agents access documents (including complex document structures) in different formats and located in different places (subblock 1572). Examples of documents with complex document structures include, but are not limited to, images and PDFs containing complex tables. Access may include retrieving relevant data chunks from the document repository and integrating such relevant data chunks as text into prompts for the AI ​​model in order to improve response accuracy.

[0202] One or more AI agents extract data from accessed documents (subblock 1574). Extraction may include parsing each document and converting the data into text. One or more AI agents may implement one or more different methods to parse different documents. Parsing methods may include, but are not limited to, expression, rule-based parsing, tokenization, named entity recognition (NER), part-of-speech (POS) tagging, semantic analysis, syntactic parsing, machine learning-based parsing, and other specialized document parsing techniques (table extraction and form recognition). One or more AI agents may select one or more parsing methods based on the document and the desired level of information extraction. One or more AI agents convert the data from accessed documents into text so that the text is used by the AI ​​model to provide contextual grounding. As an example, the data from accessed documents is flattened into a storable and searchable representation. More specifically, for example, a PDF contains images and complex tables, and one or more AI agents flatten them into a text representation to be used in subsequent agent-based automation parts.

[0203] In block 1540, one or more AI agents perform semantic storage operations. The semantic storage of the text extracted by one or more AI agents in block 1520 optimizes storage resources (e.g., databases) through intelligent allocation and management to improve performance. Furthermore, the well-optimized semantic storage by one or more AI agents enables essential semantic searching and retrieval of the text.

[0204] According to one or more embodiments, the intelligent assignment and management of one or more AI agents includes generating vectors for storing text (subblock 1583). The vectors enable the use of content embeddings (e.g., vector embeddings) for different phrases or words of text. Furthermore, one or more AI agents generate content embeddings (possibly using models or transformers) (subblock 1585). Content embeddings may take the form of long vectors of floating-point numbers generated by passing the text through a model or transformer. By using vectors, one or more AI agents can quickly access the text using content embeddings, thereby saving processing time and improving processing accuracy. One or more AI agents store vectors containing text and content embeddings (subblock 1587).

[0205] In block 1560, one or more AI agents perform search / calculation operations. These search / calculation operations by one or more AI agents include, as described herein, advanced agent-based searches such as semantic search and retrieve, hybrid search, and wide-area search with re-ranking. It should be noted that semantically storing vectors, including text and content embeddings, presents computational problems because conventional software automation cannot adapt to the complexity of vectors. For example, conventional software automation, being static, cannot search vectors at runtime without requiring extreme processing power and time (making conventional software automation impractical) to compute all floating-point numbers.

[0206] One or more AI agents solve this computational problem by performing an embedded search at runtime (using advanced agent-based search) by calculating a vector from specialized or recent queries. Thus, advanced agent-based search retrieves information grounded in the context of specialized or recent queries.

[0207] Figure 16 is a flowchart of process 1600 relating to advanced agent-based extraction and retrieval for context grounding within automation, according to one or more embodiments. Process 1600 performed in Figure 16 may be executed by automation described herein (e.g., AI agent, AOP, or RPA robot) implemented by a computer program according to one or more embodiments. The computer program may be embodied on a non-temporary computer-readable medium. The computer-readable medium may include, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other medium or combination of media used to store data. The computer program may also be stored on a computer-readable medium and may include coded instructions for controlling a processor of a computing system (see, for example, Figure 11) to implement all or part of process 1600 as described in Figure 16.

[0208] According to one or more embodiments, the advanced agent-based extract and search process 1600 provides context grounding as a mechanism to solve the problem where an AI model, i.e., an LLM (e.g., an NLP model such as word2vec, BERT, GPT-3, ChatGPT), is questioned about recent events or data (whether private chats, public chats, or a combination thereof), and the LLM responds, "Sorry, I don't know about that." This problem can arise because the LLM's training has been discontinued at some point. Context grounding is a way to supplement the information the LLM knows from pre-training with private information accumulated within the enterprise. Furthermore, the advanced agent-based extract and search process 1600 for context grounding within automation can be executed "on the fly." That is, the advanced agent-based extract and search process 1600 is a way to retrieve chunks from documents and add those chunks to prompts sent to the LLM's chat.

[0209] According to one or more embodiments, the agent-based extraction and search process 1600 is implemented by one or more AI agents and generates and provides context grounding for the AI ​​model. The advanced agent-based extraction and search process 1600 begins in block 1605 with one or more AI agents receiving a complex query (e.g., a specialized or recent query). The complex query requires information that is inaccessible to the LLM. That is, the complex query may include complex syntax and multiple parts as well as parameters. The complex query requires extensive and sophisticated logic for data transformation, handling of subqueries, context analysis, filtering, and accurate response generation.

[0210] One or more AI agents can receive complex queries via prompts. For example, complex queries are received via LLM's chat prompts (whether private, public, or a combination thereof). Thus, complex queries ask LLM about recent events or data outside of LLM, triggering advanced agent-based extraction and search processes 1600.

[0211] In block 1610, one or more AI agents implement query decomposition. Query decomposition simplifies a specialized or recent query into subqueries (e.g., smaller parts) for better search and response aggregation. Query decomposition may include, but is not limited to, translating the complex syntax and multiple parts and parameters of a specialized or recent query into a relational algebra query, verifying the syntactic and semantic correctness of the relational algebra query, and splitting the relational algebra query into multiple distinct subqueries.

[0212] When a specialized or recent query is broken down into multiple distinct subqueries, the advanced agent-based extract and search process 1600 includes extract, semantic storage, and search / calculation operations. General examples of extract, semantic storage, and search / calculation operations are described herein, for example, with respect to Figure 15 and blocks 1520, 1540, and 1560, but the advanced agent-based extract and search process 1600 describes these operations in more detail.

[0213] In block 1615, one or more AI agents access a document repository. One or more AI agents access the document repository based on multiple separate subqueries. The document repository may contain one or more documents that provide data. The data may be provided as a complex document structure of the one or more documents. Documents with a complex document structure may include images or PDFs containing complex tables. The document repository may store at least two of the one or more documents in different locations that are inaccessible to the AI ​​model. The at least two of the one or more documents may be in different formats (e.g., PDFs, text files, and source code files for web pages).

[0214] In block 1620, one or more AI agents extract data from the one or more documents. One or more AI agents extract data from the one or more documents. One or more AI agents extract data from an accessed document repository based on a plurality of distinct subqueries. According to one or more embodiments, one or more AI agents extract data by parsing each document (subblock 1622) and converting the data to text (subblock 1624) as described herein. According to one or more embodiments, the text conversion may include generating one or more vectors and context embeddings for different phrases or words of the text. Content embeddings may include floating-point numbers generated by passing the text through a model or transformer. Content embeddings may be inserted into vectors.

[0215] In block 1630, one or more AI agents implement the semantic storage of the one or more vectors. The semantic storage includes the results of the extraction. The semantic storage is then used for advanced agent-type search, providing context grounding to the LLM. According to one or more embodiments, one or more AI agents semantically store text in agent-type memory as vectors including context embeddings.

[0216] In block 1640, one or more AI agents implement advanced agent-based search. Advanced agent-based search searches agent-based memory to generate context grounding. Thus, advanced agent-based search retrieves information grounded in the context of complex queries.

[0217] It should be noted that semantically storing vectors containing text and content embedding presents computational problems for conventional software automation, as it cannot adapt to the complexity of vectors. For example, because conventional software automation is static, it cannot search for vectors at runtime without requiring extreme processing power and time (making conventional software automation impractical) to compute all floating-point numbers.

[0218] One or more AI agents solve this computational problem by performing embedded searches at runtime (using advanced agent-based search) by calculating vectors from specialized or recent queries. One or more AI agents use contextual embedding to find vectors with similar meanings to specialized or recent queries, going beyond exact matches or keywords.

[0219] As described herein, advanced agent-based search may include, but is not limited to, semantic search and search retrieval 1643, hybrid search 1645, and wide-area search 1647 with re-ranking.

[0220] Semantic search and retrieve 1643 matches queries to embeddings and obtains contextual relevance. According to one or more embodiments, semantic search and retrieve 1643 includes, but is not limited to, accessing and recalling text in storage resources based on meaning and relation (rather than mere keywords in conventional software automation). For example, one or more AI agents may determine how close their vector is to other vectors in the storage resources, and this is how one or more AI agents perform semantic search and retrieve 1643.

[0221] Hybrid search 1645 combines semantic search and keyword search to improve results. As another example, one or more AI agents perform hybrid search 1645, which implements vector computation and keyword search of semantic search and search retrieval 1643, in order to complement vector computation with keyword search.

[0222] Broad search with re-ranking 1647 performs a broad search that is refined by LLM for accuracy. As another example, one or more AI agents perform a broad search with LLM re-ranking (for example, when a very large number of search results are requested and the broad search returns all relevant but scattered, one or more AI agents can achieve accuracy by having LLM re-rank all relevant so that the most relevant vector comes first).

[0223] In block 1660, one or more AI agents output context grounding. Context grounding may be output to an LLM (e.g., an AI model). According to one or more embodiments, one or more AI agents provide context grounding to the LLM along with a complex query. According to one or more embodiments, one or more AI agents provide context grounding within a prompt containing a complex query. Context grounding improves the accuracy of the LLM's response to the complex query.

[0224] In block 1670, a response to a complex query is received. The response is output by the LLM as a reply to the context query. The response may be provided within a prompt. For example, the response may be received via the LLM's chat prompt, and one or more AI agents may obtain the response from the chat prompt. The response may also be received directly from the LLM to one or more AI agents.

[0225] In block 1680, one or more AI agents implement tethering. Generally, the advanced agent-based extraction and search process 1600 functions well enough, but there may still be some items that one or more AI agents may err on. Tethering proposes intervening with the user and providing dynamic and / or direct user input.

[0226] According to one or more embodiments, tethering may include connecting one or more specialized or recent queries and / or subqueries (e.g., any complex queries) to a context for advanced response generation. Tethering may include a human-in-the-loop operation that specifies a desired response for a particular query, improving accuracy for similar queries in the future. Tethering may include automated tethering, in which one or more AI agents perform targeted searches on one or more specialized or recent queries and / or subqueries, aggregating the results to improve response quality.

[0227] According to one or more embodiments, tethering may include automated tethering. Automated tethering may include performing targeted searches, aggregating results, and further breaking down one or more specialized or recent queries and / or subqueries into simpler components in order to improve context grounding and, consequently, response quality. Automated tethering leverages the fact that good extract results are obtained for a particular query and performance costs for similar queries in the future can be avoided (for example, automated tethering saves costs by storing / caching the results of very expensive searches in agent memory so that those results can be retrieved next time). As an example, agent memory automatically tethers context grounding and agent memory based on instructions from one or more AI agents.

[0228] The advanced agent-based extract and search process 1600 generally provides a context-grounding technique that improves AI models by integrating enterprise-specific information with pre-trained knowledge, enabling accurate responses to specialized or recent queries. Some of the technical effects, benefits, and advantages of the advanced agent-based extract and search process 1600 include, but are not limited to, one or more AI agents acting independently, adaptively, and dynamically to make decisions and take action to address the key challenge of ensuring effective search and semantic matching for specialized or recent queries. Thus, one or more AI agents can decipher specific industry terminology and complex document structures, provide precise chunking of documents, and extract and retrieve relevant information from them while reducing processing and memory resources.

[0229] Figure 17 is an architectural diagram showing agent-type memory within a hyper-automation system 1700 in one or more embodiments. In the hyper-automation system 1700, each computing system 1702, 1704, and 1706 has automations 1710, 1712, and 1714 that are implemented therein, such as RPA robots, AI agents, AOPs, etc. In some embodiments, automations 1710, 1712, and 1714 may be stored in a remote database 1740 via a network 1720 and can access one or more AI models 1772 via a cloud environment 1770. Automations 1710, 1712, and 1714 can further access agent-type memory 1790.

[0230] According to one or more embodiments, the agent-type memory 1790 can provide a dynamic caching (i.e., storage) system for automations 1710, 1712, and 1714. The agent-type memory 1790 can provide a dynamic caching (i.e., storage) system for automations 1710, 1712, and 1714 by being used to cache context grounding, escalation, tool calls, dynamic and / or direct user input generated by enhanced extraction and retrieval techniques, as well as other results generated by automations 1710, 1712, and 1714 and one or more AI models 1772. In this way, the agent-type memory 1790 solves the shortcomings of one or more AI models 1772 by providing an improved alternative storage approach for semantically mapping context grounding and other results for use by automations 1710, 1712, and 1714 throughout the hyper-automation system 1700. In some cases, agent-type memory 1790 functions as long-term memory for automation 1710, 1712, and 1714.

[0231] For example, if automation 1710 is running and encounters a problem that cannot proceed without input, automation 1710 performs an escalation for human-in-the-loop operation. After automation 1710 receives dynamic and / or direct user input, automation 1710 can store the escalation, user input, and solution in agent memory 1790. Therefore, if automation 1712 encounters the same or similar problem while running, automation 1712 semantically searches agent memory 1790 to find a solution, and human-in-the-loop operation is never performed by automation 1712. The technical effects, advantages, and superiorities of agent memory 1790 are improvements in latency and processing, as well as overall performance. According to one or more embodiments, one or more AI models 1772 can be trained with data from agent memory 1790 over time. However, even if one or more AI models 1772 are trained every 3(3) months or 6(6) months, the agent-type memory 1790 can always provide support during the training period. Over time, the agent-type memory 1790 reduces the number of escalations and the amount of user work. In contrast, traditional software automation increases the user's workload over time, subsequently losing efficiency and increasing costs due to the need for more processing power and processing time.

[0232] Figure 18 is a flowchart of a process 1800 for agent-based extraction according to one or more embodiments. The process 1800 performed in Figure 18 may be executed by automation as described herein (e.g., an AI agent, AOP, or RPA robot) implemented by a computer program according to one or more embodiments. The computer program may be embodied on a non-temporary computer-readable medium. The computer-readable medium may include, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other medium or combination of media used to store data. The computer program may also be stored on a computer-readable medium and may include coded instructions for controlling a processor of a computing system (see, for example, Figure 11) to implement all or part of the process 1800 described in Figure 18.

[0233] Process 1800 is implemented by one or more AI agents (e.g., AI agent 1801) that utilize one or more AI models 1802, 1803, and 1804. According to one or more embodiments, process 1800 for agent-based extraction is an example of extraction in block 1520 of Figure 15. Process 1800 is implemented by one or more AI agents 1801 to generate and provide context grounding. The agent-based extraction method is triggered in response to receiving a complex query that requires information inaccessible to the AI ​​models.

[0234] Process 1800 begins in block 1805 with the AI ​​agent 1801 receiving one or more documents. The one or more documents include a complex document structure. For example, the one or more documents include a PDF containing a complex document structure (e.g., images, timelines, callouts, and / or complex tables). According to one or more embodiments, at least two of the one or more documents are in different formats and are located in different locations that are inaccessible to one or more AI models 1802, 1803, and 1804.

[0235] In block 1810, the AI ​​agent 1801 extracts text data from the one or more documents. In block 1815, the AI ​​agent 1801 captures one or more images from the one or more documents.

[0236] In block 1819, AI agent 1801 performs multi-stage processing of text data and one or more images. The multi-stage processing utilizes one or more AI models 1802, 1803, and 1804, which may include one or more large-scale language models. One or more large-scale language models may include at least one multimodal large-scale language model (MLLM). An MLLM is an AI model capable of processing and generating data from multiple modalities such as text, images, audio, or video. According to one or more embodiments, AI model 1802 is a first MLLM, AI model 1803 is a second MLLM, and AI model 1804 is an LLM. The multi-stage processing generates an output set, which may include contextualization and one or more keywords.

[0237] At arrow 1821, AI agent 1801 provides / feeds text data and one or more images to the first MLLM 1802. AI agent 1801 can provide / feed text data and one or more images to the first MLLM 1802 for reformatting based on three or more draft output requests. The first MLLM 1802 generates one or more draft outputs from the text data and one or more images. In this regard, the first MLLM 1802 processes the text data and one or more images together to understand not only text alone, but also text with orientation within one or more documents, and generates one or more draft outputs. The one or more draft outputs are provided to AI agent 1901. That is, as shown by arrow 1822, AI agent 1801 receives the one or more draft outputs from the first MLLM 1802.

[0238] In decision block 1830, the AI ​​agent 1801 compares each of the one or more draft outputs with a text threshold. The text threshold may be a set value. For example, the set value may be selected from the range of 50% to 100%. In one example, the set value may be 90%. Furthermore, the comparison by the AI ​​agent 1801 includes determining whether each draft output contains at least 90% of the text provided to the first MLLM 1802 at arrow 1821 (e.g., whether it retains at least 90% of the words). If the text threshold is not met for a particular draft output, the AI ​​agent 1801 may involve the first MLLM 1802 again, resulting in reprocessing indicated by the NO arrow, which leads back to block 1820 via block 1835. In block 1835, the AI ​​agent 1801 discards the particular draft output that does not meet the text threshold. If the draft output meets the text threshold, AI agent 1801 accumulates the draft output for further checks, indicated by a YES arrow to proceed to decision block 1840.

[0239] In decision block 1840, AI agent 1801 determines whether at least two draft outputs exist (for example, whether the first MLLM 1802 provided them and AI agent 1801 received two or more draft outputs). If at least two draft outputs do not exist, as indicated by the NO arrow returning from decision block 1840 to block 1820, AI agent 1801 re-involves the first MLLM 1802 to further process the text and one or more images. If at least two draft outputs exist, as indicated by the YES arrow proceeding from decision block 1840 to block 1860, AI agent 1801 provides / feeds at least two draft outputs to the second MLLM 1803. AI agent 1801 may provide / feed at least two draft outputs to the second MLLM 1803 for decision.

[0240] In block 1860, the second MLLM 1803 selects the best draft output from the one or more draft outputs. The best draft output, compared to the remaining draft outputs, most closely matches or best describes the one or more documents. The best draft output is provided to the AI ​​agent 1901. That is, as indicated by arrow 1862, the AI ​​agent 1801 receives the best draft output selected from the one or more draft outputs from the second MLLM 1803.

[0241] In block 1870, AI agent 1801 extracts the final text from the best draft output. As indicated by arrow 1871, AI agent 1801 provides / feeds the final text to the third LLM 1804. By sending the final text to the third LLM 1804, the third LLM 1804 generates one or more contextualizations and / or one or more keywords from the final text in block 1880. As indicated by arrow 1881, AI agent 1801 receives the one or more contextualizations and / or the one or more keywords from the third LLM 1804.

[0242] In block 1890, AI agent 1801 generates an output set. AI agent 1801 can generate an output set by concatenating the final text from the best draft output, the one or more contextualizations, and the one or more keywords.

[0243] After the completion of process 1800 for agent-based extraction, the output set may be converted into one or more vectors containing context embedding to provide context grounding.

[0244] The computer program may be implemented as a hardware, software, or hybrid implementation. The computer program may consist of modules that communicate with each other in an operable manner and may be designed to pass information or instructions to a display. The computer program may be configured to run on a general-purpose computer, an ASIC, or other suitable device.

[0245] It will be readily apparent that the components of the various embodiments of the present invention, as generally described herein and shown in the drawings, can be arranged and designed in a wide variety of different configurations. Therefore, the detailed description of the embodiments of the present invention shown in the accompanying drawings is not intended to limit the claimed scope of the invention, but merely represents selected embodiments.

[0246] The features, structures, or characteristics of the present invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, any reference throughout this specification to “a particular embodiment,” “some embodiments,” or similar phrases means that the particular features, structures, or characteristics described in relation to that embodiment are included in at least one embodiment of the present invention. Therefore, any appearance of “a particular embodiment,” “some embodiments,” “other embodiments,” or similar phrases throughout this specification does not necessarily refer to the same group of embodiments, and the features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0247] Furthermore, it should be noted that throughout this specification, references to features, advantages, or similar terms do not imply that all features and advantages that may be realized by the present invention should be realized, or are realized, in any single embodiment of the invention. Rather, terms referring to features and advantages should be understood to mean that certain features, advantages, or characteristics described in relation to embodiments are included in at least one embodiment of the invention. Accordingly, discussions of features and advantages, etc., throughout this specification do not necessarily refer to the same embodiment.

[0248] Furthermore, the features, advantages, and characteristics of the present invention described herein can be combined in any suitable manner in one or more embodiments. Those skilled in the art will understand that the present invention can be implemented without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages that are not present in all embodiments of the present invention may be recognized in a particular embodiment.

[0249] Those skilled in the art will readily understand that the present invention discussed above may be implemented in a different sequence of steps and / or with hardware elements of a different configuration than those disclosed. Therefore, although the present invention has been described based on these preferred embodiments, those skilled in the art may see certain modifications, variations, and alternative configurations that remain within the spirit and scope of the invention. Accordingly, to define the scope of the present invention, refer to the appended claims.

Claims

1. An agent-type memory method implemented by one or more artificial intelligence (AI) agents, wherein the agent-type memory method is In agent-type memory, one or more vectors containing context embeddings generated by advanced agent-type extraction in response to one or more queries are stored, In order to determine context grounding, an advanced agent-based search is used to receive and match complex queries to the one or more vectors in the agent-based memory, In order to improve the accuracy of the AI ​​model's response, the context grounding is provided from the agent-type memory to the AI ​​model along with the complex query. Receiving a response to the complex query from the AI ​​model, Methods that include...

2. An agent-type memory method according to claim 1, wherein the agent-type memory evolves and stores user interactions, feedback, modifications, and solutions.

3. An agent-type memory method according to claim 1, wherein the agent-type memory enables the one or more AI agents to learn efficient solutions to the complex query by eliminating human-in-the-loop operations.

4. An agent-type memory method according to claim 1, wherein the AI ​​model includes a large-scale language model.

5. An agent-type memory method according to claim 1, wherein the agent-type memory stores one or more of the context grounding, escalation, tool invocation, and user input.

6. An agent-type memory method according to claim 1, wherein the agent-type memory includes semantic storage of one or more vectors including text and context embeddings.

7. An agent-type memory method according to claim 1, wherein the context grounding includes relevant information from proprietary industry terminology and complex document structures that are not available to the AI ​​model.

8. An agent-type memory method according to claim 1, wherein the advanced agent-type search is a method for retrieving information of one or more vectors grounded in the context of the complex query as the context grounding.

9. An agent-type search method according to claim 1, wherein the agent-type memory automatically tethers the context grounding and the agent-type memory.

10. An agent-type search method according to claim 1, wherein the agent-type memory stores contextualization for context grounding and one or more keywords by multi-stage processing of the one or more documents.

11. A computer program product that implements agent-type memory on a non-temporary medium, wherein the agent-type memory is accessible to one or more artificial intelligence (AI) agents executed by one or more processors, and the agent-type memory is In agent-type memory, one or more vectors containing context embeddings generated by advanced agent-type extraction in response to one or more queries are stored. To determine context grounding, access is provided to the one or more AI agents so that advanced agent-based search can match complex queries to the one or more vectors in the agent-based memory. In order to improve the accuracy of the AI ​​model's response, the context grounding is provided to the AI ​​model from the agent-type memory along with the complex query. The AI ​​model receives a response to the complex query. A computer program product configured in such a way.

12. A computer program product that implements the agent-type memory described in claim 11, wherein the agent-type memory evolves and stores user interactions, feedback, modifications, and solutions.

13. A computer program product implementing the agent-type memory described in claim 11, wherein the agent-type memory enables the one or more AI agents to learn efficient solutions to the complex queries by eliminating human-in-the-loop operations.

14. A computer program product that implements the agent-type memory described in claim 11, wherein the AI ​​model includes a large-scale language model.

15. A computer program product that implements the agent-type memory described in claim 11, wherein the agent-type memory stores one or more of the context grounding, escalation, tool calls, and user input.

16. A computer program product that implements the agent-type memory described in claim 11, wherein the agent-type memory includes semantic storage of one or more vectors including text and context embeddings.

17. A computer program product implementing the agent-type memory described in claim 11, wherein the context grounding includes relevant information from proprietary industry terminology and complex document structures that are not available to the AI ​​model.

18. A computer program product that implements the agent-type memory described in claim 11, wherein the advanced agent-type search retrieves the information of the one or more vectors grounded in the context of the complex query as the context grounding.

19. A computer program product that implements the agent-type memory described in claim 11, wherein the agent-type memory automatically performs context grounding and tethers the agent-type memory.

20. A computer program product that implements the agent-type memory described in claim 11, wherein the agent-type memory stores contextualization for context grounding and one or more keywords by multi-stage processing of the one or more documents.