Predictive healing artificial intelligence agents that proactively detect impending workflow errors, failures, and / or defects and automatically attempt to recover
By using predictive healing AI agents and leveraging generative AI models to monitor the object repository, workflow errors are detected and automatically repaired. This addresses the issue of poor interoperability among AI agents and improves the system's automated repair capabilities and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UIPATH INC
- Filing Date
- 2025-12-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing AI agents lack effective interoperability with other applications and software functionalities, and struggle to predict and automatically correct workflow errors, failures, and defects.
A predictive healing AI agent is adopted to monitor the object repository through a generative AI model, detect workflow changes and automatically repair them, use generative AI models to repair workflows, integrate them into existing systems, simulate repairs in a sandbox environment, and call in human interventions to solve problems when necessary.
It enables proactive detection and automatic repair of workflow errors, faults, and defects, reducing manual intervention and improving system stability and automation.
Smart Images

Figure CN122431922A_ABST
Abstract
Description
Technical Field
[0001] This invention relates generally to software automation, and more specifically to predictive healing artificial intelligence (AI) agents that proactively detect impending workflow errors, failures and / or defects and automatically attempt to recover. Background Technology
[0002] Existing AI agents possess memory, knowledge bases (e.g., repositories of policy and background knowledge, a set of data, documents, etc. provided by an organization for its specific business and needs), generative AI, and large language models (LLMs) capabilities, enabling natural language communication and decision-making based on historical information. For example, users can send natural language queries to AI agents through prompts, and the agents can provide responses based on their knowledge bases and LLMs. These AI agents can also possess human-in-the-loop capabilities, allowing them to pose queries to human users that they cannot effectively resolve. However, these AI agents are often independently developed, standalone software applications, lacking effective interoperability with other applications and software functionalities. Therefore, improved and / or alternative approaches can be beneficial. Summary of the Invention
[0003] Certain embodiments of the present invention can provide solutions and / or useful alternatives to problems and needs in the prior art that are not yet fully identified, understood, or resolved by current software automation technologies. For example, some embodiments of the present invention relate to predictive healing AI agents that proactively detect impending workflow errors, failures, and / or defects and automatically attempt to recover.
[0004] In one embodiment, one or more non-transitory computer-readable media store one or more computer programs. The one or more computer programs are configured to cause at least one processor to initiate an agent loop for healing an AI agent, the agent loop including monitoring an object repository using a generative AI model. The generative AI model is then fine-tuned using data stored in the object repository. The one or more computer programs are also configured to cause at least one processor to determine that one or more changes have occurred in the object repository that will affect the workflow of a robotic process automation (RPA) robot, another AI agent, or an agent orchestration process (AOP) using a generative AI model. The one or more computer programs are also configured to cause at least one processor to attempt to repair the workflow using the generative AI model.
[0005] In another embodiment, the computer-implemented method includes initiating an agent loop for a healing AI agent by a computing system. This agent loop includes monitoring an object repository using a generative AI model. The generative AI model is then fine-tuned using data stored in the object repository. The computer-implemented method also includes the healing AI agent determining that one or more changes have occurred in the object repository that will affect an RPA bot, another AI agent, or an AOP workflow using the generative AI model. The computer-implemented method further includes the healing AI agent attempting to repair the workflow using the generative AI model.
[0006] In another embodiment, one or more computing systems include a memory storing computer program instructions; and at least one processor configured to execute the stored computer program instructions. The computer program instructions are configured to cause the at least one processor to initiate an agent loop for healing an AI agent, the agent loop including monitoring an object repository using a generative AI model. The generative AI model is fine-tuned using data stored in the object repository. The computer program instructions are also configured to cause the at least one processor to determine that one or more changes have occurred in the object repository that will affect an RPA bot, another AI agent, or an AOP workflow using the generative AI model. The computer program instructions are further configured to cause the at least one processor to attempt to repair the workflow using the generative AI model. Furthermore, the computer program instructions are configured to cause the at least one processor, in response to a successful attempted workflow repair, to send the repaired workflow to an orchestrator application to deploy the repaired workflow to one or more computing systems on which an RPA bot, another AI agent, or an AOP engine executing the workflow is deployed. The computer program instructions are also configured to cause at least one processor to collect data relating to one or more changes in response to an unsuccessful attempt at workflow repair, and to send the collected data to one or more computing systems that are in the human loop and / or configured to retrain generative AI models. Attached Figure Description
[0007] To facilitate understanding of the advantages of certain embodiments of the present invention, the invention briefly described above will be described in more detail with reference to the specific embodiments shown in the accompanying drawings. While it should be understood that these drawings depict only typical embodiments of the invention and should not be considered as limiting its scope, the invention will be described and explained with additional features and details using the drawings, in which:
[0008] Figure 1 This is an architectural diagram illustrating a hyperautomation system configured to perform agent automation and orchestration according to an embodiment of the present invention.
[0009] Figure 2 The illustration shows some combined capabilities of artificial intelligence (AI) agents and robotic process automation (RPA) robots according to embodiments of the present invention.
[0010] Figure 3 The illustration shows an AI agent, an RPA robot, an agent orchestration process (AOP), and a pool of applications according to an embodiment of the present invention.
[0011] Figure 4A and Figure 4B The illustration shows an example AI agent service interface according to an embodiment of the present invention.
[0012] Figure 5 The illustration shows an example AOP development interface according to an embodiment of the present invention.
[0013] Figure 6 An example RPA development interface according to an embodiment of the present invention is illustrated.
[0014] Figure 7 The illustration depicts an end-to-end AI agent, RPA robot, and AOP development and deployment system according to an embodiment of the present invention.
[0015] Figure 8 This is an architectural diagram illustrating an agent automation and RPA system according to an embodiment of the present invention.
[0016] Figure 9 This is an illustration of the architecture of a deployed RPA system according to an embodiment of the present invention.
[0017] Figure 10 This is an architecture diagram illustrating the relationship between the designer, activities, and drivers according to an embodiment of the present invention.
[0018] Figure 11 This is an architectural diagram illustrating a computing system configured to perform preemptive repair using an AI agent according to an embodiment of the present invention.
[0019] Figure 12A An example of a neural network trained to supplement a healing AI agent according to an embodiment of the present invention is illustrated.
[0020] Figure 12B An example of a neuron according to an embodiment of the present invention is illustrated.
[0021] Figure 13 This is an architectural diagram illustrating a generative AI model reference architecture according to an embodiment of the present invention.
[0022] Figure 14This is a flowchart illustrating a process for training one or more AI / ML models according to an embodiment of the present invention.
[0023] Figure 15 This is a flowchart illustrating a process for healing an AI agent according to an embodiment of the present invention.
[0024] Figure 16 This is a flowchart illustrating a process according to an embodiment of the present invention, which is used to proactively detect impending workflow errors, faults and / or defects, and automatically attempt to recover by a healing AI agent.
[0025] Unless otherwise stated, similar reference numerals in all figures always denote corresponding features. Detailed Implementation
[0026] Some embodiments involve predictive healing AI agents that proactively detect impending workflow errors, failures, and / or defects and automatically attempt to recover. Predictive healing AI agents can repair workflows associated with agent orchestration processes (AOP), agent loops of other AI agents, and / or robotic process automation (RPA) workflows using RPA robots. The healing capabilities of the AI agent can be seamlessly integrated into existing systems and workflows with minimal disruption. Developers may build user interface (UI) automation workflows that work at design time but fail at runtime for various reasons. For example, automation may be deployed to a computing system with graphical features different from those presented when the automation was designed; software updates may change aspects of the UI; the computing system may be too busy to run the automation; connectivity issues may exist; third-party systems on which the automation depends may be down, and so on. According to some embodiments, the AI agent can detect these problems that will occur during the automation's runtime and attempt to automatically correct them.
[0027] AI agents can access an object store that groups UI elements on a screen (e.g., text fields, buttons, labels, menus, checkboxes, etc.) by application, application version, application screen, and set of UI elements. Workflows and their corresponding activities can also be stored in the object store. As used herein, a "screen" is an image of an application UI or a portion of an application UI at a given point in time. In this context, an "application" or a version of a given application can be a union of screens. In some embodiments, each UI element can be described by one or more UI descriptors, which can also be stored in the object store. UI elements, UI descriptors, applications, and application screens are all UI objects. In some embodiments, UI elements and screens can be further categorized into specific types of UI elements (e.g., buttons, checkboxes, text fields, etc.) and screens (e.g., top windows, modal windows, pop-ups, etc.).
[0028] To make UI objects reusable, they can be extracted to a UI object store that can be referenced by AI agents and / or RPA. For example, when selectors or other UI descriptors are modified due to a new version of the application, the object store can include the modified UI descriptors. AI agents and / or RPA using the object store can then leverage the modified version of the UI descriptors.
[0029] In some embodiments, a selector is a UI descriptor that can be used to detect UI elements. In some embodiments, a selector has the following structure: <node_1 / ><node_2 / > ...<node_N / >
[0030] The last node represents the GUI element of interest, and when it is necessary to identify the element, the preceding nodes represent some or all of the element's parent elements.<node_1> It is usually referred to as the root node and represents the top window of the application.
[0031] Each node may have one or more attributes that help correctly identify a specific level of the selected application. In some embodiments, each node has the following format: <ui_system attr_name_1='attr_value_1' ... attr_name_N='attr_value_N' / >
[0032] Each attribute can have a specified value, and attributes with constant values can be selected. This is because changes to attribute values each time the application is started can cause selectors to fail to correctly identify associated elements.
[0033] UI descriptors can be directly added to RPA workflow activities, saving developers the time required to create custom selectors for activities. A collection of UI descriptors corresponding to one or more screens from a particular version of the application can be stored. A UI descriptor is a set of instructions for finding UI elements. In some embodiments, UI descriptors are encapsulated data / structure formats that include UI element selectors (one or more), anchor selectors (one or more), computer vision (CV) descriptors (one or more), unified target descriptors (one or more), screen image captures (context), element image captures, other metadata (e.g., application and application version), semantic selectors that understand the context and can be semantically associated using an AI model trained to perform semantic matching (e.g., see U.S. Patent No. 12,124,806), combinations thereof, etc. The encapsulated data / structure can use any suitable UI descriptor for identifying UI elements on a screen without departing from the scope of the invention. The unified target descriptor links together multiple types of UI descriptors. The unified target descriptor can function like a finite state machine (FSM), where a first UI descriptor mechanism is applied in a first context, a second UI descriptor is applied in a second context, and so on.
[0034] In some embodiments, when creating new UI descriptors and / or modifying existing UI descriptors, a shareable, collaborative, and potentially open-source object repository can be established and added. In some embodiments, taxonomy and ontology can be used. Applications, versions, screens, UI elements, descriptors, etc., can be defined as categories, which are hierarchical structures of subcategories.
[0035] However, many real-world concepts are not easily categorized and organized. Instead, they can be closer to concepts in mathematical ontology. In ontology, the relationships between categories are not necessarily hierarchical. For example, the situation where a button on a screen, when clicked, takes the user to another screen cannot be easily captured by the category of that screen because the next screen is not in a hierarchical structure. When constructing graphs representing such situations, ontology can be used, which allows for the creation of interactions between UI elements on the same or different screens and provides more information about how UI elements relate to each other.
[0036] Consider the example of clicking the OK button to bring up an employee screen. An ontology structure allows the designer application to suggest employees to the user on the next screen. The ontology information about the relationships between these screens, conveyed via the OK button, allows the designer application to do so. More complex and richer relationships can be captured by defining a graphical structure that isn't necessarily a tree, but rather related to what the application is actually doing.
[0037] An object store can include natural language descriptions of workflows, activities, and / or other objects stored within it. This helps generative AI models used by AI agents understand what the objects in the object store are, what tasks the workflow intends to accomplish, what a given action does, and what version of software (one or more) the automation is designed to use. This is particularly beneficial for graphical elements identified by semantic selectors, which use generative AI to understand the context in which the described graphical elements are used. By fine-tuning generative AI models using object stores and generating vector embeddings, pre-trained generative AI models, such as GPT-4, Transformers Bidirectional Encoder Representation (BERT), and LaMDA (LaDevice Language Model), can learn to deliver better results for the desired tasks—in this case, proactively detecting and fixing problems with workflows deployed in the runtime environment.
[0038] To fix workflows before errors, malfunctions, or defects occur, AI agents can periodically analyze deployed automated workflows to identify problems. For example, when executing automation associated with a given workflow, the AI agent can determine from the object store whether there is a new UI descriptor corresponding to a new version of the software application interface that the RPA robot interacts with. The AI agent can then automatically replace the corresponding UI descriptor in the workflow with the new version, so that when the RPA robot performs automation, it will be able to find the corresponding graphical elements in the UI.
[0039] In another example, the AI agent could periodically attempt to connect to a third-party application used for automation. If the AI agent cannot do so, it can escalate to a human (e.g., an IT expert, network administrator, employee of the company developing the third-party software, etc.) to make him or her aware of the problem. In yet another example, the AI agent could learn to monitor computing system resource usage. If memory and / or processor usage exceeds a certain amount, the AI agent could instruct the RPA bot running on the computing system not to perform automation.
[0040] In some embodiments, more than one healing AI agent may be used. This can be beneficial for a number of reasons, including providing the ability to fine-tune multiple AI models for AI agents specifically designed for certain problems, in order to handle large RPA bot monitoring workloads, etc. For example, one AI agent may utilize an AI model that detects connectivity problems, another AI agent may utilize an AI model that has been trained to recognize changes in graphical elements on a screen, and yet another AI model may be trained to determine if user credentials and / or licenses are no longer valid, and so on.
[0041] In some embodiments, the healing AI agent can run in a sandbox or other test environment to check for impending workflow errors, failures, and / or defects and automatically attempt to recover. This allows developers to potentially fix such issues before the automation runs in a production environment. Sandbox testing in production environments is difficult (if not impossible).
[0042] Figure 1 This is an architectural diagram illustrating a hyperautomation system 100 configured to perform agent automation and process orchestration according to an embodiment of the present invention. As used herein, “hyperautomation” refers to an automation system that combines components of process automation, agent automation, integration tools, and technologies that enhance the capabilities of automation. Some examples of these components include, but are not limited to, AI agents, AOP, and RPA robots.
[0043] Generally speaking, as used in this article, an "AI agent" is an AI-enhanced probabilistic automation capable of independent, dynamic, decision-making, execution, and adaptive actions. This can be due to the AI agent's use of a Large Language Model (LLM). AI models are, in essence, typical probabilistic models. "AOP" is automation that allows users to describe entire business processes. AOP can be created using an interface that allows the creation of business process diagrams described using Business Process Model and Notation (BPMN), which is an Extensible Markup Language (XML) description of the business process. For example, see... Figure 5 "RPA robots" are rule-based, deterministic automation that can act predictably and make deterministic decisions.
[0044] For example, in some embodiments, RPA can be used at the core of a hyperautomated system, and in others, automation capabilities can be extended through AI / machine learning (ML), process mining, analytics, agent automation, and / or other advanced tools. For instance, as the hyperautomated system learns processes, trains AI / ML models, and employs analytics, more and more knowledge work can be automated, and computing systems within an organization, such as personally used and autonomously running systems, can participate in hyperautomated processes. Some embodiments of the hyperautomated system allow users and organizations to efficiently and effectively discover, understand, and scale automation.
[0045] In this type of implementation, AI agents coexist in series with RPA robots that perform RPA and AOP. As mentioned in this paper, the AI agents are automated, enhanced with AI skills, and can act independently, make dynamic decisions, perform actions, and adapt to their performance. AI agents can dynamically utilize tools available through these RPA robots to perform document processing (e.g., see U.S. Patent Application Publication No. 2021 / 0097274), user interface (UI) automation (e.g., see U.S. Patent Nos. 10,654,166, 10,990,876, 11,080,548, 11,507,259, 11,733,668, and 11,748,069), semantic copying and pasting between source and target (e.g., see U.S. Patent No. 12,124,806 and U.S. Patent Application Publications Nos. 2023 / 0107316, 2023 / 0415338, and 2024 / 0220581), etc. The AI agent can dynamically select these tools and execute them in a pipeline.
[0046] Generally, agent automation is probabilistic automation performed by one or more AI agents. It focuses not only on individual tasks but also on the entire end-to-end process, thus expanding an organization's automation potential. AI agent-guided RPA robot teams can enable one employee to do the work of many. Agent automation, enabled by AI agents, provides managers with guidance space, doctors with more time to care for patients, developers with the ability to fine-tune their work, engineers with the freedom to innovate, and customers with seamless, personalized experiences.
[0047] In some embodiments, agent automation can achieve various technical effects, benefits, and advantages. Agent automation improves memory utilization by reducing data storage requirements and improves processor efficiency by reducing the number of calls and actions. Agent automation also potentially provides the ability to process gigabytes, terabytes, petabytes, or more of data, which is impossible with human-implemented processes, whether mental or manual. Agent automation also potentially uses fewer triggers and models through dynamic decision support. In an example scenario, RPA itself might require 100 actions, while with agent automation, this can be significantly reduced (e.g., to 15 actions). Contextual basis can also be used to tether AI agents to the context required for agent automation. Thus, contextual basis “constrains” LLM to relevant contexts.
[0048] AI agents can possess agent memory that evolves and stores user interactions, feedback, corrections, and solutions (e.g., dynamic user input from human-in-the-loop operations). As used herein, "human-in-the-loop" or human-in-the-loop operation can include AI agents and RPA robots that work in conjunction with users to receive dynamic, direct user input. As agent memory grows, AI agents can become increasingly autonomous, reducing the need for dynamic, direct human input and increasing efficiency. AI agents can also learn more effectively based on agent memory if more efficient solutions are contained within or derived from it. For example, AI agents can periodically process their memory to analyze patterns for greater autonomy.
[0049] As used in this article, "agent memory" is a dynamic caching (i.e., storage) system used to manage upgrades and tool calls. Through example operations, when an AI agent encounters a problem at runtime, it can prompt or otherwise request one or more user interactions or feedback to overcome the problem, store / cachise these interactions or feedback, and learn from them to reduce the need for repetitive human input. Based on one or more technical effects, benefits, and advantages, agent memory provides enhanced efficiency by storing solutions to common problems and minimizing potentially costly tool calls. The collaborative operation of the AI agent and agent memory has a potentially "bend curve," meaning that the need for human interaction decreases as the AI agent learns continuously via agent memory.
[0050] Generally, agent orchestration is implemented by an orchestrator application to implement one or more AOPs utilizing AI agents and RPA bots. In some embodiments, agent orchestration orchestrates AI agents (e.g., UiPath Agents™), third-party agents, RPA bots (e.g., UiPath Robots™), AOPs, and people executing agent workflows (e.g., if human approval is required). Therefore, agent orchestration supports the automation, modeling, and monitoring of complex business processes from start to finish. Agent orchestration also provides unique capabilities to orchestrate RPA bots, AI agents, third-party agents, and people across end-to-end agent workflows. Agent orchestration facilitates the successful scaling of agent automation.
[0051] AI agents used for agent automation are based on AI models, as described above, enabling them to work independently of humans and to implement these agent automata. AI agents are also goal-oriented, using context for probabilistic decision-making. Furthermore, AI agents are well-suited for specialized tasks requiring high adaptability. AI agents learn how tasks are performed and improve over time. AI agents can use and select various tools to complete tasks, gather context, and take actions (typically via RPA bots used by the AI agent). In some embodiments, AI agents can build workflows and generate automations for execution by RPA bots and / or other AI agents, such as by leveraging UiPath Autopilot™ for developers or another application that helps developers accelerate the creation and testing of automations. For example, an AI agent can leverage a designer application via API to generate another AI agent or RPA workflow, followed by human intervention in the loop to resolve any issues with the generated workflow. If correct, the workflow can be deployed. AI agents can also have varying degrees of autonomy, which is managed by agent orchestration.
[0052] AI agents generate dynamic plans by executing an "agent loop," using provided tools and context to achieve a goal according to instructions. Once the dynamic plan is generated, the AI agent utilizes its efficient execution paths. If the dynamic plan has two or more steps that can be executed in parallel, the AI agent will execute these steps in parallel based on available resources. After each step is completed, the AI agent retrieves the output of that step and regenerates the next one or more steps. Thus, the agent loop continues until the goal is achieved. The parallel execution of dynamic plan steps and the use of ecosystem tools and context are advanced capabilities of agent orchestration.
[0053] As described in this article, RPA bots are rule-based, their actions are predictable, and they make deterministic decisions. RPA bots are highly reliable and efficient, making them well-suited for performing routine tasks. RPA bots and AI agents can manage anomalies using human intervention within the loop. According to some embodiments, AI agents are more flexible, abstract, and autonomous than RPA bots and AOPs. RPA bots are generally more stable, concrete, and easier to manage than AI agents and AOPs. AOP processes typically fall between the corresponding flexibility / stability, abstraction / concreteness, and self-determination / controllability of AI agents and RPA bots.
[0054] As referenced in this article Figure 3Furthermore, AI agents and RPA robots can potentially find and use each other as tools to accomplish tasks. AI agents and RPA robots can also access and use various applications (e.g., via application programming interfaces (APIs)). Tools can be manually configured by developers for automation, and / or AI agents and RPA robots can discover and use tools at runtime.
[0055] According to some embodiments, AI agents, AOPs, and RPA robots can work collaboratively with users (e.g., in a human-in-the-loop manner), enabling them to make faster, more consistent, and more informed decisions. Furthermore, the use of AI agents, AOPs, and RPA robots allows people to accomplish more tasks because they can take on additional repetitive, mundane, and special tasks on a scale inaccessible to human users. When AI agents, AOPs, or RPA robots encounter anomalies, people can make necessary decisions. Therefore, people can be elevated and focused on becoming regulators, decision-makers, and organizational leaders.
[0056] AI models provide AI agents with the ability to infer, plan, create, and make autonomous decisions. AI models can also be used by RPA robots for task-specific activities, such as processing documents or analyzing data. AI models can be enhanced with business-specific content and context (e.g., a collection of context stores from an enterprise), thereby improving the accuracy and results of the AI models. AI models can be applied individually or simultaneously, depending on the complexity of the task. AI model selection can come from the model library of an RPA vendor, third-party models, and the Bring Your Own Model (BYOM) option (see, for example, U.S. Patent Nos. 11,738,453 and 11,748,479).
[0057] The hyper-automation system 100 includes user computing systems such as desktop computers 102, tablet computers 104, and smartphones 106. However, any desired user computing system can be used without departing from the scope of the invention, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc. Furthermore, although... Figure 1 Three user computing systems are shown, but any suitable number of user computing systems can be used without departing from the scope of the invention. For example, in some embodiments, dozens, hundreds, thousands, or even millions of user computing systems may be used. The user computing systems may be actively used by the user or run automatically without much or any user input.
[0058] As disclosed herein, in some embodiments there are three types of automation: (1) agent automation implemented by a corresponding AI agent; (2) RPA implemented by a corresponding RPA robot; and (3) composite automation to accomplish a more complex overall task through a combination of AI agents (one or more) and RPA robots (one or more). Automation 110, 112, 114 may include, but is not limited to, those performed by RPA robots and / or AI agents, whether performed individually or to achieve greater composite automation. Other processes, such as listeners, may also be implemented. These processes may be independent applications, subprocesses of another application, part of an operating system, any other suitable software and / or hardware, or any combination thereof, without departing from the scope of the invention. In fact, in some embodiments, the logical portion or entirely of the process (one or more) is implemented via physical hardware.
[0059] Each user computing system 102, 104, 106 has a corresponding automation 110, 112, 114 running thereon, such as automation implemented by RPA robots, AI agents, etc. In some embodiments, automation 110, 112, 114 may be remotely stored (e.g., stored on server 130 or database 140 and accessed via network 120) and loaded by RPA robots and / or AI agents to implement automation 110, 112, 114. Database 140 may store structured and / or unstructured data, although RPA typically requires the former. RPA automation may exist as scripts (e.g., Extensible Markup Language (XML), Extensible Application Markup Language (XAML), etc.) or be compiled into machine-readable code (e.g., as a digital link library). For example, in the case of AI agents, agent automation may be generated based on a plain text description of the desired objective.
[0060] The listener monitors and records data related to user interactions with the corresponding computing system and / or the operation of the unattended computing system, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), mobile communication network, satellite communication network, the Internet, any combination thereof, etc.). The data may include, but is not limited to, which buttons were clicked, where the mouse moved, text entered in fields, one window being minimized and another opened, applications associated with the windows, etc. In some embodiments, data from the listener may be sent periodically as part of a heartbeat message. In some embodiments, once a predetermined amount of data has been collected, after a predetermined time period has elapsed, or in both cases, the data may be sent to the core hyper-automation system 120. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0061] In the case of automations 110, 112, and 114 being RPA, automations 110, 112, and 114 can execute the logic developed in the workflow during design. A workflow can include a set of steps, defined herein as "activities," which are executed sequentially or through some other logical flow. Each activity can include actions such as clicking a button, reading a document, writing to a log panel, etc. In some embodiments, workflows can be nested or embedded.
[0062] In some embodiments, the long-running workflow of RPA is the main project supporting service orchestration, human-in-the-loop, and long-running transactions in unattended environments. For example, see U.S. Patent No. 10,860,905, the entire contents of which are incorporated herein by reference. Humans in the loop play a role when certain processes require human input (e.g., dynamic direct user input) to handle exceptions, approvals, or verifications before proceeding to the next step of an activity. In this case, process execution is paused, releasing the RPA robot until the human-in-the-loop portion of the task is completed.
[0063] Long-running workflows can support workflow segmentation via continuous activities and can be combined with invoking flows and non-user interactive activities to orchestrate human-in-the-loop tasks through RPA bot tasks. In some embodiments, multiple or many computing systems can participate in executing the logic of long-running workflows. Long-running workflows can run in a single session to accelerate execution. In some embodiments, long-running workflows can orchestrate background processes that may contain activities that perform API calls and run within the long-running workflow session. In some embodiments, these activities may be invoked by invoking flow activities. Processes with user interactive activities running in a user session can be invoked by starting a job from an orchestrator activity (the orchestrator, which will be described in more detail later). In some embodiments, the user can interact with a task that requires completing a form in the orchestrator. Activities may be included that cause the RPA bot to wait for the form task to complete before resuming the long-running workflow.
[0064] One or more of automation systems 110, 112, and 114 communicate with the core hyperautomation system 120. In some embodiments, the core hyperautomation system 120 may run an orchestrator application on one or more servers (e.g., server 130). Although a server 130 is illustrated for illustrative purposes, multiple or many servers may be employed in close proximity to each other or in a distributed architecture without departing from the scope of the invention. For example, one or more servers may be provided for orchestrator functionality, AI / ML model services, authentication, governance, and / or any other suitable functionality without departing from the scope of the invention. In some embodiments, the core hyperautomation system 120 may be incorporated into, or be part of, a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In some embodiments, the core hyperautomation system 120 may host multiple software-based servers on one or more computing systems such as server 130. In some embodiments, one or more servers of the core hyperautomation system 120, such as server 130, may be implemented via one or more virtual machines (VMs).
[0065] In some embodiments, one or more of automation systems 110, 112, and 114 may invoke one or more AI / ML models 132, which are deployed on or accessible by the core hyper-automation system 120 and trained to perform various tasks. For example, AI / ML model 132 may include models trained to find various application versions, perform computer vision (CV), perform optical character recognition (OCR), generate user interface (UI) descriptors, provide suggestions for the next activity or sequence of activities in an RPA workflow, perform semantic matching, perform natural language processing (NLP), generate or modify code and / or RPA workflows, etc. AI / ML models can be trained using labeled data, including but not limited to elements from data sources (e.g., web pages, tables, scanned documents, application interfaces, screens, etc.), previously created RPA workflows, screenshots of various application screens for various versions and their corresponding UI elements, libraries of UI objects, etc. AI / ML model 132 can be trained to achieve a desired confidence threshold without overfitting the given training dataset. Generally speaking, UI elements, UI descriptors, applications, and application screens can be considered UI objects.
[0066] Without departing from the scope of the invention, the AI / ML model 132 can be trained for any suitable purpose, as will be discussed in more detail below. In some embodiments, two or more AI / ML models 132 can be linked (e.g., in series, in parallel, or a combination thereof) such that they collectively provide collaborative output (one or more). The AI / ML model 132 can perform or assist in CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated RPA workflow generation, sequence extraction, cluster detection, audio-to-text translation, NLP, semantic matching, and any combination thereof. However, any desired number and / or type (one or more) of AI / ML models can be used without departing from the scope of the invention. For example, using multiple AI / ML models can allow the system to develop a global picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, and so on. Patterns can be determined individually by the AI / ML models or jointly by multiple AI / ML models. In some embodiments, one or more AI / ML models are locally deployed on at least one of the computing systems 102, 104, 106.
[0067] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that runs on data, and the AI / ML model itself may be, for example, a deep learning neural network (DLNN) of trained artificial “neurons” trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers performing various functions, such as statistical modeling (e.g., Hidden Markov Model (HMM)) and utilizing deep learning techniques (e.g., Long Short-Term Memory (LSTM) deep learning, encoding of previously hidden states, etc.) to perform the desired functionality.
[0068] In some embodiments, the hyper-automation system 100 can provide four main sets of functionalities: (1) discovery; (2) building automation; (3) management; and (4) collaboration. For example, in some embodiments, automation (e.g., running on a user's computing system, server, etc.) can be performed by an RPA robot, AOP, or AI agent and can provide any of the functionalities described herein. By way of example, an RPA robot can include a manned robot, an unmanaged robot, and / or a test robot. A manned robot assists the user in completing a task (e.g., via UiPathAssistant™). An unmanaged robot works independently of the user and may potentially run in the background without the user's knowledge. A test robot runs test cases against an application or RPA workflow. In some embodiments, a test robot can run in parallel on multiple computing systems.
[0069] Discovery functionality identifies various opportunities for business process automation and provides automated recommendations. Such functionality can be implemented by one or more servers, such as server 130. In some embodiments, discovery functionality may include providing an automation hub, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath AutomationHub™) can provide a mechanism for managing automation showcases with visibility and control. For example, automation ideas can be crowdsourced from employees via form submissions. Feasibility and ROI calculations for automating these ideas can be provided, documentation for future automations can be collected, and collaboration can be provided to accelerate the transition from automation discovery to implementation.
[0070] Process mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI C-Enter™) refers to the collection and analysis of data from applications (e.g., Enterprise Resource Planning (ERP) applications, Customer Relationship Management (CRM) applications, email applications, call center applications, etc.) to identify which end-to-end processes exist within an organization and how to effectively automate them, as well as indicating the impact of automation. This data may be collected from user computing systems 102, 104, 106 by, for example, listeners, and processed by servers such as server 130. In some embodiments, one or more AI / ML models 132 may be employed for this purpose. This information may be exported to an automation hub to accelerate implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of process mining goals include, but are not limited to, increasing profits, improving customer satisfaction, regulatory and / or contractual compliance, and improving employee productivity.
[0071] Task mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Enter™) identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in daily tasks, scoring these tasks for automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 can be employed to reveal repetitive task patterns in the data. Repetitive tasks suitable for automation can then be identified. In some embodiments, this information may initially be provided by a listener and analyzed on a server (e.g., server 130) of the core hyper-automation system 120. The results of task mining (e.g., XAML process data) can be exported to process documents or designer applications such as UiPath Studio™ for faster creation and deployment of automation. In some embodiments, task mining may include taking screenshots of user actions (e.g., mouse click locations, keyboard inputs, application windows and graphical elements that the user is interacting with, timestamps of interactions, etc.), collecting statistics (e.g., execution time, number of actions, text entries, etc.), editing and annotating screenshots, specifying the types of actions to be recorded, etc.
[0072] Task capture (e.g., via UiPath Automation Cloud™ and / or UiPath AI C Enter™) automatically records engaged processes while users are working, or provides a framework for unattended processes. This type of documentation can include Process Definition Documents (PDDs), framework workflows, actions captured for each part of the process, recorded user actions, and automatically generated comprehensive workflow diagrams including details about each step, as well as Microsoft Word documents. ® Desired tasks to be automated can be in the form of documents, XAML documents, etc. In some embodiments, build-ready workflows can be exported directly to designer applications, such as UiPath Studio™. Task capture can simplify the requirements gathering process for subject matter experts who interpret the process and members of Centers of Excellence (CoEs) that provide production-level automation.
[0073] Building automation can be achieved through designer applications (such as UiPath Studio™, UiPath StudioX™, or UiPath StudioWeb™). For example, developers in RPA developer facility 150 can use designer application 154 of computing system 152 to build and test for various applications and environments (such as web, mobile, SAP). ®This includes agent automation (and virtual desktops), RPA, AOP, and / or composite automation. Developers can also build AOP. For example, developers can create automation executed by RPA bots, AI agents, AOP, combinations thereof, etc., providing API integration for various applications, technologies, and platforms. Predefined activities, drag-and-drop modeling, and workflow loggers make automation easier with minimal coding. Document understanding functionality can be provided via drag-and-drop AI skills for calling one or more AI / ML models132 for data extraction and interpretation. This type of automation can handle virtually any document type and format, including tables, checkboxes, signatures, and handwriting. When data is validated or anomalies are handled, this information can be used to retrain the corresponding AI / ML models, improving their accuracy over time.
[0074] Designer application 152 may be designed to invoke one or more trained AI / ML models 132 on server 130 and / or generative AI models 172 in a cloud environment via network 120 (e.g., local area network (LAN), mobile communication network, satellite communication network, Internet, any combination thereof, etc.) to assist in automating the development process. In some embodiments, one or more of the AI / ML models may be packaged with designer application 152 or otherwise stored locally on computing system 150.
[0075] In some embodiments, the designer application 152 and one or more AI / ML models 132 may be configured to use an object store stored in the database 140. See, for example, U.S. Patent No. 11,748,069, the entire contents of which are incorporated herein by reference. Generally, an object store is a storage mechanism used for the automation of images, text, semantic data, classification associations, ontology associations, UI objects, etc. For example, an object store may include a library of UI objects that can be used to develop RPA workflows via the designer application 152. The object store can be used for activities that add UI descriptors to the workflow of the designer application 152 for UI automation. In some embodiments, one or more of the AI / ML models 132 may generate new UI descriptors and add them to the object store in the database 140.
[0076] After automation is completed in the designer application 152, it can be published to server 130 and pushed to computing systems 102, 104, 106, etc. For example, when creating new UI descriptors and / or modifying existing UI descriptors, a global repository of UI objects can be built, which can be shared and collaborated on by all automations. Regarding the object repository, taxonomy and ontology can be used. Taxonomy is a hierarchical structure of subcategories. Ontology is a formal representation of a knowledge domain, including concepts, attributes, and the relationships between them. In ontology, the relationships between categories are not necessarily hierarchical, and ontology relationships can span multiple screens of the application.
[0077] For example, integration services allow developers to seamlessly combine UI automation with API automation. Automation can be built that requires APIs or traverses both API and non-API applications and systems, such as any of the types described herein. Repositories (e.g., UiPath Object Repository™) or marketplaces (e.g., UiPath Marketplace™) for pre-built automation templates and solutions can be provided to allow developers to automate a wide variety of processes more quickly. Therefore, when building automation, the HyperAutomation System 100 can provide a user interface, development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. In some embodiments, the HyperAutomation System 100 enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots and AI agents, which provides automation for the HyperAutomation System 100.
[0078] In some embodiments, components of the hyperautomation system 100, such as designer applications (one or more) and / or external rule engines, support governance policies for managing and implementing various functionalities provided by the hyperautomation system 100. Governance is the ability of an organization to place policies in place to prevent users from developing automations (e.g., RPA bots and / or AI agents) capable of taking actions that could harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the US Health Insurance Portability and Accountability Act (HIPAA), third-party application terms of service, etc. Because developers can otherwise create automations that violate privacy laws, terms of service, etc., when performing their automations, some embodiments implement access controls and governance restrictions at the bot and / or bot design application level. In some embodiments, this provides an additional level of security and compliance by preventing developers from relying on unapproved software libraries to develop pipelines for automated processes, which could introduce security risks or operate in a manner that violates policies, regulations, privacy laws, and / or privacy policies. See, for example, U.S. Patent No. 11,733,668, the entire contents of which are incorporated herein by reference.
[0079] Management functionality can provide automated management, deployment, and optimization across the organization. In some embodiments, management functionality may include orchestration, test management, AI functionality, and / or insights. The management functionality of the hyper-automation system 100 can also serve as an integration point for third-party solutions and applications for automation applications and / or RPA robots. The management capabilities of the hyper-automation system 100 may include, but are not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots and / or AI agents.
[0080] An orchestrator application, such as UiPath Orchestrator™ (in some embodiments, it may be provided as part of UiPath Automation Cloud™, or deployed on-premises, in a VM, in a private or public cloud, in a Linux™ VM, or via UiPath Automation Suite™ as a cloud-native single-container suite), provides orchestration capabilities to deploy, monitor, optimize, scale, and secure the deployment of RPA bots and / or AI agents. A test suite (e.g., UiPath Test Suite™) provides test management to monitor the quality of deployed automation. The test suite facilitates test planning and execution, meets requirements, and ensures defect traceability. The test suite may include comprehensive test reports.
[0081] Analytics software (such as UiPath Insights™) tracks, measures, and manages the performance of deployed automation. It can link automated operations to an organization's specific key performance indicators (KPIs) and strategic outcomes. The results can be presented in dashboard formats for better understanding by human users.
[0082] For example, data services (e.g., UiPath Data Service™) can be stored in database 140, and data can be placed into a single, scalable, secure location via a drag-and-drop storage interface. Some embodiments can provide low-code or no-code data modeling and storage to automation while ensuring seamless data access, enterprise-grade security, and scalability. AI functionality can be provided by an AI hub (e.g., UiPath AI C-Enter™), which facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options even allow non-data scientists to use these functionalities. Deployed automation (e.g., RPA bots) can invoke AI / ML models, such as AI / ML model 132, from the AI hub. The performance of AI / ML models can be monitored, trained, and improved using human-verified data, such as data provided by data review center 160. Human reviewers can provide tagged data to core hyper-automation system 120 via review application 152 on computing system 154. For example, a human reviewer can verify that the predictions of AI / ML model 132 and / or generative AI model 172 are accurate, or otherwise provide corrections. The human reviewer can also provide dynamic, direct user input to the AI agent (e.g., within the scope of human-in-loop operation), and the responses and corrections provided by the human reviewer can be used to train one or more LLMs used by the AI agent to be more accurate. In other words, this dynamic input can be saved as training data for retraining AI / ML model 132 and / or generative AI model 172, and can be stored in a database such as database 140. The AI center can then schedule and execute training jobs to train new versions of the AI / ML models using the training data. Both positive and negative samples can be stored and used for the retraining of AI / ML model 132 and / or generative AI model 172.
[0083] Collaborative functionality brings humans and automation together as a team to seamlessly collaborate on desired processes. Low-code applications can be built (e.g., via UiPath Apps™) to connect browser tabs and legacy software, even in implementations lacking APIs. For example, applications can be quickly created using a web browser with a rich library of drag-and-drop controls. An application can connect to one or more automations.
[0084] Action centers (such as UiPath Action C-Enter™) provide a direct and efficient mechanism for transferring processes from automation to humans and vice versa. Humans can provide approvals or escalations, make exceptions, etc. Automation can then perform the automatic functionality of a given workflow.
[0085] A local assistant can be provided as a launch panel for users to initiate automations (e.g., UiPath Autopilot™). Such assistants may also provide semantic cut and paste functionality (e.g., UiPath Clipboard AI™). See, for example, U.S. Patent No. 12,124,806 and U.S. Patent Application Publications Nos. 2023 / 0107316, 2023 / 0415338, and 2024 / 0220581. This functionality may be provided in the system tray provided by the operating system and may allow users to interact with RPA bots and RPA-driven applications on their computing system. The interface may list automations approved by a given user and allow the user to run them. These may include off-the-shelf automations from automation marketplaces, internal automation storage from automation hubs, etc. When automations are running, they can run as local instances in parallel with other processes on the computing system, allowing the user to utilize the computing system while the automation performs its actions. In some embodiments, the assistant is integrated with task capture functionality, allowing users to record processes they are about to automate from the assistant launch panel.
[0086] In some embodiments, the HyperAutomation System 100 provides end-to-end measurement and control of automation programs of any scale. Based on the foregoing, analytics can be employed to understand the performance of automation (e.g., via UiPath Insights™). Data modeling and analysis using any combination of available business metrics and operational insights can be used for a wide range of automation processes. Custom-designed, pre-built dashboards allow visualization of data across desired metrics, discovery of new analytical insights, tracking of performance metrics, discovery of automation ROI, performance of telemetry monitoring on user computing systems, detection of errors and anomalies, and debugging of automation. An automation management console (e.g., UiPath Automation Ops™) can be provided to manage the entire lifecycle of automation. An organization can control how automation is built, what users can do with it, and which automations users can access.
[0087] In some embodiments, the hyper-automation system 100 provides an iterative platform. Processes can be discovered, automation can be built, tested, and deployed, performance can be measured, automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the process can repeat itself. This facilitates more robust and efficient automation suites.
[0088] In some embodiments, generative AI models are used in accordance with the foregoing. For example, the AI agent utilizes a generative AI model. Generative AI models can generate various types of content, such as text, images, audio, and synthetic data. Various types of generative AI models can be used, including but not limited to LLMs, generative adversarial networks (GANs), diffusion models, stream-based models, variational autoencoders (VAEs), transformers, etc. In the case of LLMs, for example, NLP models such as word2vec, BERT, GPT-3, ChatGPT, etc., can be used in some embodiments to facilitate semantic understanding and provide more accurate and human-like answers. These models can be part of the AI / ML model 132 hosted on server 130. For example, generative AI models can be trained on a large corpus of textual information to perform semantic understanding, understand the nature of content presented on the screen from the text, automatically generate code, etc. The AI agent can use such generative AI models. In some embodiments, models can be adopted and trained by existing cloud ML service providers (such as OpenAI). ® Google ® Amazon ® Microsoft ® IBM ® Nvidia ® Meta ® Generative AI models 172 provided by (e.g., [companies]) can provide this type of functionality. In generative AI embodiments where generative AI models (one or more) 172 are remotely hosted, server 130 may be configured to integrate with third-party APIs, allowing server 130 to send requests to generative AI models (one or more) 172 including necessary input information and receive returned responses (e.g., semantic matching of fields between application versions, classification of application types on screen, responses to natural language queries from users, etc.). Such embodiments can provide a more advanced and sophisticated user experience, as well as access to state-of-the-art NLP and other ML capabilities offered by these companies.
[0089] In some embodiments, one aspect of generative AI models is the use of transfer learning. In transfer learning, a pre-trained generative AI model, such as an LLM (Limited Language Model), is fine-tuned for a specific task or domain. This allows the LLM to leverage knowledge already learned during initial training and adapt it to a specific application. In the case of an LLM, the pre-training phase involves training the LLM on a large text corpus, typically consisting of billions of words. During this phase, the LLM learns relationships between words and phrases, enabling it to generate coherent and human-like responses to text-based input. The output of this pre-training phase is an LLM with a high degree of understanding of the underlying patterns in natural language.
[0090] During the fine-tuning phase, the pre-trained LLM is adapted to a specific task or domain by training it on a smaller, task-specific dataset. For example, in some embodiments, the LLM may be trained to analyze one or more types of data sources to improve its accuracy regarding their content. This data may include, but is not limited to, cue tuning or instruction tuning, where the model is specifically trained to better understand and follow specific types of instructions or cues, thereby improving its ability to perform a specific task when given appropriate instructions. This type of information can be provided as part of the training data, and the LLM can learn to focus on these areas and more accurately identify the data elements within them. Fine-tuning allows the LLM to learn the nuances of the task or domain, such as the specific vocabulary and grammar used in that domain, without requiring as much data as would be needed to train the LLM from scratch. By leveraging the knowledge learned in the pre-training phase, a fine-tuned LLM can achieve state-of-the-art performance on a specific task with a relatively small amount of training data.
[0091] LLM can utilize vector databases. Vector databases index, store, and provide access to structured or unstructured data (e.g., text, images, time-series data, etc.) and their vector embeddings. Data such as text can be tokenized, where individual letters, words, or sequences of words are parsed from the text into tokens. These tokens are then "embedded" into vector embeddings, which are numerical representations of the data. Vector databases enable LLM to quickly and massively find and retrieve similar objects in production environments, which is impossible through manual processes.
[0092] AI and ML allow for the digital representation of unstructured data without losing its semantics in vector embeddings. A vector embedding is a long string of numbers, each describing a feature of the data object it represents. Similar objects are grouped together in a vector space. In other words, the more similar the objects, the closer their vector embeddings are. Similar objects can be found using vector search, similarity search, or semantic search. The distance between vector embeddings can be calculated using various techniques, including but not limited to squared Euclidean or L2 squared distance, Manhattan or L1 distance, cosine similarity, dot product, Hamming distance, etc. Choosing the same metric for training AI / ML models can be beneficial.
[0093] Vector indexing can be used to organize vector embeddings for efficient data retrieval. With a large number of data points, calculating the distance between a vector embedding and all other vector embeddings in a vector database using the k-nearest neighbor (kNN) algorithm can be computationally expensive, as the required computation increases linearly (O(n)) with the dimensionality and number of data points. Finding similar objects using the approximate nearest neighbor (ANN) method is more efficient. The distances between vector embeddings are pre-computed, and similar vectors are organized and stored close to each other (e.g., in clusters or graphs). Similar objects can be found much faster. This process is called "vector indexing." ANN algorithms that can be used in some embodiments include, but are not limited to, cluster-based indexing, proximity graph-based indexing, tree-based indexing, hash-based indexing, compression-based indexing, etc.
[0094] Figure 2The illustration shows some combined capabilities 200 of an AI agent 210 and an RPA robot 220 according to an embodiment of the present invention. The AI agent 210 is configured to process natural language instructions and thereby achieve a desired goal; to execute through dynamic decision-making or dynamic flow control with self-healing capabilities; to store information in long-term memory and evaluate its own performance; and to learn from human-in-the-loop interactions and its own performance during execution. AI agent 210 can utilize RPA robot 220 to respond to triggers (e.g., orchestrator applications from UiPath Orchestrator™), respond based on context (i.e., RPA robot 220 can retrieve information from the context to perform deterministic steps, such as updating a document based on information retrieved from the context; alternatively, agent 210 can use the retrieved context to update a dynamic plan and execute subsequent steps to complete the goal as instructed), utilize AI models (e.g., CV models, document processing models, speech-to-text models, OCR models, etc.), utilize RPA tools (e.g., utilize tools available in the RPA ecosystem, such as fully automated workflows, workflows within automation, integration service connector calls for third-party and first-party services, RPA designer application activities, LLM calls, etc.); and perform actions that RPA robot can take based on input from the AI agent (i.e., using RPA robot as a tool). Agent 210 can also take actions to update its memory, update the plan to complete its goals according to instructions, self-evaluate and learn from actions, self-correct when it encounters obstacles, and upgrade to human assistance when it needs help.
[0095] As discussed above, various technical effects, benefits, and advantages can be achieved through agent automation in some embodiments. Agent automation improves memory utilization by reducing data storage requirements and improves processor efficiency by reducing the number of calls and operations. Agent automation also potentially provides the ability to process gigabytes, terabytes, petabytes, or more of data, which is impossible with human-implemented processes (whether mental or manual). It also potentially allows for the use of fewer triggers and models through dynamic decision-making. In an example scenario, RPA itself might require 100 actions, while with agent automation, this can be significantly reduced (e.g., to 15 actions). Contextual basis can also be used to tether AI agents to the context required for agent automation. This "constrains" the LLM within the relevant context.
[0096] As used herein, “context-based” refers to a method of improving models (such as LLMs) by combining enterprise-specific information with pre-trained knowledge, enabling accurate responses to specific or recent queries. In some embodiments, context-based approaches use external data to enhance LLM responses and obtain responses that the LLM itself is unaware of, answering queries on top of the provided context. By way of example, because unique industry terminology and complex document structures pose challenges in ensuring effective retrieval and semantic matching, context-based approaches address these challenges by providing precise document chunking to ensure that relevant information (e.g., from unique industry terminology and complex document structures) is delivered to the LLM without noise. As another example, context-based approaches offer enhanced extraction and search techniques tailored to different industry and application domains (e.g., tailored for unique industry terminology and complex document structures), which improves LLM responses.
[0097] Figure 3 The illustration shows a pool 300 of AOP, AI agents, RPA robots, and applications according to an embodiment of the present invention. The AOP pool 310 includes AOPs that implement business processes. As mentioned above, AOP can be implemented as a BPMN, which is executed by an AOP execution engine (such as Temporal). ® AOP can leverage AI agents and / or RPA robots to execute parts of business processes.
[0098] AI Agent Pool 320 includes AI agents They have been trained to perform a variety of tasks, such as investigating claims, seeking solutions with human employees, and summarizing strategies and technical specifications. The RPA robot pool 330 includes RPA robots. They perform various forms of automation, such as UI automation, semantic matching automation, and form filling automation. Application pool 340 includes applications that AI agents and / or RPA robots can interact with. For example, applications may include CRM applications, invoice applications, payroll applications, banking applications, web applications, legacy system applications, word processing applications, spreadsheet applications, email applications, etc. AI agents, RPA robots, and applications can reside on a single computing system or on multiple or many computing systems. In some embodiments, AOP is typically located in the cloud or on a server and can reside on one or more of the same computing systems as the orchestrator application 350.
[0099] AOP can trigger or invoke AI agents and RPA robots via orchestrator application 350. AI agents and RPA robots can also trigger or invoke each other via orchestrator application. For example, to invoke an RPA robot, an AI agent can make a "Start Job" call in orchestrator application 350. It should be noted that the RPA robot is deployed as an automation controlled by orchestrator application 350. AI agents and RPA robots can also trigger or invoke certain applications. For example, using information collected from human-in-the-loop actions, an AI agent can dynamically learn which RPA robot, other AI agents, and / or applications to trigger or invoke to complete a task. For example, an AI agent can learn to trigger an RPA robot via orchestrator application 350 to fill out and submit web forms. An AI agent can also learn to open Microsoft Excel. ® It also inputs form information into appropriate labels, opens and updates the payroll application, etc. The AI agent can also learn to invoke or trigger emails via the orchestrator application 350 to resolve issues; if a problem occurs, the orchestrator application 350 contacts a human customer service representative of the bank. In some embodiments, the technical effects, benefits, and advantages can be similar to those described above. Figure 1 and Figure 2 Those that are discussed.
[0100] To enable AI agents and RPA robots to find each other, AI agents can belong to the same tenant. The designer application can invoke the orchestrator to obtain a list of available RPAs. In some embodiments, there are three ways to obtain automation capabilities: (1) when a workflow is created in the designer application, the user provides a description of what the automation does; (2) the AI agent and ML technology are used to generate a summary of what a given workflow does; or (3) the developer can describe what the automation does in the designer application. The orchestrator application can also list the applications available for a given AI agent and RPA robot. In other words, the descriptions of available AI agents, RPA robots, and / or applications are derived or specified by the AI agent, ML technology, or the user.
[0101] Figure 4A and Figure 4B An example intelligent agent service interface 400 according to an embodiment of the present invention is illustrated. (Reference) Figure 4A The agent answers questions about the policy document provided in the context. The agent instruction pane 410 includes a natural language description, entered by the user, of what the AI agent intends to do. If needed, user prompts 420 allow developers to enter user prompts in the content field 422. The tools dropdown menu 430 allows developers to select tools the AI agent will utilize, such as using an API for the application, invoking an RPA bot to perform RPA, etc.
[0102] The context dropdown menu 440 allows developers to configure the context base for the AI agent. In this example, the context configuration pane 442 allows developers to provide a description via the description field 444 and a Flexible Common Schema (ECS) index for specific policy documents containing information such as contracts, agreements, and what to do via the ECS index field 446. Developers can also add additional context 450 to further supplement the context base. Human upgrade options can be configured via the dropdown menu 460.
[0103] Query field 470 allows the user to provide a query that the AI agent will respond to. When the user clicks the run button 480, the AI agent runs the query. Go to Figure 4B The results of the AI agent's execution are then displayed in the execution pane 490 when the AI agent retrieves and outputs them.
[0104] Figure 5 An example AOP development interface 500 according to an embodiment of the present invention is illustrated. The AOP development interface 500 includes AOP components 510, AI agents 520, and RPA 530 that a user can select when developing a business process. These can be selected and dragged onto a canvas 540 where the user can manually develop AOP. In this example, a credit check is implemented by retrieving customer data from a database and invoking an AI agent to determine the customer type (e.g., highly likely to pay, likely to miss payments, frequent transactions, etc.) by analyzing the customer data. This type is then provided to the RPA robot, which takes this information into account when performing the credit check. Alternatively, the AOP developer can type a description of the business process in field 550 and click the generate button 560. This text is provided to an LLM (Local Management System), which attempts to understand the requested business process and automatically create an AOP workflow. The AOP developer can edit the AOP workflow as needed.
[0105] Figure 6 An example RPA development interface 600 according to an embodiment of the present invention is illustrated. The RPA development interface 600 includes RPA components 610 that a user can select when developing an RPA workflow. These can be selected and dragged onto a canvas 620. Alternatively, the RPA developer can type a description of the RPA in field 630 and click a generate button 640. This text is provided to an LLM (Linux Virtual Machine), which attempts to understand the requested business process and automatically create the RPA workflow. The developer can then edit the RPA workflow as needed. It should be noted that in some embodiments, regarding... Figure 4A , Figure 4B , Figure 5 and Figure 6 The functionality illustrated and described can be provided in a single designer application.
[0106] Figure 7 An end-to-end AI agent, RPA robot, and AOP development and deployment system 700 is illustrated according to an embodiment of the present invention. Designer application 710 allows developers to design AOP, AI agent, and RPA workflows. Once these have been tested and validated, they are packaged and published to automation database 720.
[0107] The orchestrator application 730 manages these automations, as well as the deployment of AOP, AI agents, and RPA robots. When a human user or software process 732 requests to run an AOP, the orchestrator application 730 sends a start job request to the AOP engine 740, which selects from AOP 742 and begins the appropriate automation. While executing AOP 742, steps may be encountered that are implemented by the AI agent 750 or the RPA robot 760. When this occurs, the AOP engine 740 pauses the AOP workflow execution and sends a request to the orchestrator application 730 to send a start job request to the appropriate AI agent 750 or RPA robot 760 to execute the step.
[0108] Upon requesting an AI agent, the orchestrator application 730 sends a job initiation request to the appropriate AI agent 750. This request may include natural language text or other information provided to the orchestrator application 730 by the AOP engine 740. The AI agent 750 then assists in performing the task by executing LLM 752. The AI agent 750 then sends task-related information (e.g., the requested information, indications that the step has been completed, indications that the step failed, etc.) to the orchestrator 730, which provides this information to the AOP engine 740. The AOP engine 740 then resumes its operation.
[0109] In the event of a requested RPA robot, the orchestrator application 730 sends a job start request to the appropriate RPA robot 760. The RPA robot 760 then executes the requested RPA 762. The RPA robot 760 then sends task-related information (e.g., requested information, indications of completed steps, indications of failed steps, etc.) to the orchestrator application 730, which provides this information to the AOP engine 740. The AOP engine 740 then resumes its operation.
[0110] In some cases, AOP 742, AI agent 750, or RPA 762 may require human intervention. In this case, AOP engine 740, AI agent 750, or RPA robot 760 contacts human 770 in the automated human-in-the-loop section. After the human completes the task, AOP engine 740, AI agent 750, or RPA robot 760 resumes the automated section of the automation.
[0111] Figure 8 This is an architectural diagram illustrating an agent automation and RPA system 800 according to an embodiment of the present invention. In some embodiments, the agent automation and RPA system 800 is... Figure 1 This is part of a hyper-automation system 100. The AI agent automation and RPA system 800 includes a designer 810, which allows developers to design automations (e.g., workflows, natural language instructions for AI agents, contextual bases, tool configurations, etc.) for AI agents and RPA bots. The designer 810 can provide solutions for application integration and automating third-party applications, managing information technology (IT) tasks, and business IT processes. The designer 810 facilitates the development of automation projects, which are graphical representations of business processes. In short, the designer 810 facilitates the development and deployment of automation for RPA bots and AI agents. In some embodiments, the designer 810 can be an application running on a user's desktop, an application running remotely in a VM, a web application, etc.
[0112] Automation projects automate rules-based processes by giving developers control over the execution order and the relationships between a set of custom steps (i.e., the "activities" mentioned above) developed within the workflow. A commercial example of an embodiment of Designer 810 is UiPath Studio™. Each activity can include actions such as clicking a button, reading a document, writing to a log panel, etc. In some embodiments, the workflow can be nested or embedded.
[0113] Certain types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. Sequences are particularly well-suited for linear processes, enabling flow from one activity to another without making the workflow chaotic. Flowcharts are particularly well-suited for more complex business logic, implementing the integration of decisions and the connection of activities in more diverse ways through multiple branching logic operators. FSMs are particularly well-suited for large workflows. FSMs can use a finite number of states in their execution, triggered by conditions (i.e., transitions) or activities. Global exception handlers are particularly well-suited for determining workflow behavior and debugging processes when execution errors are encountered.
[0114] Once the workflow and / or other configurations of the AI agents are developed in designer 810, the execution of the business process is orchestrated by orchestrator 820. Orchestrator 820 orchestrates one or more robots 830, one or more AI agents 850, and / or one or more AOPs 870 to execute the workflow developed in designer 810. A commercial example of an embodiment of orchestrator 820 is UiPath Orchestrator™. Orchestrator 820 facilitates the creation, monitoring, and deployment of resources in the management environment. Orchestrator 820 can act as an integration point with third-party solutions and applications. As described above, in some embodiments, orchestrator 820 may be... Figure 1 The core of the hyper-automation system is part of 120.
[0115] It should be noted that the RPA robot 830 can operate independently for deterministic processes. The AI agent 850 and AOP 870 can also operate independently (e.g., for non-deterministic processes), or utilize one or more RPA robots 830s or other AI agents 850s as tools to complete part of their agent automation. The AI agent 850 can drive composite automation utilizing both the RPA robot 830 and the AI agent 850, or vice versa, and AOP 870 can include such composite automation.
[0116] Orchestrator 820 manages a fleet of robots 830 and AI agents 850, connecting and executing RPA robots 8530 and AI agents 850 from a central point (e.g., as required by an AOP engine implementing AOP). The types of manageable RPA robots 830 include, but are not limited to, manned robots, unmanned robots, development robots (similar to unmanned robots but used for development and testing purposes), and non-production robots (similar to manned robots but used for development and testing purposes). Manned robots are triggered by user events and work alongside humans on the same computing system. Manned robots can be used with orchestrator 820 for centralized process deployment and recording media. Manned robots can assist human users in performing various tasks and can be triggered by user events. In some embodiments, processes cannot be initiated from orchestrator 820 on this type of robot, and / or they cannot run under a locked screen. In some embodiments, managed robots can only be started from a robot tray or from a command prompt. In some embodiments, manned robots should operate under human supervision.
[0117] Unattended robots operate unattended in virtual environments, automating numerous processes. They can handle remote execution, monitoring, scheduling, and support for work queues. In some embodiments, debugging for all robot types can be performed in Designer 810. Both manned and unattended robots can automate various systems and applications, including but not limited to mainframes, web applications, VMs, and enterprise applications (e.g., those developed by SAP). ® Salesforce ® Oracle ® Applications in production and computing systems (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0118] Orchestrator 820 may have various capabilities, including but not limited to provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining connections between robot 830, AI agent 850, and / or AOP 870 and orchestrator 820 (e.g., a web application). Deployment may include ensuring that package versions are correctly delivered to designated robot 830, AI agent 850, and / or AOP for execution. Configuration may include maintaining and delivering RPA robot and AI agent environment and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include tracking robot and AI agent identification data and maintaining user permissions. Logging may include storing and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a "not just" SQL (NoSQL) database) and / or another storage mechanism (e.g., ElasticSearch, which provides the ability to store and quickly query large datasets). ® The orchestrator 820 can provide interconnectivity by acting as a centralized communication point for third-party solutions and / or applications.
[0119] Robot 830 is the execution agent of the workflow built into the implementation designer 810. A commercial example of some embodiments of robot(s) 830 is UiPath Robots™. In some embodiments, the RPA robot 830 is installed with Windows by default. ® Services managed by the Service Control Manager (SCM). As a result, this type of RPA robot 830 can open interactive Windows applications under the local system account. ® Session, and has Windows ® Service permissions.
[0120] In some embodiments, the RPA robot 830 can be installed in user mode. For such robots 830, this means they have the same permissions as the user who installed the given RPA robot 830. This feature can also be used for high-density (HD) robots, ensuring that the maximum potential of each machine is fully utilized. In some embodiments, any type of RPA robot 830 can be configured in an HD environment.
[0121] In some embodiments, the RPA robot 830 is divided into several components, each dedicated to a specific automation task. Robot components in some embodiments include, but are not limited to, SCM-managed robot services, user-mode robot services, actuators, agents, and command lines. SCM-managed robot services manage and monitor Windows. ® The session acts as an agent between the orchestrator 820 and the execution host (i.e., the computing system on which the robot 830 executes). These services are trusted and manage the credentials of the RPA robot 830. The SCM launch console application is located on the local system.
[0122] In some embodiments, user-mode robot service management and monitoring Windows ® The session acts as an intelligent agent between the orchestrator 820 and the execution host. The user-mode robot service can be trusted and manages the credentials of the RPA robot 830. If the SCM-managed robot service is not installed, Windows... ® The application can start automatically.
[0123] The actuator can be used in Windows ® A given job runs within a session (i.e., they execute workflows). The executor can know the dots per inch (DPI) setting for each monitor. The agent can be a Windows instance that displays available jobs in a system tray window. ® Demonstrates a basic (WPF) application. Note that these agents are different from AI Agent 850. Agents can be clients of services and can request to start or stop jobs and change settings. The command line is a client of the service. The command line is a console application that can request to start a job and wait for its output.
[0124] As described above, breaking down the components of robot 830 helps developers, support users, and computing systems more easily run, identify, and track what each component is performing. This allows for the configuration of specific behaviors for each component, such as setting different firewall rules for executors and services. In some embodiments, the executor can always know the DPI setting of each monitor. As a result, workflows can execute at any DPI, regardless of the configuration of the computing system that created them. In some embodiments, projects from designer 810 can also be independent of browser scaling levels. For applications that do not know the DPI or are intentionally marked as not knowing it, DPI can be disabled in some embodiments.
[0125] The agent automation and RPA system 800 in this embodiment is part of a hyper-automation system, such as... Figure 1 The hyper-automation system 100. Developers can use designer 810 to build and test RPA, AOP, and AI agents that utilize AI / ML models deployed in the core hyper-automation system 840 (e.g., as part of its AI hub). Such RPA robots can send inputs to execute one or more AI / ML models and receive outputs from them via the core hyper-automation system 840.
[0126] As described above, one or more RPA robots 830 can be listeners. These listeners can provide the core hyperautomation system 840 with information about what users are doing while using their computing systems. This information can then be used by the core hyperautomation system for process mining, task mining, task capture, etc.
[0127] An assistant / chatbot (not shown) can be provided on the user's computing system to allow the user to launch a local RPA bot. For example, the assistant / chatbot could be located in the system tray. The chatbot can have a user interface so that the user can see the text within it. Alternatively, the chatbot can run in the background without a user interface, using the computing system's microphone to listen to the user's voice.
[0128] In some embodiments, data annotation may be performed by a user of the computing system on which the RPA robot or AI agent executes, or on another computing system on which the robot or AI agent provides information. For example, if the robot invokes an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly label buttons on the screen, the user can draw rectangles around incorrectly labeled or unlabeled parts, potentially providing text with correct labeling. This information can be provided to the core hyper-automation system 540 and then later used to train a new version of the AI / ML model.
[0129] Figure 9 This is an architectural diagram illustrating a deployed RPA system 900 according to an embodiment of the present invention. In some embodiments, the RPA system 900 may be... Figure 8 Intelligent agent automation and RPA systems 800 and / or Figure 1 This is part of a hyper-automated system 100. It should be noted that in some embodiments, the architecture of the deployed RPA system 900 may not be used. The deployed RPA system 900 can be a cloud-based system, an on-premises system, a desktop-based system, etc., providing enterprise-level, user-level, or device-level automation solutions for automating different computing processes.
[0130] It should be noted that the client, server, or both may include any desired number of computing systems without departing from the scope of the invention. On the client side, the robot application 910 includes an actuator 912, an executive agent 914, and a designer 916. However, in some embodiments, the designer 916 may not run on the same computing system as the actuator 912 and the executive agent 914. The actuator 912 is running a process. Several business items can run simultaneously. In this embodiment, the executive agent 914 (e.g., Windows...) ® The service is the single point of contact for all executors 912. All messages in this embodiment are logged in orchestrator 940, which further processes them via database server 950, AI / ML server 960, indexer server 970, or any combination thereof. (See above regarding...) Figure 8 The actuator 912 may be a robot component.
[0131] In some embodiments, an RPA robot represents an association between a machine name and a username. The robot can manage multiple actuators simultaneously. This is possible in computing systems that support multiple interactive sessions running concurrently (e.g., Windows). ® On Windows Server 2012, multiple bots can run simultaneously, each using a unique username on a separate Windows instance. ® Running in the session. This is the HD robot mentioned above.
[0132] The executive agent 914 is also responsible for sending the robot's status (e.g., periodically sending "heartbeat" messages indicating that the robot is still running) and downloading the required version of the package to be executed. In some embodiments, communication between the executive agent 914 and the orchestrator 940 is always initiated by the executive agent 914. In a notification scenario, the executive agent 914 may open a WebSocket channel that will later be used by the orchestrator 940 to send commands (e.g., start, stop, etc.) to the robot.
[0133] It should be noted that, although in order to reduce Figure 9 The chaos within is not shown in this article, but the AI agent can also interact with the orchestrator 940, for example, as mentioned above. Figure 1 and Figure 8 The orchestrator 940 can orchestrate the operations of the AI agent. The orchestrator 940 can also facilitate interaction between the AI agent and the AI / ML model through the AI / ML server 960, which can store and / or facilitate access to the generative AI model.
[0134] Listener 930 monitors and records data related to user interactions with the manned computing system and / or unattended computing system operations. Listener 930 may be an RPA robot, part of an operating system, a downloadable application for the corresponding computing system, or any other software and / or hardware, without departing from the scope of the invention. In fact, in some embodiments, the logic portion of the listener may be implemented entirely through physical hardware.
[0135] On the server side, the system includes a presentation layer (web application 942, Open Data Protocol (oData) Representation State Transport (REST) Application Programming Interface (API) endpoint 944, and notification and monitoring 946), a service layer (API implementation / business logic 948), and a persistence layer (database server 950, AI / ML server 960, and indexer server 970). The orchestrator 940 includes the web application 942, the oData REST API endpoint 944, notification and monitoring 946, and the API implementation / business logic 948. In some embodiments, most actions performed by the user in the interface of the orchestrator 940 (e.g., via a browser 920) are performed by calling various APIs. These actions may include, but are not limited to, starting a job on a robot, adding / removing data from a queue, scheduling unattended job execution, etc., without departing from the scope of the invention. The web application 942 is the visual layer of the server platform. In this embodiment, the web application 942 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of the invention. In this embodiment, the user interacts with a webpage from a web application 942 via a browser 920 to perform various actions to control the orchestrator 940. For example, the user can create robot groups, assign packages to robots, analyze logs for each robot and / or each process, start and stop robots, etc.
[0136] In addition to the web application 942, the orchestrator 940 also includes a service layer that exposes the oData REST API endpoint 944. However, other endpoints may be included without departing from the scope of the invention. The REST API is consumed by the web application 942 and the executive agent 914. In this embodiment, the executive agent 914 is the supervisor of one or more robots on the client computer.
[0137] The REST API in this embodiment covers configuration, logging, monitoring, and queuing functionality. In some embodiments, the configuration endpoint can be used to define and configure application users, licenses, bots, assets, versions, and environments. For example, the logging REST endpoint can be used to log various information, such as errors, explicit messages sent by bots, and other environment-specific information. Bots can use the deployment REST endpoint to query the package version that should be executed when a start job request is used in orchestrator 940. The queuing REST endpoint can be responsible for queue and queue item management, such as adding data to the queue, retrieving transactions from the queue, and setting the status of transactions.
[0138] The monitoring REST endpoint can monitor the web application 942 and the executive agent 914. The notification and monitoring API 946 can be a REST endpoint used to register the executive agent 914, deliver configuration settings to the executive agent 914, and send / receive notifications from the server and the executive agent 914. In some embodiments, the notification and monitoring API 946 can also use WebSocket communication.
[0139] In some embodiments, APIs in the service layer can be accessed by configuring appropriate API access paths, for example, based on whether the orchestrator 940 and the entire hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The APIs for the orchestrator 940 can provide customized methods for querying statistics about various entities registered in the orchestrator 940. In some embodiments, each logical resource can be an oData entity. Components such as robots, processes, and queues can have attributes, relationships, and operations. In some embodiments, the APIs of the orchestrator 940 can be consumed by the web application 942 and / or the executive agent 914 in two ways: (1) by obtaining API access information from the orchestrator 940; or (2) by using an oAuth flow through registration with an external application.
[0140] In this embodiment, the persistence layer includes three servers: a database server 950 (e.g., an SQL server), an AI / ML server 960 (e.g., a server providing AI / ML model services, such as AI hub functionality), and an indexer server 970. In this embodiment, the database server 950 stores configurations of robots and AI agents, robot and AI agent groups, AOPs, associated processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 942. The database server 950 can manage queues and queue items. In some embodiments, the database server 950 can store messages recorded by robots and AI agents (in addition to or instead of the indexer server 970). The database server 950 can also store, for example, process mining, task mining, and / or task capture related data received from a listener 930 installed on a client. Although no arrow is shown between the listener 930 and the database 950, it should be understood that in some embodiments, the listener 930 is capable of communicating with the database 950, and vice versa. This data can be stored in the form of PDDs, images, XAML documents, etc. It should be noted that structured and / or unstructured data can be stored. Listener 930 can be configured to intercept user actions, processes, tasks, and performance metrics on the corresponding computing system where it resides. For example, listener 930 can record user actions (e.g., clicks, typed characters, locations, applications, active elements, times, etc.) on its corresponding computing system and then convert these into a suitable format for provision to database server 950 and storage therein.
[0141] The AI / ML Server 960 facilitates the integration of AI / ML models into automation systems. Pre-built AI / ML models, model templates, and various deployment options make these functionalities accessible even to non-data scientists. Deployed automation (e.g., RPA bots and / or AI agents) can invoke AI / ML models from the AI / ML Server 960. The performance of AI / ML models can be monitored, trained, and improved using human-verified data. The AI / ML Server 960 can schedule and execute training jobs to train new versions of AI / ML models. The AI / ML model server can also store and / or access generative AI models.
[0142] The AI / ML Server 960 can store data related to AI / ML models and ML packages, used to configure various ML skills for users during development. As used herein, ML skills are pre-built and trained ML models used in processes that can be used, for example, by automation. The AI / ML Server 960 can also store data about document understanding technologies and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, NLP, speech analysis, and different types of AI / ML models.
[0143] Indexer server 970, optional in some embodiments, stores and indexes information recorded by the robot. In some embodiments, indexer server 970 can be disabled via configuration settings. In some embodiments, indexer server 970 uses ElasticSearch. ® It is an open-source full-text search engine project. Messages recorded by the bot (e.g., using activities like log messages or writing lines) can be sent to indexer server 970 via logging REST endpoints (one or more), where they are indexed for future use.
[0144] Figure 10 This is an architecture diagram illustrating the relationship 1000 between a designer 1010, activities 1020, 1030, 1040, 1050, a driver 1060, an API 1070, and an AI / ML model 1080 according to an embodiment of the present invention. As described above, developers use the designer 1010 to develop workflows and automations executed by RPA robots, AI agents, and an AOP engine. Developers can design and configure RPA robot workflows 1012, design and configure agent automation 1014 for AI agents (e.g., providing natural language descriptions, contextual basis, tools, etc. for AI agents), and design and configure AOP 1016. See, for example, [link to relevant documentation]. Figure 4A , Figure 4B , Figure 5 and Figure 6In some embodiments, various types of activities can be displayed to the developer. The designer 1010 can be located locally on the user's computing system or remotely on the user's computing system (e.g., accessed via a VM or a local web browser interacting with a remote web server). The RPA robot's workflow can include user-defined activities 1020, API-driven activities 1030, AI / ML activities 1040, and / or UI automation activities 1050. User-defined activities 1020 and API-driven activities 1040 interact with the application through their APIs. In some embodiments, user-defined activities 1020 and / or AI / ML activities 1040 can invoke one or more AI / ML models 1080, which can be located locally and / or remotely on the computing system on which the robot operates.
[0145] Some embodiments are capable of identifying non-textual visual parts in an image, referred to herein as CV. However, it should be noted that in some embodiments, CV is combined with OCR. CV can be performed at least in part by one or more AI / ML models 1080. Some CV activities associated with these parts may include, but are not limited to, extracting text from segmented label data using OCR, fuzzy text matching, cropping segmented label data using ML, comparing text extracted from label data with basic fact data, etc. In some embodiments, there may be hundreds or even thousands of activities that can be implemented in user-defined activities 1020. However, any number and / or type of activities may be used without departing from the scope of the invention.
[0146] UI automation activities 1050 are a subset of specific low-level activities written in low-level code and designed for easy interaction with the screen. UI automation activities 1050 facilitate these interactions via driver 1060, which allows the robot to interact with desired software. For example, driver 1060 may include operating system (OS) driver 1062, browser driver 1064, VM driver 1066, enterprise application driver 1068, etc. In some embodiments, UI automation activities 1050 may use one or more AI / ML models 1080 to perform interactions with the computing system. In some embodiments, AI / ML models 1080 may enhance or completely replace drivers 1060. In fact, in some embodiments, driver 1060 is not included.
[0147] Driver 1060 can interact with the OS at a lower level via OS driver 1062, searching for hooks, monitoring keys, etc. Driver 1060 can facilitate integration with Chrome. ® IE ® Citrix ® SAP® Integration of features such as "click" activities. For example, the "click" activity performs the same role in these different applications via driver 1060.
[0148] Figure 11 This is an architectural diagram illustrating a computing system 1100 configured to perform preemptive repair using an AI agent according to an embodiment of the present invention. In some embodiments, the computing system 1100 may be one or more computing systems depicted and / or described herein. In some embodiments, the computing system 1100 may be part of a hyper-automated system, such as... Figure 1 and Figure 8 As shown. The computing system 1100 includes a bus 1105 or other communication mechanism for transmitting information, and processors (one or more) 1110 coupled to the bus (one or more) 1105 for processing information. The processors (one or more) 1110 can be any type of general-purpose or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The processors (one or more) 1110 may also have multiple processing cores, and at least some cores may be configured to perform specific functions. In some embodiments, multi-parallel processing may be used. In some embodiments, at least one of the processors (one or more) 1110 may be a neuromorphic circuit that includes processing elements that simulate biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0149] The computing system 1100 also includes a memory 1115 for storing information and instructions to be executed by a processor (one or more) 1110. The memory 1115 may consist of any combination of random access memory (RAM), read-only memory (ROM), flash memory, cache, static memory such as a disk or optical disk, or any other type of non-transitory computer-readable medium or combinations thereof. The non-transitory computer-readable medium may be any available medium accessible by the processor (one or more) 1110 and may include volatile media, non-volatile media, or both. The medium may also be removable, non-removable, or both. The computing system 1100 includes a communication device 1120, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, the communication device 1120 may include one or more antennas that are single, arrayed, phased, switched, beamformed, beamcontrolled, combinations thereof, and / or any other antenna configuration without departing from the scope of the invention.
[0150] Processors (one or more) 1110 are also coupled to a display 1125 via a bus 1105. Any suitable display device and tactile I / O can be used without departing from the scope of the invention. A keyboard 1130 and a cursor control device 1135, such as a computer mouse, touchpad, etc., are further coupled to the bus 1105 to enable user interaction with the computing system 1100. However, in some embodiments, a physical keyboard and mouse may be absent, and the user may interact with the device solely through the display 1125 and / or a touchpad (not shown). Any type and combination of input devices can be used, depending on design choice. In some embodiments, no physical input device and / or display is present. For example, a user may interact remotely with the computing system 1100 via another computing system with which it communicates, or the computing system 1100 may operate autonomously.
[0151] Memory 1115 stores software modules that provide functionality when executed by processor(s) 1110. These modules include the operating system 1140 of computing system 1100. These modules also include a healing AI agent module 1145, which is configured to perform all or part of the processes described herein or derivatives thereof. Computing system 1100 may include one or more additional functional modules 1150, which include additional functionality.
[0152] Those skilled in the art will understand that, without departing from the scope of this invention, a "computing system" can be embodied as a server, embedded computing system, personal computer, console, personal digital assistant (PDA), mobile phone, tablet computing device, smartwatch, quantum computing system, or any other suitable computing device, or a combination of such devices. Presenting the foregoing functionality as being performed by a "system" is not intended to limit the scope of the invention in any way, but rather to provide one example among many embodiments of the invention. In fact, the methods, systems, and apparatuses disclosed herein can be implemented in a localized and distributed manner consistent with computing technologies, including cloud computing systems. The computing system may be part of or accessible from a LAN, mobile communication network, satellite communication network, the Internet, public or private cloud, hybrid cloud, server cluster, any combination thereof, etc. Any localized or distributed architecture may be used without departing from the scope of this invention.
[0153] It should be noted that some of the system features described in this specification are presented as modules to more specifically emphasize their implementation independence. For example, modules can be implemented as hardware circuits, including custom-designed very large-scale integrated circuits (VLSI) or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Modules can also be implemented in programmable hardware devices such as field-programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, etc.
[0154] Modules can also be implemented, at least partially, in software and executed by various types of processors. Identifying units of executable code can include, for example, one or more physical or logical blocks of computer instructions, which can be organized, for example, into objects, procedures, or functions. However, the executable code of the identified module does not need to be physically located together, but can include different instructions stored in different locations that, when logically combined, constitute the module and perform its intended purpose. Furthermore, without departing from the scope of the invention, modules can be stored on a computer-readable medium, such as a hard disk drive, flash memory device, RAM, magnetic tape, and / or any other such non-transitory computer-readable medium for storing data.
[0155] In practice, an executable code module can be a single instruction or multiple instructions, and may even be distributed across multiple different code segments, different programs, and multiple memory devices. Similarly, operational data can be identified and illustrated within the module, and can be materialized in any suitable form and organized within any suitable type of data structure. Operational data can be collected as a single dataset, or it can be distributed across different locations, including different storage devices, and can exist at least in part simply as electronic signals on a system or network.
[0156] Various types of AI / ML models can be trained and deployed without departing from the scope of this invention. For example, according to embodiments of the invention, Figure 12A The illustration shows an example of a neural network 1200 trained to supplement a healing AI agent. Neural network 1200 includes many hidden layers. Both DLNNs and shallow learning neural networks (SLNNs) typically have multiple layers, although SLNNs may have only one or two layers in some cases, and are generally fewer than DLNNs. Typically, a neural network architecture includes an input layer, multiple intermediate layers, and an output layer, as is the case in neural network 1200.
[0157] DLNNs typically have many layers (e.g., 10, 50, 200, etc.), and subsequent layers often reuse features from previous layers to compute more complex general functions. SLNNs, on the other hand, tend to have only a few layers and are relatively fast to train because expert features are created beforehand from the original data samples. However, feature extraction is laborious. DLNNs, on the other hand, generally do not require expert features but often require longer training times and more layers.
[0158] For both methods, layers are trained simultaneously on the training set, and overfitting is typically checked on isolated cross-validation sets. Both techniques produce excellent results, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of layers vary depending on the problem the corresponding neural network is solving.
[0159] return Figure 12A Workflow, graphical elements, applications, screen and version information, natural language descriptions of objects in the object store, etc., are provided as input to the J neurons of hidden layer 1. Various other inputs are possible, including but not limited to computational system state information, published automation, business rules, information about what RPA workflows and / or tasks are related to, initial definitions of automation, process automation documentation, etc. All these inputs are fed into each neuron. In this example, various architectures that can be used individually or in combination are possible, including but not limited to feedforward networks, radial basis function networks, deep feedforward networks, deep convolutional inverse graph networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks, without departing from the scope of the invention.
[0160] Hidden layer 2 receives input from hidden layer 1, hidden layer 3 receives input from hidden layer 2, and so on for all hidden layers until the last hidden layer uses its output as input to the output layer. While multiple suggestions are illustrated as outputs here, in some embodiments, only a single output suggestion is provided. In some embodiments, suggestions are ranked based on confidence scores. In this embodiment, the output is a workflow correction, a notification to a person, data related to a problem that the AI agent cannot solve, the corresponding confidence score, etc. AI agent actions may include, but are not limited to, determining another AI agent or RPA robot to trigger, determining a software application to interact with, attempting to solve a problem for a person or software application, requesting help from a person in the loop, etc.
[0161] It should be noted that the number of neurons I, J, K, and L are not necessarily equal. Therefore, any desired number of layers can be used for a given layer of the neural network 1200 without departing from the scope of the invention. In fact, in some embodiments, the types of neurons in a given layer may not be exactly the same.
[0162] A neural network 1200 is trained to assign confidence scores (one or more) to appropriate outputs. To reduce inaccurate predictions, in some embodiments, only those results whose confidence scores meet or exceed a confidence threshold can be provided. For example, if the confidence threshold is 80%, outputs with confidence scores exceeding that threshold can be used, while the remaining outputs are ignored.
[0163] Neural networks are probabilistic constructs that typically have confidence scores (one or more). These scores can be learned by an AI / ML model based on the frequency with which similar inputs are correctly identified during training. Some common types of confidence scores include decimal numbers between 0 and 1 (which can also be interpreted as confidence percentages), numbers between negative ∞ and positive ∞, a set of expressions (e.g., "low," "medium," and "high"), etc. Various post-processing calibration techniques can also be employed to attempt to obtain more accurate confidence scores, such as temperature scaling, batch normalization, weight decay, negative log-likelihood (NLL), etc.
[0164] In neural networks, "neurons" are implemented as mathematical functions through algorithms, typically based on the functionalization of biological neurons. Neurons receive weighted inputs and have activation functions that sum and control whether they pass their output to the next layer. This activation function can be a non-linear threshold activity function, where nothing happens if the value is below a threshold, but the function responds linearly above the threshold (i.e., corrected linear unit (ReLU) non-linearity). Summation and ReLU functions are used in deep learning because real neurons can have approximately similar activity functions. Information can be subtracted, added, etc., via linear transformations. Essentially, neurons act as gating functions, passing output to the next layer, governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some neurons.
[0165] Figure 12B The image shows an example of neuron 1210. Input from the previous layer. Assigned corresponding weights Therefore, the collective input from the previous neuron 1 is These weighted inputs are used for neuron summation functions modified by bias, such as: (1)
[0166] The sum and the activation function Comparisons are made to determine whether neurons are "activated." For example, It can be given by the following formula: (2)
[0167] Therefore, the output y of neuron 1210 can be given by the following formula.
[0168] In this case, neuron 1210 is a single-layer perceptron. However, any suitable neuron type or combination of neuron types can be used without departing from the scope of the invention. It should also be noted that, in some embodiments, the ranges of weight values and / or the output values (one or more) of the activation function may differ without departing from the scope of the invention.
[0169] A goal, or "reward function," is often employed. The reward function explores intermediate transitions and steps with short-term and long-term rewards to guide the search of the state space and attempt to achieve the goal (e.g., finding the most accurate answer to a user query based on associated metrics). During training, various labeled data are fed through the neural network. Successful labels strengthen the weights of the neuron inputs, while unsuccessful labels weaken them. Cost functions such as mean squared error (MSE) or gradient descent can be used to penalize slightly incorrect predictions that are much smaller than very incorrect predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, data scientists can modify the reward function to provide corrections for incorrect predictions, etc.
[0170] Backpropagation is a technique for optimizing synaptic weights in a feedforward neural network. It allows you to "open" the hidden layers of the neural network to see how much loss each node is responsible for, and then update the weights to minimize the loss by giving lower weights to nodes with higher error rates, and vice versa. In other words, backpropagation allows data scientists to iteratively adjust the weights to minimize the difference between the actual and expected outputs.
[0171] The mathematical foundation of the backpropagation algorithm is optimization theory. In supervised learning, training data with known outputs is passed through a neural network, and the error is calculated using a cost function derived from the known target output; this gives the error for backpropagation. The error is then calculated at the output and converted into correction values for the network weights to minimize the error.
[0172] In the context of supervised learning, an example of backpropagation is provided below. Column vector input. x Through each layer of the network A series of N A nonlinear activity function The process involves multiplying the output of a given layer by a synaptic matrix. And add a bias vector Network output o It is given by the following formula (4)
[0173] In some embodiments, o With target output t Comparison can lead to errors. This is the expectation that is minimized.
[0174] Optimization in the form of gradient descent can be used by modifying the synaptic weights at each layer. To minimize the error. The gradient descent procedure requires, given a known target output... t Inputx Calculate the output under the following circumstances o And it produces errors. o -- t This global error is then propagated backward, giving the local error for the weight update, the calculation of which is similar to, but not exactly the same as, the calculation used for forward propagation. In particular, the backpropagation step typically requires the form: The activity function, where It is a layer j Network activities at the location (i.e., ),in and apostrophe Represents the activity function f The derivative of .
[0175] Weight updates can be calculated using the following formula: (5) (6) (7) (8) (9)
[0176] in This represents the Hadamard product (i.e., the element-wise product of two vectors). T This represents the matrix transpose, and express ,in Here, learning rate The choice was based on considerations related to machine learning. Below, This relates to the neural Hebbian learning mechanism used in neural implementations. Note the synapse. W and b These can be combined into a large synaptic matrix, where it is assumed that the input vector has additional synapses, and represents... b The extra column for synapses is categorized into W .
[0177] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., 97% or better after approximately 2000 epochs using an F2 or F4 threshold). In some embodiments, this level of accuracy can be determined using F1, F2, F4 scores, or any other suitable technique without departing from the scope of the invention. Once trained on the training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This facilitates ensuring that the AI / ML model does not "overfit," causing it to perform well on the training data but poorly on other data.
[0178] In some embodiments, the achievable accuracy level of the AI / ML model may be unknown. Therefore, if the accuracy of the AI / ML model begins to decline when analyzing evaluation data (i.e., the model performs well on training data but starts to perform poorly on evaluation data), the AI / ML model may undergo more training iterations on the training data (and / or new training data). In some embodiments, an AI / ML model is deployed only if a certain level of accuracy is achieved or if the accuracy of the trained AI / ML model is superior to that of existing deployed AI / ML models. In some embodiments, a collection of trained AI / ML models can be used to perform a task. For example, one AI / ML model may be trained to recognize images, another to recognize text, and yet another to recognize semantics and / or ontology associations, etc.
[0179] It should be noted that, in addition to or instead of neural networks, some implementations may use transformer networks such as SentenceTransformers™, a Python™ framework for state-of-the-art sentence, text, and image embeddings. These transformer networks learn associations between words and phrases with high and low scores. This trains AI / ML models to determine what is close to the input and what is not. Transformer networks can use not only pairs of words / phrases but also field length and field type.
[0180] In some embodiments, NLP models such as word2vec, BERT, GPT-3, ChatGPT, and other LLMs can be used to facilitate semantic understanding and provide more accurate and human-like answers. Other techniques, such as clustering algorithms, can be used to find similarities between groups of elements. Clustering algorithms can include, but are not limited to, density-based, distribution-based, centroid-based, and hierarchical algorithms. Examples include k-means clustering, DBSCAN clustering, Gaussian mixture model (GMM) algorithms, balanced iterative reduction, and hierarchical clustering (BIRCH) algorithms. These techniques can also aid in classification.
[0181] Figure 13This is an architecture diagram illustrating a reference architecture 1300 for a generative AI model according to an embodiment of the present invention. The architecture consists of several layers: an API plugin, a hint library, vector data source ingestion, access processing control, a model training pipeline, an evaluation layer for evaluation illusion / telemetry / evaluation, a BYOM embedding layer, and an LLM orchestration layer. There are also retrieval plugins, access control plugins, and API plugins integrated into enterprise systems.
[0182] This embodiment has three main processes:
[0183] Data ingestion and training process Data is read from multiple data stores, preprocessed, chunked, and trained through embedding models (e.g., Retrieval Augmentation (RAG)) and training pipelines (i.e., fine-tuning). Vector databases store chunked document embeddings, allowing for better semantic and similarity-based data retrieval.
[0184] Enhanced prompts for data retrieval Once a user query reaches the API layer, a prompt is selected, and data is then retrieved via a vector database or API plugin to obtain the correct contextual data before the prompt is passed to the LLM layer.
[0185] LLM inference This is the choice between using a general-purpose base model or a self-hosted base model. A fine-tuned model can be used when tailoring to a specific task or use case. Evaluate the accuracy of the response and other metrics, including hallucinations.
[0186] It should be noted that in some embodiments, generative AI models with multiple "heads" may be used. A head refers to the output layer of a generative AI model. Generative AI models, such as... Figure 1 Generative AI models172 typically have a series of layers, and each head usually shares the first few layers of the model before branching into its own different layers.
[0187] Figure 14 This is a flowchart illustrating a process 1400 for training one or more AI / ML models according to an embodiment of the present invention. In some embodiments, as described above, the AI / ML model (one or more) may be a generative AI model. In the case of a neural network, the architecture typically includes multiple layers of neurons, including input, output, and hidden layers. See, for example, [link to documentation]. Figure 12A and Figure 12B The hidden layers in the middle process the input data and generate intermediate representations of the input used to generate the output. These hidden layers can include various types of neurons, such as convolutional neurons, recurrent neurons, and / or transform neurons. Generative AI models can also have different layers.
[0188] In some embodiments, at 1410, the training process begins by providing workflow, graphical elements, application, screen and version information, natural language descriptions of objects in the object repository, etc., whether or not they are labeled. In the case of generative AI models that are typically trained in general, the training process can be skipped unless fine-tuning of the model is desired, as discussed in more detail below. Then at 1420, the AI / ML model is trained over multiple epochs, and the results are checked at 1430. While various types of AI / ML models can be used, LLMs and other generative AI models are typically trained (fine-tuned) using a process called “supervised learning,” which has also been discussed above. Supervised learning involves providing the model with a large dataset that the model uses to learn the relationship between inputs and outputs. During the training process, the model adjusts the weights and biases of neurons in the neural network to minimize the difference between the predicted outputs and the actual outputs in the training dataset.
[0189] In some embodiments, one aspect of the model is the use of transfer learning. For example, transfer learning can leverage a pre-trained model, such as ChatGPT, which is fine-tuned for a specific task or domain in step 1420. This allows the model to utilize knowledge already learned from the pre-training phase and adapt it to a specific application via the training phase of step 1420.
[0190] The pre-training phase involves training the model on a more general initial training dataset. During this phase, the model learns relationships within the data. In the fine-tuning phase (e.g., if the pre-trained model is used as the initial basis for the final model, in some embodiments, performed during step 1420 in addition to or instead of the initial training phase), the pre-trained model is adapted to a specific task or domain by training it on a smaller, task-specific dataset. For example, in some embodiments, the model may focus on certain types (one or more) of data sources. This can help the model identify data elements more accurately than a generative AI model pre-trained alone. Fine-tuning allows the model to learn nuances in the sources, such as specific vocabulary and grammar, certain graphical features, certain data formats, etc., without requiring as much data as it would need to train the model from scratch. By leveraging the knowledge learned in the pre-training phase, a fine-tuned model can achieve state-of-the-art performance on a specific task with relatively little additional training data.
[0191] In some embodiments, if the AI / ML model fails to meet the expected confidence threshold at 1440, then at 1450, supplementary training data and / or modification of the reward function are performed to help the AI / ML model better achieve its goal, and the process returns to step 1420. If the AI / ML model meets the confidence threshold at 14140, then at 1460, the AI / ML model is tested on evaluation data to ensure that the AI / ML model generalizes well and is not overfitting relative to the training data. The evaluation data includes information that the AI / ML model has not previously processed. If the evaluation data meets the confidence threshold at 1470, then the AI / ML model is deployed at 1480. If not, the process returns to step 1450, and the AI / ML model is trained further.
[0192] Figure 15 This is a flowchart illustrating a healing AI agent 1510 according to an embodiment of the present invention. The process begins with initiating an agent loop 1510 for the healing AI agent. In this embodiment, the agent loop includes monitoring the object store 1520 and monitoring connectivity issues, but other monitoring tasks can be performed by the AI agent. Besides changes to the object store UI descriptor definition, various settings, policies, etc., can be used to set the values of UI elements. This can also be observed in one environment and applied in another.
[0193] For example, consider a regular input field that uses a standard "type" activity. In UiPath... ® In an RPA workflow, this activity sends a key to a UI element. However, the new version of the application has converted this component to a dropdown list, now requiring clicking the activity, typing the activity, waiting for the activity, and selecting the item activity to correctly select the desired value. The Heal AI agent can monitor the UI changes, observe these changes, and apply appropriate fixes. The Heal AI agent can then search the object repository for other RPA workflows that will encounter this problem and proactively fix them.
[0194] When the healing AI agent 1510 determines that there are changes in the object store 1520 that could affect the workflow of the RPA robot 1530, the healing AI agent 1510 repairs the RPA workflow (e.g., replaces UI descriptors, changes activity parameters, etc.) and sends the updated workflow to the RPA robot 1530. If a connectivity issue affecting automation is detected, the healing AI agent 1540 notifies the appropriate person. However, in some embodiments, the healing AI agent first instructs the RPA robot to pause its operation for a period of time before notifying a human to check if the connectivity issue has been resolved. Furthermore, if the healing AI agent 1510 cannot find a solution to the problem, it collects data about the problem and notifies the human 1540. For example, the healing AI agent 1510 may collect data on the object store object and / or activity parameter changes it attempted to make, screenshots from the sandbox environment in which the healing AI agent 1510 attempted to repair the automation, etc. Other relevant information includes, but is not limited to, how UI elements respond to interactions (e.g., an additional click / confirmation is now required before submitting a value), the input methods used to interact with UI elements (e.g., whether via hardware events, emulation, or debugger APIs), and other settings (e.g., before, after, after pressing Enter after a value, how to clear the value of an input box). This information can then be used for the future training of the AI model utilized by the healing AI agent 1510.
[0195] Figure 16 This is a flowchart illustrating process 1600 according to an embodiment of the present invention. Process 1600 is used to proactively detect impending workflow errors, failures, and / or defects and automatically attempt to recover by a healing AI agent. The process begins at 1610 with an agent loop for the healing AI agent. The healing agent loop includes monitoring an object repository using a generative AI model that has been fine-tuned using data stored in the object repository. In some embodiments, the agent loop also includes monitoring user interface changes in the application.
[0196] At 1620, the healing AI agent determines that one or more changes have occurred in the object store that will affect the RPA bot, another AI agent, or an AOP workflow using a generative AI model. At 1630, the healing AI agent then attempts to repair the workflow using the generative AI model. In some embodiments, workflow repair includes at least one of the following: replacing the UI descriptor, changing the parameters used for RPA activities, changing the method used to interact with UI elements, changing the context basis of another AI agent, and adding, deleting, or changing AOP steps.
[0197] If the repair is successful at 1640, the healing AI agent sends the repaired workflow to the orchestrator application at 1650 to deploy the repaired workflow to one or more computing systems, on which RPA bots, another AI agent, or an AOP engine executing the workflow are deployed. Then, at 1660, the healing AI agent uses a generative AI model to search the object store for other workflows affected by the changes (one or more), repairs these other workflows using the generative AI model, and sends the repaired workflows to the orchestrator application. The healing AI agent then continues the agent loop at 1670.
[0198] However, if the repair is unsuccessful, at 1680, the healing AI agent collects data related to one or more changes and sends the collected data to one or more computing systems that are in the human loop and / or configured to retrain the generative AI model. In some embodiments, the collected data includes at least one of the following: identified object store objects, changes in RPA workflow activity parameters attempted by the healing AI agent, screenshots from the sandbox environment, how UI elements respond to interactions, input methods used to interact with UI elements, delays before interacting with UI elements, delays after interacting with UI elements, whether Enter was pressed after inputting a value, and how to clear the value from the input box in the UI. The healing AI agent then continues the agent loop at 1670.
[0199] According to an embodiment of the present invention, Figure 15 and Figure 16 The process steps executed in the process can be performed by a computer program that encodes instructions from one or more processors to perform the operation. Figure 15 and Figure 16 The computer program may be contained on a non-transitory computer-readable medium. This computer-readable medium may be, but is not limited to, hard disk drives, flash memory devices, RAM, magnetic tape, and / or any other such medium or combination of media for storing data. The computer program may include processors (one or more) for controlling a computing system (e.g., ...). Figure 11 The computing system 1100 has one or more processors 1110 to implement Figure 15 and Figure 16 The coded instructions for all or part of the process steps described herein may also be stored on a computer-readable medium.
[0200] Computer programs can be implemented in hardware, software, or a hybrid approach. A computer program can consist of modules that are operatively communicable to each other and is designed to transmit information or instructions to a display. A computer program can be configured to run on a general-purpose computer, an ASIC, or any other suitable device.
[0201] It is readily understood that, as generally described and illustrated in the figures herein, the components of various embodiments of the invention can be arranged and designed in many different configurations. Therefore, as shown in the accompanying drawings, the detailed description of embodiments of the invention is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the invention.
[0202] The features, structures, or characteristics of the invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, throughout the specification, references to "certain embodiments," "some embodiments," or similar language mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in some embodiments," "in some embodiments," "in other embodiments," or similar language appearing throughout the specification do not necessarily refer to the same set of embodiments, and the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner.
[0203] It should be noted that references to features, advantages, or similar language in this specification do not imply that all features and advantages achievable with the invention should be included in any single embodiment of the invention. Rather, language relating to features and advantages is to be understood as indicating that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Therefore, the discussion of features and advantages throughout this specification, as well as similar language, may refer to, but do not necessarily refer to the same embodiment.
[0204] Furthermore, the features, advantages, and characteristics described in this invention may be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize that this invention can be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of the invention may be recognized in certain embodiments.
[0205] It will be readily understood by those skilled in the art that the present invention described above can be practiced through steps in a different order and / or with hardware elements different from those disclosed in the configuration. Therefore, although the invention has been described based on these preferred embodiments, certain modifications, variations, and alternative constructions will be apparent to those skilled in the art while remaining within the spirit and scope of the invention. Therefore, reference should be made to the appended claims to define the scope of the invention.
Claims
1. One or more non-transitory computer-readable media, said one or more non-transitory computer-readable media storing one or more computer programs, said one or more computer programs being configured to cause at least one processor to: Initiate an agent loop for healing artificial intelligence (AI) agents, the agent loop including monitoring an object repository using a generative AI model and fine-tuning the generative AI model using data stored in the object repository; The generative AI model is used to determine that one or more changes have occurred in the object store, and these changes will affect the workflow of a Robotic Process Automation (RPA) robot, another AI agent, or an Agent Orchestration Program (AOP). Try using the generative AI model to fix the workflow.
2. The one or more non-transitory computer-readable media of claim 1, wherein in response to a successful attempted workflow repair, the one or more computer programs are configured to cause the at least one processor to: The repaired workflow is sent to an orchestrator application to deploy the repaired workflow to one or more computing systems on which the RPA robot, the other AI agent, or the AOP engine that executes the workflow is deployed.
3. The one or more non-transitory computer-readable media according to claim 2, wherein the one or more computer programs are configured to cause the at least one processor to: The generative AI model is used to search the object store for other workflows that will be affected by the one or more changes. The other workflows are repaired using the generative AI model; and The repaired other workflows are sent to the orchestrator application.
4. One or more non-transitory computer-readable media according to claim 1, wherein in response to an unsuccessful attempt to repair the workflow, the one or more computer programs are configured to cause the at least one processor to: Collect data relating to the one or more of the aforementioned changes; and The collected data is sent to one or more computing systems that are in the human loop and / or configured to retrain the generative AI model.
5. One or more non-transitory computer-readable media according to claim 4, wherein the collected data includes at least one of the following: identified object repository objects, RPA workflow activity parameter changes attempted by the healing AI agent, and screenshots from the sandbox environment.
6. The one or more non-transitory computer-readable media of claim 4, wherein the collected data includes at least one of the following: the manner in which a user interface (UI) element responds to an interaction, the input method used to interact with the UI element, the delay before interacting with the UI element, the delay after interacting with the UI element, whether enter is pressed after an input value, and how the value is cleared from an input box in the UI.
7. The one or more non-transitory computer-readable media according to claim 1, wherein the agent loop further includes monitoring connectivity issues.
8. The one or more non-transitory computer-readable media of claim 7, wherein in response to a connectivity problem being detected, the one or more computer programs are configured to cause the at least one processor to: Send a message to one or more RPA bots, one or more other AI agents and / or one or more AOP engines, instructing the one or more RPA bots, the one or more other AI agents and / or the one or more AOP engines to pause workflow execution for a period of time before retrying the execution of the corresponding workflow of the one or more RPA bots, the one or more other AI agents and / or the one or more AOP engines.
9. One or more non-transitory computer-readable media according to claim 8, wherein in response to the connectivity problem not being resolved within the time period, the one or more computer programs are configured to cause the at least one processor to: Contact in the loop regarding the aforementioned connectivity issues.
10. One or more non-transitory computer-readable media according to claim 1, wherein the agent loop further includes monitoring user interface changes in the application.
11. The one or more non-transitory computer-readable media of claim 1, wherein the repair of the workflow comprises at least one of the following: replacing the UI descriptor, changing the parameters for the RPA activity, changing the method for interacting with UI elements, changing the context basis for the other AI agent, and adding, deleting, or changing the AOP.
12. A computer-implemented method, comprising: An agent loop for healing artificial intelligence (AI) agents is initiated by a computing system. The agent loop includes monitoring an object repository using a generative AI model, which is fine-tuned using data stored in the object repository. The healing AI agent determines that one or more changes have occurred in the object store, and these changes will affect the workflow of a Robotic Process Automation (RPA) robot, another AI agent, or an agent using the generative AI model to orchestrate process AOPs; and The healing AI agent attempts to repair the workflow using the generative AI model.
13. The computer-implemented method of claim 12, wherein in response to the successful completion of the attempted workflow repair, the method further comprises: The repaired workflow is sent to an orchestrator application by the healing AI agent to deploy the repaired workflow to one or more computing systems on which the RPA robot, the other AI agent, or an AOP engine executing the workflow is deployed.
14. The computer-implemented method according to claim 13, further comprising: The healing AI agent uses the generative AI model to search the object store for other workflows that will be affected by the one or more changes; The healing AI agent uses the generative AI model to repair the other workflows; as well as The healing AI agent sends the repaired other workflows to the orchestrator application.
15. The computer-implemented method of claim 12, wherein in response to the unsuccessful attempt to repair the workflow, the method further comprises: Collect data related to the one or more changes; as well as The collected data is sent to one or more computing systems that are in the human loop and / or configured to retrain the generative AI model.
16. The computer-implemented method of claim 15, wherein the collected data includes at least one of the following: an identified object repository object, changes in RPA workflow activity parameters attempted by the healing AI agent, screenshots from the sandbox environment, the manner in which user interface (UI) elements respond to interactions, the input method used to interact with the UI elements, the delay before interacting with the UI elements, the delay after interacting with the UI elements, whether enter is pressed after inputting a value, and how to clear the value from the input box in the UI.
17. The computer-implemented method of claim 12, wherein the agent loop further includes monitoring connectivity problems, and in response to a connectivity problem being detected, the method further includes: The healing AI agent sends a message to one or more RPA robots, one or more other AI agents, and / or one or more AOP engines, instructing the one or more RPA robots, the one or more other AI agents, and / or the one or more AOP engines to pause workflow execution for a period of time before retrying the execution of the corresponding workflow of the one or more RPA robots, the one or more other AI agents, and / or the one or more AOP engines.
18. The computer-implemented method of claim 17, wherein in response to the connectivity problem not being resolved within the time period, the one or more computer programs are configured to cause the at least one processor to: Contact in the loop regarding the aforementioned connectivity issues.
19. The one or more non-transitory computer-readable media of claim 12, wherein the repair of the workflow comprises at least one of the following: replacing the UI descriptor, changing the parameters for the RPA activity, changing the method for interacting with UI elements, changing the context basis for the other AI agent, and adding, deleting, or changing the AOP.
20. One or more computing systems, including: The memory stores computer program instructions; as well as At least one processor, the at least one processor being configured to execute the computer program instructions, wherein the computer program instructions are configured to cause the at least one processor to: Initiate an agent loop for healing the AI agent, the agent loop comprising monitoring an object repository using a generative AI model, the generative AI model being fine-tuned using data stored in the object repository. The generative AI model is used to determine one or more changes that have occurred in the object store, and these changes will affect the workflow of a Robotic Process Automation (RPA) robot, another AI agent, or an agent orchestration process (AOP). Try using the generative AI model to repair the workflow. In response to a successful workflow repair attempt, the repaired workflow is sent to an orchestrator application to deploy the repaired workflow to one or more computing systems, on which the RPA bot, the other AI agent, or an AOP engine executing the workflow is deployed. In response to the failure of the attempted workflow repair, data related to the one or more changes is collected and the collected data is sent to one or more computing systems that are in the human loop and / or configured to retrain the generative AI model.
21. One or more computing systems according to claim 20, wherein in response to the successful completion of the attempted workflow repair, the computer program instructions are further configured to cause the at least one processor to: The generative AI model is used to search the object store for other workflows that will be affected by the one or more changes. The other workflows are repaired using the generative AI model. as well as The repaired other workflows are sent to the orchestrator application.
22. One or more computing systems according to claim 21, wherein the collected data includes at least one of the following: identified object repository objects, RPA workflow activity parameter changes attempted by the healing AI agent, screenshots from the sandbox environment, the way user interface (UI) elements respond to interactions, the input method used to interact with the UI elements, the delay before interacting with the UI elements, the delay after interacting with the UI elements, whether enter is pressed after inputting a value, and how to clear the value from the input box in the UI.
23. The computing system of claim 21, wherein the repair of the workflow includes at least one of the following: replacing the UI descriptor, changing the parameters used for RPA activities, changing the method used for interacting with UI elements, changing the context basis used for the other AI agent, and adding, deleting, or changing the AOP.
Citation Information
Patent Citations
Automation windows for robotic process automation
US10654166B1
Long running workflows for document processing using robotic process automation
US10860905B1
Detecting user interface elements in robotic process automation using convolutional neural networks
US10990876B1
Text detection, caret tracking, and active element detection
US11080548B1
Graphical element detection using a combined series and delayed parallel execution unified target technique, a default graphical element detection technique, or both
US11507259B2