Design time smart analyzer and runtime smart handler for robotic process automation
The integration of a cognitive AI layer within a smart analyzer and smart handler addresses the challenges of RPA workflow development and runtime issues by providing real-time suggestions and automated corrections, leading to improved accuracy and efficiency.
Patent Information
- Application Number
- JP2024066695
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-04-17
- Publication Date
- 2025-06-30
AI Technical Summary
During the development of RPA workflows, developers often introduce errors and inefficiencies, and issues can arise at runtime due to changes in the user interface, applications, or operating systems, necessitating improved design-time and run-time automation solutions.
A design-time smart analyzer and run-time smart handler that utilize a cognitive AI layer to provide suggestions for repairing and improving RPA workflows. The smart analyzer monitors workflow development and receives suggestions from the cognitive AI layer to automatically change the workflow or provide recommendations to users. The smart handler detects errors and performance issues at runtime, provides information to the cognitive AI layer, and automatically attempts to repair the automation based on the AI's suggestions.
The solution enhances the accuracy and efficiency of RPA workflow development and execution by providing real-time suggestions and automated corrections, thereby reducing errors and improving performance.
Smart Images

Figure 2025097251000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to artificial intelligence (AI), and more specifically, to a design-time smart analyzer and / or a run-time smart handler for robotic process automation (RPA) that uses a cognitive AI layer to provide suggestions.
Background Art
[0002] During the development of an RPA workflow, developers may introduce errors and / or inefficiencies into the RPA workflow logic. Also, problems may occur with the automation deployed at run-time due to unexpected situations in the user interface (UI), changes in the application or operating system on which the automation operates, etc. Therefore, an improved and / or alternative approach to design-time RPA workflow development and / or run-time automation execution may be beneficial.
Summary of the Invention
[0003] Certain embodiments of the present invention may provide solutions to problems and needs in the art that are not yet fully specified, evaluated, or solved by current RPA technologies. For example, some embodiments of the present invention relate to a design-time smart analyzer and / or a run-time smart handler for RPA that uses a cognitive AI layer to provide suggestions.
[0004] In an embodiment, the non-transitory computer-readable medium stores a computer program for a smart analyzer. The computer program is configured such that at least one processor monitors RPA workflow development in an RPA designer application and provides information regarding the RPA workflow development to the cognitive AI layer. The computer program is also configured such that at least one processor receives from the cognitive AI layer an output including one or more suggestions for repairing the RPA workflow, for improving the performance of the RPA workflow, or for both. The computer program is further configured such that at least one processor provides one or more suggestions from the cognitive AI model to a user of the RPA designer application, automatically changes the RPA workflow using one or more suggestions from the cognitive AI model, or both.
[0005] In another embodiment, one or more computing systems include a memory storing computer program instructions and at least one processor configured to execute the computer program instructions. The computer program instructions are configured such that at least one processor develops an RPA workflow in an RPA designer application and provides information regarding the RPA workflow development to the cognitive AI layer. The computer program instructions are also configured such that at least one processor receives from the cognitive AI layer an output including one or more suggestions for fixing the RPA workflow, improving the performance of the RPA workflow, or both. The computer program instructions further are configured such that at least one processor provides one or more suggestions from the cognitive AI model to a user of the RPA designer application, automatically changes the RPA workflow using one or more suggestions from the cognitive AI model, or both. The cognitive AI layer includes a generative AI model configured to facilitate an understanding of the intent of the RPA workflow, how previous activities within the RPA workflow and / or the logical flow of the RPA workflow affect a particular activity, one or more best courses of action to take to repair or improve the RPA workflow, or combinations thereof. The generative AI layer also includes one or more other AI / ML models configured to provide an intelligent analysis function to a smart analyzer using the output from the generative AI model.
[0006] In yet another embodiment, a computer-implemented method for a smart analyzer includes monitoring, by a computing system, RPA workflow development in an RPA designer application. The computer-implemented method also includes providing, by the computing system, information regarding the RPA workflow development to a cognitive AI layer. The computer-implemented method further includes receiving, by the computing system, an output from the cognitive AI layer that includes one or more suggestions for repairing the RPA workflow, improving the performance of the RPA workflow, or both. Further, the computer-implemented method includes providing, by the computing system, one or more suggestions from the cognitive AI model to a user of the RPA designer application, automatically changing the RPA workflow using one or more suggestions from the cognitive AI model, or both.
[0007] In yet another embodiment, a non-transitory computer-readable medium stores a computer program for a smart handler. The computer program is configured such that at least one processor monitors an automation performed by an RPA robot during execution and detects errors and / or one or more performance issues during the execution of the automation. The computer program is also configured such that at least one processor provides information for dealing with the errors and / or one or more performance issues to a cognitive AI layer and receives an output from the cognitive AI layer that includes one or more suggestions for repairing the automation. The computer program is further configured such that at least one processor automatically attempts to repair the automation based on the output from the cognitive AI layer, or provides one or more suggestions to a user of the computing system performing the automation, receives a selection from the user, and attempts to repair the automation based on the selected suggestions.
[0008] In another embodiment, the computing system includes a memory storing computer program instructions for a smart handler and at least one processor configured to execute the computer program instructions. The computer program instructions are configured such that at least one processor monitors the automation performed by the RPA robot at runtime and detects errors and / or one or more performance issues during the execution of the automation. The computer program instructions also are configured such that at least one processor provides information for dealing with the errors and / or one or more performance issues to the cognitive AI layer and receives an output from the cognitive AI layer that includes one or more proposals for repairing the automation. The computer program instructions further are configured such that at least one processor automatically attempts to repair the automation based on the output from the cognitive AI layer, or provides one or more proposals to the user of the computing system performing the automation, receives a selection from the user, and attempts to repair the automation based on the selected proposal. The cognitive AI layer includes a generative AI model configured to facilitate understanding of the intent of the automation, the previous activities within the RPA workflow associated with the automation and / or how the logical flow of the RPA workflow impacts a particular activity, one or more best courses of action to take to repair or improve the automation, or combinations thereof. The cognitive AI layer also includes one or more other AI / ML models configured to use the output from the generative AI model to provide an intelligent analysis function to the smart handler.
[0009] In yet another embodiment, a computer-implemented method for a smart handler includes monitoring, by a computing system, the automation performed by an RPA robot at runtime and detecting errors and / or one or more performance issues during the execution of the automation. The computer-implemented method also includes providing, by the computing system, information for handling the errors and / or one or more performance issues to a cognitive AI layer and receiving, from the cognitive AI layer, an output including one or more proposals for repairing the automation. The computer-implemented method further includes automatically attempting, by the computing system, to repair the automation based on the output from the cognitive AI layer, or providing one or more proposals to a user of the computing system performing the automation, receiving a selection from the user, and attempting to repair the automation based on the selected proposal.
Brief Description of the Drawings
[0010] To facilitate an understanding of the advantages of particular embodiments of the present invention, a more particular description of the invention briefly described above is depicted with reference to the particular embodiments illustrated in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the invention and are not to be considered limiting of its scope, but that the invention will be described and explained with additional particularity and detail by use of the following accompanying drawings.
[0011]
Figure 1
[0012]
Figure 2
[0013]
Figure 3
[0014]
Figure 4
[0015]
Figure 5
[0016]
Figure 6A
[0017]
Figure 6B
[0018]
Figure 7
[0019]
Figure 8
[0020]
Figure 9
[0021]
Figure 10
[0022] Unless otherwise noted, similar reference characters denote corresponding features consistently throughout the accompanying drawings.
DETAILED DESCRIPTION OF THE INVENTION
[0023] (Detailed Description of Embodiments) Some embodiments relate to a design-time smart analyzer and / or a runtime smart handler for RPA that uses a cognitive AI layer to provide suggestions. At design time (e.g., using UiPath Studio (trademark)), the smart analyzer analyzes the RPA workflow under development and checks for errors and inefficiencies. For example, the smart analyzer can provide suggestions for improving the workflow, such as fixing mistakes in logic that may not necessarily be functional (e.g., related to crashes). The smart analyzer can also provide suggestions to the user for improving the efficiency of the RPA workflow to be executed at runtime. For example, the smart analyzer can propose further subdividing the workflow instead of using a relatively computationally expensive loop, or propose separating nested loops. Thus, some embodiments of the smart analyzer can do more than just find errors alone and can be actively done at design time.
[0024] There are existing RPA automations that are 4, 5, 6 years old or more. If a developer were to open the RPA workflows for these automations today, the smart analyzer can help the developer improve these automations even if the tasks designed by the developer are already being achieved. For example, the smart analyzer analyzes the activities of the RPA workflow and provides suggestions for shortening the process execution time, such as leveraging new application programming interfaces (APIs) of the application or operating system, separating loops, updating the workflow to replace specific activities with more efficient versions (e.g., updated selectors or other UI descriptors), using different building blocks. The developer may not even be aware of the existence of these improvement points before they are proposed or automatically implemented by the smart analyzer.
[0025] Security can also be improved by replacing old dependencies and code that pose potential threats, or blocking potentially malicious, insecure, or unapproved-by-the-company websites. Such problems can be identified based on risky behaviors such as third-party dependencies or access to suspicious websites. In some embodiments, the smart analyzer can read logs containing information on how automation was executed during production (runtime). This information includes, but is not limited to, timestamps of the execution of each step (activity), variable values, user-defined logs, etc. For example, the smart analyzer may determine during operation that automation did not execute one or more potential paths of the RPA workflow. The smart analyzer may then propose deleting these unused paths (if any) from the RPA workflow.
[0026] On the other hand, the smart handler is executed at runtime and can be used for optimization and / or self-repair purposes in some embodiments. For example, the smart handler can analyze the code within the published automation and evaluate how the code can operate using best coding practices (e.g., no nested loops, do not assign values to uninitialized variables, etc.). The user can be warned that these problems exist, or the smart handler can also automatically correct these problems (e.g., separating nested loops, deleting assignments of values to uninitialized variables, or initializing variables first, etc.).
[0027] The smart handler can also monitor the published automation, and if an error occurs, the smart handler will attempt to "repair" the error. For example, if the target application of the automation is not displayed on the screen, the pop-up can be closed, the appropriate window can be opened, the smart handler can recognize that the connection is slow, and it can wait until certain information is received or the application startup is completed. The monitoring by the smart handler may be somewhat similar to the try / catch statements in programming languages. However, unlike these statements, the smart handler monitors the automation while it is attempting to execute a task (e.g., trying to create an email, trying to fill in and submit a form, etc.). The smart handler may observe that part of the task is not completed and intervene to find a solution. This can occur when an email is not sent, when a business rule is not satisfied, when a UI element is not clicked, etc.
[0028] To address such issues, the smart handler may pause the automation, check the location of the automation within the RPA workflow and the content on the screen, check the system logs, and may check the initial definition of the automation, process automation documents, etc. This may be done, in some embodiments, using a cognitive AI layer (e.g., incorporating generative AI) that can also include design-time information such as what the automation is intended for (e.g., sending an email, submitting a form, populating a spreadsheet, etc.). In some embodiments, the smart handler may attempt to roll back the actions performed by the automation, such as restoring the computing system to its original state, reverting the automation's operations, rolling back the actions, and rolling back database commits. For example, the smart handler may do the following: (1) propose code changes to avoid the failure while maintaining the logic underlying the automation; (2) provide proposals for application changes to avoid the problem while maintaining the underlying business logic; (3) provide proposals for business rules to improve the performance of the workflow from the perspective of performance and / or return on investment (ROI), which may be essentially process improvement proposals based on execution data; (4) provide test cases that need to be added to more quickly detect failure situations in the future; or (5) any combination thereof.
[0029] The cognitive AI layer is one or more AI / ML models that provide proposals to the design-time smart analyzer and / or the run-time smart handler. In the case of multiple AI / ML models, the cognitive AI layer may include a generative AI model that provides results to other cognitive AI layer models. These AI / ML models may, in some embodiments, be in a serial configuration, a parallel configuration, or a combination of serial and parallel configurations (plural available).
[0030] In some embodiments, the smart handler can first ask the user questions before performing an action. For example, the smart handler can display a prompt to notify the user of what problems (if any) occurred during automation and the reasons therefor. In certain embodiments, the smart handler can also provide multiple suggestions to a human and allow selection therefrom, or the user can terminate the automation. The RPA workflow of the automation can then be repaired at design time and a new automation can be generated. This can be done automatically in some embodiments with the assistance of a smart analyzer.
[0031] To find errors, in some embodiments, the generative AI model of the cognitive AI layer may be provided with formatted data from other cognitive AI layer models. The data provided to the generative AI model may include, additionally or alternatively, the following: (1) the original automation blueprint (code) representing the intent of automation (e.g., process definition documents, automation code, snapshots during design and debugging, etc.), (2) the current execution log representing a trace that shows in text how the execution has been performed so far, (3) the current execution record that visually shows how the execution has been performed so far, (4) the interpretation of code, logs, and records by the cognitive AI layer model (e.g., text output, JavaScript Object Notation (JSON), Extensible Markup Language (XML), etc.), (5) an RPA automation language such as UiPath® automation language that includes activities and building blocks (e.g., in JSON / XML), (6) a screen ontology such as UiPath® screen ontology that represents how screen elements and the screen are structured and how they interact with each other in text, programming language code, JSON, XML, etc. (e.g., from an object repository), (7) the technical representation of the current screen such as the Document Object Model (DOM) of a web page, (8) boundaries or rules that prevent an RPA robot from performing a specific action or accessing specific information, e.g., in text form, or (9) any combination thereof. The above all represent input data to the generative AI model and other AI / ML models of the potential cognitive AI layer. This data may be in Hypertext Markup Language (HTML), XML, Extensible Application Markup Language (XAML), or another structured format in some embodiments.
[0032] The generative AI model and / or other models of the cognitive AI layer may process this data, provide suggestions, and / or perform the above repairs automatically or based on user confirmation from a prompt.
[0033] For example, consider the case where an RPA workflow is developed for the current version of a target application. After production automation is created for that RPA workflow, a new version of the application is released that moves graphical elements to a different location in the UI or another application screen, changes the text labels of fields within the application to another semantically similar word, and changes the appearance of buttons that perform the same function as the function the activity is scheduled to execute. The generative AI model and / or other cognitive AI layer models can be called by the smart handler to search for elements that were not found within a specific period or elements that no longer match the original version. This problem can be automatically resolved by the smart handler or the smart handler can present recommendations to the user.
[0034] To perform a repair, the cognitive AI layer needs to understand the following: (1) the purpose of the RPA process, (2) how the RPA process has been performed so far, (3) what a human would do if the RPA process fails before completion, and what the best course of action regarding the repair is. However, in some embodiments, generative AI and CV alone may not be sufficient to create such a cognitive AI layer. To address this problem, some embodiments employ one or more additional AI / ML models in the cognitive AI layer to facilitate this understanding. In some embodiments, proposals may be provided regarding what should be changed so that the same problem does not occur in future versions of the automation. For recovery during live (runtime) execution, blocks of code may be automatically generated and / or automation may be created that understands the intent and reason for the failure. In other words, in some embodiments, rather than providing proposals, the smart handler may attempt to address the problem itself.
[0035] In some embodiments, a specific minimum confidence score (e.g., 95%, 99%, etc.) may be required to accept a repair recommendation from the cognitive AI layer. Alternatively, the recommendations from the cognitive AI layer model may be tried regardless of the confidence score, and an attempt can also be made to see if it works to correct the activity. In certain embodiments, different types of generative AI models and / or other AI / ML models can be used simultaneously or sequentially to attempt repairs.
[0036] Figure 1 is an architectural diagram showing a hyper-automation system 100 according to an embodiment of the present invention. As used herein, "hyper-automation" refers to an automation system that combines components of process automation, integration tools, and technologies that amplify the ability to automate work. For example, in some embodiments, robotic process automation (RPA) is used at the core of the hyper-automation system, and in certain embodiments, the automation capabilities can be extended by AI / machine learning (ML), process mining, analytics, and / or other advanced tools. When the hyper-automation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work can be automated, and both computing systems within an organization, such as those used by individuals and those that operate autonomously, can all engage as participants in the hyper-automation process. The hyper-automation systems of some embodiments enable users and organizations to discover, understand, and expand automation efficiently and effectively.
[0037] The hyper-automation system 100 includes user computing systems such as desktop computer 102, tablet 104, and smartphone 106. However, any desired user computing system, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc., can be used without departing from the scope of the present invention. Also, although three user computing systems are shown in FIG. 1, any appropriate number of user computing systems can be used without departing from the scope of the present invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of user computing systems can be used. The user computing system may be actively used by the user or may be automatically executed without much or any user input.
[0038] Each user computing system 102, 104, 106 has its respective automation process(es) 110, 112, 114 running thereon. In some embodiments, the automation process is stored remotely (e.g., accessed via network 120 on server 130 or database 140) and loaded by an RPA robot to implement the automation. The automation can exist as a script (e.g., XML, XAML, etc.) or can be compiled into machine-readable code (e.g., as a digital link library).
[0039] Automation processes 110, 112, 114 (plural possible) may include, but are not limited to, RPA robots, a part of an operating system, downloadable applications (plural possible) for respective computing systems, any other suitable software and / or hardware, or any combination thereof without departing from the scope of the present invention. In some embodiments, one or more processes 110, 112, 114 may be listeners. A listener may be, without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the listener(s) is partially or fully implemented via physical hardware.
[0040] The listener monitors and records data related to user interactions with respective computing systems and / or the operation of unattended computing systems, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data may include, but is not limited to, which button was clicked, where the mouse moved, text entered in a field, one window was minimized and another window was opened, the application associated with the window, etc. In certain embodiments, the data from the listener may be transmitted periodically as part of a heartbeat message. In some embodiments, the data may be transmitted to the core hyper-automation system 120 when a predetermined amount of data is collected, after a predetermined period has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0041] The automation process can execute the logic developed in the workflow during design time. In the case of RPA, the workflow can include a set of steps performed in a sequence or some other logical flow, defined herein as an "activity". Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, the workflow can be nested or embedded.
[0042] The long-running workflows for RPA in some embodiments are master projects that support service orchestration, human intervention, and long-running transactions in an unattended environment. See, for example, U.S. Patent No. 10,860,905, which is hereby incorporated by reference in its entirety. Human intervention occurs when a particular process requires human input for exception handling, approval, or verification before proceeding to the next step of the activity. In this case, the execution of the process is paused and the RPA robot is released until the human task is completed.
[0043] Long-running workflows may support fragmentation of the workflow via persistence activities, combined with call processes and non-user interaction activities, and orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or a large number of computing systems may participate in the execution of the logic of the long-running workflow. Long-running workflows may be executed in a session to facilitate rapid execution. In some embodiments, long-running workflows may orchestrate background processes that execute API calls and may include activities that execute in a long-running workflow session. These activities may be called by call process activities in some embodiments. Processes having user interaction activities that execute in a user session may be called by starting a job from a conductor activity (the conductor is described in more detail later in this specification). The user may interact in some embodiments through tasks that require the user to complete a form in the conductor. Activities may be included that cause the RPA robot to wait for the form task to complete and then resume the long-running workflow.
[0044] One or more automation processes 110, 112, 114 communicate with a core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may execute a conductor application on one or more servers such as server 130. Although one server 130 is shown for illustration purposes, a plurality or a number of servers in close proximity to each other or in a distributed architecture may be employed without departing from the scope of the present invention. For example, one or more servers may be provided for conductor functionality, AI / ML model provision, authentication, governance, and / or any other suitable functionality without departing from the scope of the present invention. In some embodiments, the core hyper-automation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyper-automation system 120 may host a plurality of software-based servers on one or more computing systems such as server 130. In some embodiments, one or more servers of the core hyper-automation system 120 such as server 130 may be implemented via one or more virtual machines (VMs).
[0045] In some embodiments, one or more automation processes 110, 112, 114 may invoke one or more AI / ML models 132 that are deployed on or accessible by a core hyper-automation system 120 and are trained to accomplish various tasks. For example, the AI / ML models 132 may include models trained to search for various application versions, perform CV, perform OCR, generate UI descriptors, and provide suggestions for the next activity or sequence of activities in an RPA workflow. The AI / ML models may be trained using labeled data including elements of data sources (e.g., web pages, forms, scanned documents, application interfaces, screens, etc.), previously created RPA workflows, screenshots of various application screens of various versions, including corresponding UI elements, libraries of UI objects, etc. The AI / ML models 132 may be trained to achieve a desired confidence threshold without overfitting to a given set of training data.
[0046] The AI / ML model 132 can be trained for any suitable purpose without departing from the scope of the present invention, as will be discussed in more detail later in this specification. Two or more AI / ML models 132 may be chained in some embodiments (e.g., in series, in parallel, or a combination thereof) such that they collectively provide a collaborative output(s). The AI / ML model 132 may perform or assist with CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automatic RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, any combination of these, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of the present invention. By using multiple AI / ML models, for example, the system can develop an overall picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, etc. Patterns may be determined individually by an AI / ML model or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one of the computing systems 102, 104, 106.
[0047] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that executes on data, and the AI / ML model itself can be, for example, a deep learning neural network (DLNN) of artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions such as statistical modeling (e.g., hidden Markov model (HMM)), and may utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform the desired functions.
[0048] In some embodiments, the Hyper-Automation System 100 may provide four main functional groups: (1) Discovery, (2) Automation Construction, (3) Management, and (4) Engagement. Automation (e.g., executed on user computing systems, servers, etc.) may be performed by software robots such as RPA robots in some embodiments. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with users to assist them in tasks (e.g., via UiPath Assistant™). Unattended robots operate independently of users and may potentially execute in the background without the user's knowledge. Test robots are unattended robots that execute test cases against applications or RPA workflows. In some embodiments, test robots may be executed in parallel on multiple computing systems.
[0049] The Discovery function may discover various opportunities for automating business processes and provide automated recommendations therefor. Such a function may be implemented by one or more servers such as server 130. The Discovery function may include, in some embodiments, providing an Automation Hub, Process Mining, Task Mining, and / or Task Capture. The Automation Hub (e.g., UiPath Automation Hub™) may provide a mechanism for managing the rollout of automation with visibility and control. Automation ideas may be crowdsourced from employees, for example, via a submission form. Feasibility and ROI calculations for automating these ideas are provided, documentation for future automation is collected, and collaboration for quickly going from discovery to construction of automation may be provided.
[0050] Process mining (e.g., via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) refers to the process of collecting and analyzing data from applications (such as enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how to effectively automate them, and the impact of automation. This data can be obtained, for example, by a listener from user computing systems 102, 104, 106 and processed by a server such as server 130. In some embodiments, one or more AI / ML models 132 can be employed for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of the goals of process mining include, but are not limited to, increased profit, improved customer satisfaction, regulatory and / or compliance, and improved employee efficiency.
[0051] Task mining (e.g., via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in everyday tasks and scores such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to reveal repetitive task patterns in the data. Repetitive tasks ripe for automation can then be identified. This information may first be provided by a listener and, in some embodiments, analyzed on a server of a core hyperautomation system 120 such as server 130. Discoveries from task mining (e.g., XAML process data) are exported to a process document or a designer application such as UiPath Studio (trademark) to enable more rapid creation and deployment of automation. Task mining in some embodiments may include taking screenshots with user actions (e.g., mouse click location, keyboard input, application windows and graphical elements the user interacted with, timestamps for interactions, etc.), collecting statistical data (e.g., execution time, number of actions, text input, etc.), editing and annotating screenshots, specifying the types of actions being recorded, etc.
[0052] Task capture (via UiPath Automation Cloud™ and / or UiPath AI Center™) automatically documents attended processes as the user works or provides a framework for unattended processes. Such documentation may include process definition documents (PDDs), skeleton workflows, capture of actions for each part of the process, recording of the user's actions, and automatic generation of comprehensive workflow diagrams that include details about each step, tasks that are desirably automated in formats such as Microsoft Word® documents, XAML files, etc. Constructible workflows may, in some embodiments, be directly exported to designer applications such as UiPath Studio™. Task capture can streamline the requirements gathering process for both subject matter experts who describe the process and Center of Excellence (CoE) members who provide production-grade automation.
[0053] Automation can be achieved through designer applications (such as UiPath Studio™, UiPath StudioX™, UiPath Studio Web™, etc.). For example, RPA developers at an RPA development facility 150 can use an RPA designer application 154 on a computing system 152 to build and test automation for various applications and environments such as web, mobile, SAP®, and virtual desktop. API integration can be provided for various applications, technologies, and platforms. Pre-defined activities, drag-and-drop modeling, and workflow recorders can facilitate automation with minimal coding. Document understanding capabilities can be provided through drag-and-drop AI skills for data extraction and interpretation that call one or more AI / ML models 132. Such automation can process virtually any document type and format, including tables, checkboxes, signatures, and handwritten. When data is validated or exceptions are handled, this information may be used to retrain the respective AI / ML models, improving their accuracy over time.
[0054] The RPA designer application 152 can be designed to call one or more of the trained AI / ML models 132 on the server 130 and / or the generative AI model 172 within the cloud environment via a network 120 (such as a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.) to assist in the RPA automation development process. In some embodiments, one or more of the AI / ML models can be packaged with the RPA designer application 152 or otherwise stored locally on the computing system 150.
[0055] In some embodiments, the RPA Designer applications 152 and one or more AI / ML models 132 may be configured to use an object repository stored in the database 140. The object repository can include a library of UI objects that can be used to develop RPA workflows via the RPA Designer application 152. The object repository can be used to add UI descriptors to activities within the workflows of the RPA Designer application 152 for UI automation. In some embodiments, one or more of the AI / ML models 132 can generate new UI descriptors and add them to the object repository within the database 140. When automation is completed in the Designer application 152, the automation is published on the server 130 and can be pushed out to computing systems 102, 104, 106, etc.
[0056] With the integration service, developers can seamlessly combine, for example, UI automation and API automation. Automations that require APIs or that cross both API and non-API applications and systems can be built. A repository of pre-built RPA and AI templates and solutions (e.g., UiPath Object Repository (trademark)) or marketplace (e.g., UiPath Marketplace (trademark)) can be provided so that developers can automate a wide variety of processes more quickly. Thus, when building automations, the hyper-automation system 100 can provide a user interface, a development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. The hyper-automation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.
[0057] In some embodiments, components of the hyper-automation system 100, such as designer applications and / or external rule engines, provide support for managing and enforcing governance policies for controlling the various functions provided by the hyper-automation system 100. Governance is the ability of an organization to introduce policies to prevent users from developing automations (such as RPA robots) that can harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), the terms of use of third-party applications, etc. Otherwise, developers could create automations that violate privacy laws, terms of use, etc. during the execution of their automations. Thus, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. This can provide an additional level of security and compliance in the automation process development pipeline in some embodiments by preventing developers from introducing security risks or taking dependencies on unapproved software libraries that could operate in a way that violates policies, regulations, privacy laws, and / or privacy policies. See, for example, U.S. Patent No. 11,733,668, which is hereby incorporated by reference in its entirety.
[0058] The management function can provide management, deployment, and optimization of automation across the entire organization. The management function may include, in some embodiments, orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management function of the hyper-automation system 100 can include, among other things, but not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnectivity of RPA robots.
[0059] Conductor applications such as UiPath Orchestrator (trademark) (which may be provided as part of UiPath Automation Cloud (trademark) in some embodiments, or on-premises, VM, private or public cloud, on a Linux (trademark) VM, or as a cloud-native single-container suite via UiPath Automation Suite (trademark)) provide orchestration capabilities to deploy, monitor, optimize, scale, and secure RPA robot deployments. A test suite (e.g., UiPath Test Suite (trademark)) can provide test management for monitoring the quality of deployed automation. The test suite can facilitate test planning and execution, requirement fulfillment, and defect traceability. The test suite can include comprehensive test reports.
[0060] Analytics software (e.g., UiPath Insights (trademark)) can track, measure, and manage the performance of deployed automation. The analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes of the organization. The analytics software can present results in dashboard form for easier understanding by human users.
[0061] A data service (e.g., UiPath Data Service (trademark)) can, for example, be stored in a database 140 and bring data into a single, scalable, and secure location using a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities may be provided by an AI center (e.g., UiPath AI Center (trademark)), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may enable non-data scientists to access such capabilities. Deployed automation (e.g., an RPA robot) may call an AI / ML model from an AI center such as AI / ML model 132. The performance of the AI / ML model can be monitored and trained and improved using human-verified data such as that provided by a data review center 160. A human reviewer may provide labeled data to the core hyperautomation system 120 via a review application 152 on a computing system 154. For example, a human reviewer may verify that predictions by the AI / ML model 132 and / or generative AI model 172 are accurate or otherwise provide corrections. This dynamic input may then be saved as training data for retraining the AI / ML model 132 and / or generative AI model 172, for example, stored in a database such as database 140. The AI center can then schedule and perform a training job to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining the AI / ML model 132 and / or generative AI model 172.
[0062] The engagement function involves humans and automation as one team for seamless collaboration regarding a desired process. Low-code applications can be built (e.g., via UiPath Apps™) even if they lack APIs in some embodiments, to connect browser tabs and legacy software. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. An application can be connected to one automation or multiple automations.
[0063] The Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for passing a process from automation to humans or vice versa. Humans can provide approvals or escalations and perform exception handling, etc. Then, the automation can execute the automated functions of a given workflow.
[0064] The local assistant can be provided as a launch pad for the user to start an automation (e.g., UiPath Assistant (trademark)). This feature can be provided, for example, in the tray provided by the operating system, enabling the user to interact with RPA robots and RPA robot - enabled applications on their computing system. The interface can list the automations approved for a given user and allow the user to execute them. These can include off - the - shelf automations from an automation marketplace, an internal automation store in an automation hub, etc. While an automation is running, they can execute as a local instance in parallel with other processes on the computing system so that the user can use the computing system while the automation performs its actions. In certain embodiments, the assistant is integrated with a task capture function so that the user can document the processes that will soon be automated from the assistant's launch pad.
[0065] Chatbots (e.g., UiPath Chatbots (trademark)), social messaging applications, and / or voice commands can enable the user to execute an automation. This can simplify access to the information, tools, and resources necessary to conduct customer interactions or other activities. Conversations between people can be easily automated just like other processes. Trigger RPA robots launched in this way may be able to perform actions such as order status checks and data posting to CRM using plain - language commands.
[0066] End-to-end measurement of automation programs at any scale and governance can be provided by the hyper-automation system 100 in some embodiments. As such, analytics (e.g., via UiPath Insights™) may be employed to understand the performance of the automation. Data modeling and analytics using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for the automation, perform remote monitoring on the user's computing system, detect errors and anomalies, and debug the automation. An automation management console (e.g., UiPath Automation Ops™) may be provided to manage the automation throughout its lifecycle. An organization may govern how the automations are built, what users can do with them, and which automations users can access.
[0067] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes can be discovered, automations can be built, tested, and deployed, performance can be measured, use of the automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the process itself can be repeated. This facilitates a more robust and effective set of automations.
[0068] In some embodiments, a generative AI model is used. Generative AI can generate various types of content, such as text, images, audio, and synthetic data. Various types of generative AI models can be used, including but not limited to large language models (LLMs), generative adversarial networks (GANs), variational autoencoders (VAEs), transformers, etc. These models may be part of the AI / ML model 132 hosted on the server 130. For example, the generative AI model can be trained on a large corpus of text information to perform semantic understanding, understand the nature of what exists on the screen from text, automatically generate code, etc. In certain embodiments, a generative AI model 172 provided by an existing cloud ML service provider such as OpenAI®, Google®, Amazon®, Microsoft®, IBM®, Nvidia®, Facebook® can be adopted and trained to provide such functionality. In a generative AI embodiment where the generative AI model(s) 172 is hosted remotely, the server 130 can be configured to integrate with a third-party API, whereby the server 130 can send requests containing the necessary input information to the generative AI model(s) 172 and receive its responses (e.g., semantic matching of fields between versions of an application, classification of the type of application on the screen, etc.). Such embodiments can not only provide a more advanced and sophisticated user experience but also provide access to state-of-the-art natural language processing (NLP) and other ML capabilities offered by these companies.
[0069] One aspect of the generative AI model in some embodiments is the use of transfer learning. In transfer learning, a pre-trained generative AI model such as an LLM is fine-tuned for a specific task or domain. This allows the LLM to leverage the knowledge it has already learned during its initial training and adapt it to a specific application. In the case of an LLM, during the pre-training phase, the LLM is typically trained on a large text corpus consisting of billions of words. During this phase, the LLM learns the relationships between words and phrases, enabling it to generate consistent human-like responses to text-based inputs. The output of this pre-training phase is an LLM that highly understands the patterns underlying natural language.
[0070] In the fine-tuning phase, the pre-trained LLM is adapted to a specific task or domain by training the LLM on a smaller dataset specific to the task. For example, in some embodiments, the LLM can be trained to analyze a specific type or types of data sources to improve the accuracy regarding its content. Such information can be provided as part of the training data, and the LLM can learn to focus on these areas and more accurately identify the data elements within them. Through fine-tuning, without requiring as much data as would be necessary to train the LLM from scratch, the LLM can learn the subtle nuances of the task or domain, such as the specific vocabulary and syntax used in that domain. By leveraging the knowledge learned during the pre-training phase, the fine-tuned LLM can achieve state-of-the-art performance on a specific task with a relatively small amount of training data.
[0071] LLMs can be trained using vector databases. Vector databases index, store, and provide access to structured or unstructured data (e.g., text, images, time series data, etc.) along with their vector embeddings. Data such as text can be tokenized, where single characters, words, or sequences of words are parsed from the text into tokens. These tokens are then "embedded" into vector embeddings, which are numerical representations of this data. Vector databases allow users to search and retrieve similar objects quickly and at scale in production environments.
[0072] AI and ML allow unstructured data to be represented numerically without losing its semantic meaning in vector embeddings. A vector embedding is a long list of numbers, where each number represents a feature of the data object that the vector embedding represents. Similar objects are grouped together in the vector space. In other words, the more similar the objects are, the closer the vector embeddings that represent them are to each other. Similar objects may be found using vector search, similarity search, or semantic search. The distance between vector embeddings may be calculated using various techniques, including but not limited to Euclidean squared or L2 squared distance, Manhattan or L1 distance, cosine similarity, dot product, Hamming distance, etc. It may be beneficial to choose the same metric used to train the AI / ML model.
[0073] Vector indexing can be used to organize vector embeddings so that data can be retrieved efficiently. When the number of data points is large, using the k-nearest neighbor (kNN) algorithm to calculate the distance between a vector embedding in a vector database and all other vector embeddings can be computationally expensive because the necessary calculations increase linearly (O(n)) depending on the dimension and the number of data points. It is more efficient to use an approximate nearest neighbor (ANN) approach to find similar objects. The distances between vector embeddings are pre-computed, similar vectors are organized, and stored close to each other (e.g., within a cluster or graph), so that similar objects can be found more quickly. This process is called "vector indexing". ANN algorithms that can be used in some embodiments include, but are not limited to, clustering-based indexing, proximity graph-based indexing, tree-based indexing, hash-based indexing, compression-based indexing, etc.
[0074] FIG. 2 is an architecture diagram showing an RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 may facilitate the development of automation projects that are graphical representations of business processes. Briefly, the designer 210 facilitates the development and deployment of workflows and robots. In some embodiments, the designer 210 may be an application that runs on a user's desktop, an application that runs remotely on a VM, a web application, or the like.
[0075] Automation projects enable the automation of rule - based processes by giving developers control over the execution order and relationships between steps in a custom set of workflows defined as "activities" in this specification as described above. A commercial example of an embodiment of Designer 210 is UiPath Studio™. Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows can be nested or embedded.
[0076] Some types of workflows can include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. A sequence may be particularly suitable for a linear process that enables the flow from one activity to another without cluttering the workflow. A flowchart may be particularly suitable for more complex business logic, enabling the integration of decision - making and connection of activities in more diverse ways through multiple branching logic operators. An FSM may be particularly suitable for large - scale workflows. An FSM can use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. A global exception handler may be particularly suitable for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0077] When a workflow is developed within the designer 210, the execution of the business process is coordinated by the conductor 220, which coordinates one or more robots 230 that execute the workflow developed within the designer 210. A commercial example of an embodiment of the conductor 220 is UiPath Orchestrator™. The conductor 220 facilitates the management of the generation, monitoring, and deployment of resources in an environment. The conductor 220 can operate as an integration point with third - party solutions and applications. As such, in some embodiments, the conductor 220 can be part of the core hyper - automation system 120 of FIG. 1.
[0078] The conductor 220 can manage all robots 230 and connect and execute the robots 230 from a central point. The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non - production robots (similar to attended robots 232 but used for development and testing purposes). Attended robots 232 are triggered by user events and operate alongside humans on the same computing system. Attended robots 232 can be used with the conductor 220 for centralized process deployment and logging medium. Attended robots 232 may assist human users in achieving various tasks and may be triggered by user events. In some embodiments, a process cannot start from the conductor 220 on this type of robot and / or they cannot execute under a locked screen. In certain embodiments, attended robots 232 can only be launched from a robot tray or a command prompt. Attended robots 232 preferably operate under human supervision in some embodiments.
[0079] The unattended robot 234 can operate unmanned in a virtual environment and automate many processes. The unattended robot 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can be performed by the designer 210 in some embodiments. Both attended and unattended robots can automate various systems and applications including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0080] The conductor 220 can have various capabilities including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning can include creating and maintaining a connection between the robot 230 and the conductor 220 (e.g., a web application). Deployment can include ensuring the correct delivery of the package version to the assigned robot 230 for execution. Configuration can include maintaining and delivering the robot environment and process configuration. Queuing can include providing management of queues and queue items. Monitoring can include tracking specific data of the robot and maintaining user permissions. Logging can include saving and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a “not only” SQL (NoSQL) database) and / or another storage mechanism (e.g., ElasticSearch® which provides the ability to store large datasets and execute queries quickly). The conductor 220 can provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0081] The robot 230 is an execution agent that implements the workflows built by the designer 210. One commercial example of some embodiments of the robot(s) 230 is UiPath Robots™. In some embodiments, the robot 230, by default, installs the Microsoft Windows® Service Control Manager (SCM) management service. As a result, such a robot 230 can open an interactive Windows® session under the local system account and can have the rights of a Windows® service.
[0082] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, it means having the same rights as the user in which the given robot 230 is installed. This feature can also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0083] The robot 230 in some embodiments is divided into a plurality of components, each specialized for a specific automation task. Robot components in some embodiments include, but are not limited to, SCM management robot services, user mode robot services, an executor, an agent, and a command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 is executed). These services are entrusted with managing the qualification information of the robot 230. The console application is launched by the SCM under the local system.
[0084] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host. The user mode robot service can be entrusted with managing the qualification information of the robot 230. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0085] The executor can execute a job given under a Windows® session (i.e., can execute a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows® Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0086] As described above, the fact that the components of the robot 230 are divided helps developers, support users, and computing systems to more easily execute, identify, and track what each component is doing. In this way, special behavior can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 can be made independent of the browser zoom level. In the case of applications marked as not recognizing or intentionally not recognizing DPI, DPI can be disabled in some embodiments.
[0087] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for the execution of the AI / ML model(s) and receive outputs therefrom via the core hyper-automation system 240.
[0088] One or more robots 230 may be listeners, as described above. These listeners can provide information to the core hyper-automation system 240 about what users are doing when they use their computing systems. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0089] An assistant / chatbot 250 can be provided on the user computing system to enable the user to launch an RPA local robot. The assistant can be placed, for example, in the system tray. The chatbot can have a user interface so that the user can view the text of the chatbot. Alternatively, the chatbot can run in the background without a user interface and can listen for the user's utterances using the microphone of the computing system.
[0090] In some embodiments, data labeling can be performed by a user of the computing system that the robot is executing on, or on another computing system that the robot provides information to. For example, if the robot calls an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user can draw a rectangle around the mis-identified or non-identified component and potentially provide text with the correct identification. This information can be provided to the core hyper-automation system 240 and can then be used later for training a new version of the AI / ML model.
[0091] FIG. 3 is an architectural diagram showing an expanded RPA system 300 according to an embodiment of the present invention. In some embodiments, the RPA system 300 can be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The expanded RPA system 300 can be a cloud-based system, an on-premises system, a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes.
[0092] It should be noted that the client side, the server side, or both can include any desired number of computing systems without departing from the scope of the present invention. On the client side, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is executing a process. As shown in FIG. 3, a plurality of business projects can be executed simultaneously. The agent 314 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via a database server 350, an AI / ML server 360, an indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 can be a robot component.
[0093] In some embodiments, the robot represents an association between a machine name and a user name. The robot can manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots can be executed simultaneously, each running in a separate Windows® session using a unique user name. This is referred to as the HD robot described above.
[0094] Agent 314 is also responsible for sending the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. Communication between agent 314 and conductor 340 is, in some embodiments, always initiated by agent 314. In a notification scenario, agent 314 may open a WebSocket channel that is later used by conductor 340 to send commands (e.g., start, stop, etc.) to the robot.
[0095] Listener 330 monitors and records data related to user interactions with the operation of the attended computing system and / or unattended computing system in which listener 330 resides. Listener 330 can be, without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. In fact, in some embodiments, the listener logic is implemented partially or fully via physical hardware.
[0096] On the server side, there are a presentation layer (web application 342, Open Data Protocol (OData) Representational State Transfer (REST) Application Programming Interface (API) endpoint 344, notifications and monitoring 346), a service layer (API implementation / business logic 348), and a persistence layer (database server 350, AI / ML server 360, indexer server 370). The conductor 340 includes the web application 342, the OData REST API endpoint 344, notifications and monitoring 346, and the API implementation / business logic 348. In some embodiments, most of the actions that a user performs at the interface of the conductor 340 (e.g., via the browser 320) are performed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of the present invention. The web application 342 is the visual layer of the server platform. In this embodiment, the web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of the present invention. The user interacts with the web page from the web application 342 via the browser 320 in this embodiment to perform various operations for controlling the conductor 340. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or process, start and stop robots, etc.
[0097] In addition to the web application 342, the conductor 340 also includes a service layer that exposes an OData REST API endpoint 344. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 342 and the agent 314. The agent 314 is, in this embodiment, a supervisor for one or more robots on a client computer.
[0098] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functions. The configuration endpoints may be used, in some embodiments, to define and configure the users, permissions, robots, assets, releases, and environments of the application. The logging REST endpoints may be used to log various information, such as errors, explicit messages sent by the robots, and other environment-specific information. The deployment REST endpoints may be used by the robots to query the version of the package to be executed when a job start command is used in the conductor 340. The queuing REST endpoints may be responsible for the management of queues and queue items, such as adding data to the queue, retrieving transactions from the queue, and setting the status of transactions.
[0099] Monitoring of the REST endpoints may monitor the web application 342 and the agent 314. The notification and monitoring API 346 may be a REST endpoint used for the registration of the agent 314, the distribution of configuration settings to the agent 314, and the sending and receiving of notifications from the server and the agent 314. The notification and monitoring API 346 may use WebSocket communication in some embodiments.
[0100] In some embodiments, the API of the service layer can be accessed through the configuration of an appropriate API access path, for example, based on whether the conductor 340 and the overall hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The API for the conductor 340 can provide custom methods for querying statistics regarding various entities registered with the conductor 340. In some embodiments, each logical resource may be an OData entity. In such an entity, components such as robots, processes, queues, etc. may have properties, relationships, and operations. In some embodiments, the API of the conductor 340 can be consumed by the web application 342 and / or the agent 314 in the following two ways: by obtaining API access information from the conductor 340 or by registering an external application for using the OAuth flow.
[0101] In this embodiment, the persistent layer includes three servers: a database server 350 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides an AI / ML model providing service such as an AI center function), and an indexer server 370. The database server 350 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 342. The database server 350 may manage queues and queue items. In some embodiments, the database server 350 may store (in addition to or instead of the indexer server 370) messages recorded by robots. The database server 350 may also store, for example, process mining, task mining, and / or task capture related data received from a listener 330 installed on the client side. Although no arrow is shown between the listener 330 and the database 350, it should be understood that in some embodiments, the listener 330 can communicate with the database 350 and vice versa. This data can be stored in the form of PDD, images, XAML files, etc. The listener 330 may be configured to eavesdrop on user actions, processes, tasks, and performance metrics on each computing system where the listener 330 resides. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, time, etc.) on its respective computing system and then convert them into a form suitable for being provided to and stored in the database server 350.
[0102] The AI / ML server 360 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can enable non-data scientists to access such capabilities. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of the AI / ML models can be monitored and trained and improved using human-verified data. The AI / ML server 360 can schedule and execute training jobs to train new versions of the AI / ML models.
[0103] The AI / ML server 360 can store data related to AI / ML models and ML packages for configuring various ML skills for users during development. The ML skills used herein are, for example, pre-built and trained ML models for processes that can be used by automation. The AI / ML server 360 can also store data related to document understanding techniques and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, NLP, voice analysis, different types of AI / ML models, etc.
[0104] Optionally in some embodiments, the indexer server 370 stores information recorded by robots and creates an index. In certain embodiments, the indexer server 370 may be disabled via configuration settings. In some embodiments, the indexer server 370 uses ElasticSearch®, an open-source project full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 370 via logging REST endpoint(s), where they are indexed for future use.
[0105] Figure 4 is an architectural diagram illustrating the relationship 400 between designer 410, activities 420, 430, 440, 450, driver 460, API 470, and AI / ML model 480, according to an embodiment of the present invention. As described above, a developer uses designer 410 to develop a workflow to be performed by a robot. Various types of activities may be presented to the developer in some embodiments. Designer 410 may be local or remote to the user's computing system (e.g., accessed via a local web browser that interacts with a VM or remote web server). The workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. User-defined activities 420 and API-driven activities 440 interact with applications via their APIs. User-defined activities 420 and / or AI / ML activities 440 may call one or more AI / ML models 480 that may be located locally and / or remotely to the computing system on which the robot operates in some embodiments.
[0106] In some embodiments, non-text visual components in an image can be identified, which is referred to as CV herein. However, it should be noted that in some embodiments, CV incorporates OCR. CV can be at least partially executed by AI / ML model(s) 480. Some CV activities related to such components can include, but are not limited to, extraction of text from segmented label data using OCR, fuzzy text matching, cropping of segmented label data using ML, comparison of the extracted text in the label data with ground truth data, etc. In some embodiments, the number of activities that can be implemented in user-defined activity 420 can be hundreds or thousands. However, any number and / or type of activities can be used without departing from the scope of the present invention.
[0107] UI automation activity 450 is a subset of special low-level activities described in low-level code that facilitate interaction with the screen. UI automation activity 450 facilitates these interactions via a driver 460 that enables the robot to interact with the desired software. For example, driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, etc. In some embodiments, one or more AI / ML models 480 can be used by UI automation activity 450 to perform interactions with the computing system. In certain embodiments, the AI / ML models 480 can enhance or completely replace the drivers 460. In fact, in certain embodiments, the drivers 460 are not included.
[0108] Driver 460 can interact with the OS at a low level, such as searching for hooks and monitoring keys, via the OS driver 462. Driver 460 may facilitate integration with Chrome (registered trademark), IE (registered trademark), Citrix (registered trademark), SAP (registered trademark), etc. For example, a "click" activity serves the same role in these different applications via driver 460.
[0109] Figure 5 is an architectural diagram showing a computing system 500 configured to implement a smart analyzer or a smart handler according to an embodiment of the present invention. In some embodiments, the computing system 500 may be one or more computing systems depicted and / or described herein. In certain embodiments, the computing system 500 may be part of a hyper-automation system as shown in FIGS. 1 and 2. The computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to the bus 505 for processing information. The processor(s) 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The processor(s) 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one processor(s) 510 can be a neuromorphic circuit including processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0110] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by processor(s) 510. The memory 515 can be composed of random access memory (RAM), read-only memory (ROM), flash memory, cache, a magnetic disk or an optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media can be any available media accessible by processor(s) 510 and can include volatile media, non-volatile media, or both. Also, the media can be removable, non-removable, or both. The computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, the communication device 520 can include one or more antennas that are a single antenna, an array of antennas, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration without departing from the scope of the present invention.
[0111] The processor(s) 510 is further coupled to the display 525 via the bus 505. Without departing from the scope of the present invention, any suitable display device and tactile I / O may be used. The keyboard 530 and cursor control device 535, such as a computer mouse, touchpad, etc., are further coupled to the bus 505 to enable the user to interface with the computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user can interact with the device only via the display 525 and / or a touchpad (not shown). Any type and combination of input devices may be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, the user may interact with the computing system 500 remotely via another computing system communicating with the computing system 500, or the computing system 500 may operate autonomously.
[0112] The memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include a smart analyzer or smart handler module 545 configured to execute all or part of the AI / ML processes described herein or derivatives thereof. The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0113] One skilled in the art will understand that the "system" can be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices, without departing from the scope of the present invention. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. In fact, the methods, systems, and devices disclosed herein may be implemented in a localized form and a distributed form that is consistent with computing techniques including cloud computing systems. The computing system may be part of, or accessible by, a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public cloud or a private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of the present invention.
[0114] It should be noted that some of the system features described herein are presented as modules in order to emphasize implementation independence more. For example, a module can be implemented as a hardware circuit including off-the-shelf semiconductors such as custom very large scale integration (VLSI) circuits or gate arrays, logic chips, transistors, or other discrete components. Also, a module can be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, etc.
[0115] The module can also be at least partially implemented in software for execution by various types of processors. For example, a specified unit of executable code can include one or more physical or logical blocks of computer instructions that may be organized, for example, as objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and can include modules when logically combined and can include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module can be stored on a non-transitory computer-readable medium such as, for example, a hard disk drive, a flash device, RAM, a tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of the present invention.
[0116] In fact, a module of executable code can be a single instruction, or a number of instructions, and can even be distributed among multiple different code segments, between different programs, and among multiple memory devices. Similarly, the operational data can be specified within the module, shown here, and can be embodied in any suitable form and organized within any suitable type of data structure. The operational data can be collected as a single data set or can be distributed in different locations across different storage devices and can exist, at least in part, simply as electronic signals on a system or network.
[0117] Without departing from the scope of the present invention, various types of AI / ML models can be trained and deployed. For example, FIG. 6A shows an example of a neural network 600 trained to complement the operation of a smart analyzer or smart handler according to an embodiment of the present invention. The neural network 600 includes a number of hidden layers. Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) typically have multiple layers, but an SLNN may in some cases have only one or two layers and usually has fewer layers than a DLNN. Typically, the architecture of a neural network includes an input layer, a plurality of intermediate layers, and an output layer, as in the case of the neural network 600.
[0118] Often, DLNNs have many layers (such as 10, 50, 200, etc.), and subsequent layers typically reuse the features from the previous layer to compute more complex and general functions. On the other hand, SLNNs have only a few layers and tend to be trained relatively quickly because expert features are pre-created from raw data samples. However, feature extraction is cumbersome. On the other hand, DLNNs typically do not require expert features but take longer to train and tend to have more layers.
[0119] In either approach, the layers are trained simultaneously on the training set and typically checked for overfitting on a separate cross-validation set. Excellent results are obtained with both techniques, and there is significant enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem being addressed by each neural network.
[0120] Returning to FIG. 6A, the RPA workflow, coding practice rules, APIs, and native OS information (e.g., native calls that exist to simulate mouse clicks, key presses, etc.), logs, etc. are provided as the input layer and fed as input to the J neurons in hidden layer 1. Screenshots, state information of the computing system, information related to publicly available automation, business rules, RPA workflows and / or tasks, initial definitions of automation, process automation documents, etc., but not limited to these, various other inputs are possible. In this example, all of these inputs are fed to each neuron, but not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks that do not depart from the scope of the present invention are possible and can be used individually or in combination.
[0121] The hidden layer 2 receives inputs from the hidden layer 1, the hidden layer 3 receives inputs from the hidden layer 2, and this is done similarly for all hidden layers until the last hidden layer provides its output as the input to the output layer. Although multiple proposals are shown as outputs in this specification, in some embodiments, only a single output proposal is provided. In certain embodiments, the proposals are ranked based on a confidence score.
[0122] Note that the numbers of neurons I, J, K, and L are not necessarily equal. Thus, any desired number of layers can be used for a given layer of the neural network 600 without departing from the scope of the present invention. In fact, in certain embodiments, the types of neurons in a given layer may not all be the same.
[0123] The neural network 600 is trained to assign confidence score(s) to the appropriate output. To reduce inaccurate predictions, in some embodiments, only those results with a confidence score above a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs with a confidence score exceeding this amount may be used and the rest may be ignored.
[0124] A neural network is typically a probabilistic construct that has confidence score(s). This can be a score learned by the AI / ML model based on the frequency with which similar inputs were correctly identified during training. Some common types of confidence scores include decimal numbers between 0 and 1 (which can also be interpreted as a confidence percentage), numerical values between negative infinity and positive infinity, a series of expressions (e.g., "low", "medium", and "high"), etc. Various post - processing calibration techniques such as temperature scaling, batch normalization, weight decay, negative log - likelihood (NLL), etc. can also be used to obtain a more accurate confidence score.
[0125] The "neurons" of a neural network are usually algorithmically implemented as mathematical functions based on the functions of biological neurons. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass the output to the next layer. This activation function can be a non-linear thresholded activity function that does nothing if the value is below the threshold, and responds linearly when the function exceeds the threshold (i.e., Rectified Linear Unit (ReLU) non-linearity). Since actual neurons can have a nearly similar activity function, the sum function and the ReLU function are used in deep learning. Through linear transformation, information can be subtracted, added, etc. Essentially, neurons function as gating functions that pass the output to the next layer governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some neurons.
[0126] JPEG2025097251000002.jpg84144
[0127] JPEG2025097251000003.jpg46144
[0128] JPEG2025097251000004.jpg29132
[0129] In this case, neuron 610 is a single-layer perceptron. However, without departing from the scope of the present invention, any suitable neuron type or combination of neuron types can be used. It should also be noted that the weights of the activation function and / or the range of values of the output value(s) can be different in some embodiments without departing from the scope of the present invention.
[0130] A target, i.e., a "reward function", is often adopted. The reward function guides the search of the state space and attempts to achieve the target (e.g., finding the most accurate answer to a user's query based on relevant indicators) by using both short-term and long-term rewards to explore intermediate transitions and steps. During training, various labeled data are supplied through the neural network 600. When a particular case is successful, the weights of the inputs to the neurons are strengthened, while when a particular case fails, those weights are weakened. A cost function such as the mean squared error (MSE) or gradient descent can be used to make slightly incorrect predictions cost much less than greatly incorrect predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, the data scientist can change the reward function or correct incorrect predictions, etc.
[0131] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "pop up" the hidden layers of a neural network to check how much loss each node is bearing, and then assign low weights to nodes with a high error rate and vice versa to update the weights to minimize the loss. That is, backpropagation enables the data scientist to repeatedly adjust the weights so as to minimize the difference between the actual output and the desired output.
[0132] The algorithm of backpropagation is mathematically based on optimization theory. In supervised learning, training data with known outputs are passed through the neural network, the error is calculated using a cost function from the known target outputs, and this gives the error for backpropagation. The error is calculated at the output, and this error is converted into a correction of the network's weights to minimize the error.
[0133] JPEG2025097251000005.jpg71144
[0134] JPEG2025097251000006.jpg22144
[0135] JPEG2025097251000007.jpg123144
[0136] JPEG2025097251000008.jpg106116
[0137] JPEG2025097251000009.jpg83144
[0138] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., above 97% using an F2 or F4 threshold for detection, about 2000 epochs). This accuracy level can be determined in some embodiments using an F1 score, an F2 score, an F4 score, or any other suitable technique that does not depart from the scope of the present invention. Once trained on the training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps prevent the AI / ML model from becoming “overfitted” such that it performs well on the training data but not on other data.
[0139] In some embodiments, it may not be known what accuracy levels an AI / ML model can achieve. Thus, when the accuracy of the AI / ML model begins to decline when analyzing the evaluation data (i.e., the model performs well on the training data but its performance begins to degrade on the evaluation data), the AI / ML model can undergo additional training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when the accuracy reaches a certain level or when the accuracy of the trained AI / ML model is better than that of an existing deployed AI / ML model. In certain embodiments, a set of trained AI / ML models can be used to accomplish a task. For example, one model can be trained to recognize images, another can be trained to recognize text, and yet another can be trained to recognize semantic and / or ontology-related aspects, etc.
[0140] In some embodiments, a transformer neural network such as SentenceTransformers™, a state-of-the-art Python™ framework for sentence, text, and image embedding, can be used. Such a transformer neural network learns associations of words and phrases with both high and low scores. This trains the AI / ML model to determine what is close to the input and what is not, respectively. Instead of using only word / phrase pairs, the transformer neural network may also use field length and field type.
[0141] In some embodiments, NLP technologies such as word2vec, BERT, GPT-3, ChatGPT, and other LLMs can be used to facilitate semantic understanding and provide more accurate and human-like answers as described above. Other technologies such as clustering algorithms can be used to find similarities between groups of elements. Clustering algorithms can include, but are not limited to, density-based algorithms, distribution-based algorithms, centroid-based algorithms, and hierarchical-based algorithms. Such as the K-means clustering algorithm, DBSCAN clustering algorithm, Gaussian mixture model (GMM) algorithm, balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm, etc. Such techniques can also be useful for classification.
[0142] FIG. 7 is a flowchart showing a process 700 for training an AI / ML model(s) according to an embodiment of the present invention. In some embodiments, the AI / ML model(s) may be a generative AI model as described above. The neural network architecture of an AI / ML model typically includes multiple layers of neurons including an input layer, an output layer, and hidden layers. For example, refer to FIGS. 6A and 6B. The intermediate hidden layers process the input data and generate an intermediate representation of the input used for the generation of the output. These hidden layers can include various types of neurons such as convolutional neurons, recurrent neurons, and / or transformer neurons.
[0143] The training process begins at 710 by providing an RPA workflow, coding practice rules, API and native OS information, and labeled or unlabeled logs. The AI / ML model is then trained at 720 over multiple epochs, and the results are reviewed at 730. Various types of AI / ML models can be used, but LLM and other generative AI models are typically trained using a process called "supervised learning" as described above. Supervised learning involves providing the model with a large dataset, which the model uses to learn the relationship between inputs and outputs. During the training process, the model adjusts the weights and biases of the neurons within the neural network to minimize the difference between the predicted output and the actual output within the training dataset.
[0144] One aspect of the model in some embodiments is the use of transfer learning. For example, transfer learning can utilize a pre-trained model such as ChatGPT that is fine-tuned for a specific task or domain at step 720. This allows the model to leverage the knowledge already learned from the pre-training phase and adapt it to a specific application through the training phase at step 720.
[0145] The pre-training phase involves training the model based on an initial set of training data that may be more general. In this phase, the model learns the relationships within the data. In the fine-tuning phase (e.g., in some embodiments, when a pre-trained model is used as the initial basis for the final model, in addition to or instead of the initial training phase, executed during step 720), the pre-trained model is adapted to a specific task or domain by training the model with a smaller dataset specific to the task. For example, in some embodiments, the model may focus on a particular type(s) of data source. This can help the model more accurately identify the data elements within it than a generative AI model pre-trained alone. Through fine-tuning, the model can learn the nuances of the source, such as a particular vocabulary and syntax, particular graphical characteristics, particular data formats, etc., without requiring as much data as would be needed to train the model from scratch. By leveraging the knowledge learned in the pre-training phase, the fine-tuned model can achieve state-of-the-art performance on a particular task with relatively little additional training data.
[0146] If the AI / ML model does not meet the desired confidence threshold at 740, the training data is supplemented and / or the reward function is modified at 750 to help the AI / ML model better achieve its objective, and the process returns to step 720. If the AI / ML model meets the confidence threshold at 740, the AI / ML model is tested against the evaluation data at 760 to confirm that the AI / ML model generalizes well and does not overfit to the training data. The evaluation data includes information that the AI / ML model has not processed before. If the confidence threshold is met for the evaluation data at 770, the AI / ML model is deployed at 780. Otherwise, the process returns to step 750 and the AI / ML model is further trained.
[0147] Figure 8 shows an AI / ML model 800 of the cognitive AI layer according to an embodiment of the present invention. In some embodiments, the model of the cognitive AI layer 800 can be trained using the process 700 of FIG. 7. The generative AI model 810 provides results to other cognitive AI layer models in a serial configuration 840, a parallel configuration 842, or a combination 844 of serial and parallel configurations (plural possible). For example, the generative AI model can generate code, provide a semantic association between texts on the screen, determine actions to address RPA workflow or runtime automation issues, and the like. The CV model 820 and the OCR model also provide the detected graphical elements and the recognized text to the generative AI model 810 and other cognitive AI layer models within the configurations 840, 842, or 844, respectively.
[0148] Other cognitive AI layer models of configurations 840, 842, or 844 use the outputs from the generative AI model 810, the CV model 820, and / or the OCR model 830 to provide intelligent analysis capabilities to a smart analyzer or a smart handler. The cognitive AI layer model 800 can facilitate an understanding of the RPA workflow or automation purposes, how automation has been performed so far, previous activities within the RPA workflow and / or how the logical flow of the RPA workflow affects a particular activity, and what the best course(s) of action are regarding the repair or improvement of the RPA workflow or automation. The output from other cognitive AI layer models of configurations 840, 842, or 844 (i.e., result 850) is for proposed error corrections, efficiency improvements, improvements in the execution speed of the RPA workflow, reduction in the consumption of processing resources during RPA workflow execution, and / or reduction in memory consumption during RPA workflow execution (e.g., reduction in cloud hosting costs) for the RPA workflow or automation, proposals for what to change in future versions of the RPA workflow or automation to avoid the same problems, code and / or automatically generated blocks of automation to understand the intent and reason for failure to recover during execution or to prevent errors from occurring in the first place during the design of the RPA workflow, rollback of previously executed steps in the automation to undo a failure, a confidence score for each, any combination thereof, etc. In some embodiments, these operations can be automatically performed by a smart analyzer or a smart handler.
[0149] Figure 9 is a flowchart showing a process 900 of a smart analyzer according to an embodiment of the present invention. The process begins at 910 with the smart analyzer monitoring the development of an RPA workflow in an RPA designer application. However, in some embodiments, the execution of test cases and changes to the target application can be monitored for the purpose of proposing ways to adapt the test cases.
[0150] The smart analyzer is part of an RPA designer application, or an RPA robot, or another software process that monitors RPA workflow development. If automation related to the RPA workflow already exists, step 910 may include reading one or more logs containing information on how the automation was executed at runtime. The log information includes the timestamps of the execution of each activity of the RPA workflow, the values of variables within the RPA workflow, or combinations thereof.
[0151] The Smart Analyzer provides RPA workflow development information to the Cognitive AI layer at 920. This information includes logs regarding activities and their parameters within the RPA workflow, the relationships between activities in the RPA workflow, the most frequently used RPA workflows and activities, and how users compose the most frequently used RPA workflows and activities (where the RPA workflow may include common incremental components), how each automation was executed during operation (e.g., if a particular application slowed down after a button was clicked, whether a pause needs to be introduced before attempting the next activity or other aspects of the operation of a third - party application), how the RPA workflow is executed at runtime (e.g., if a button click causes a particular application to slow down, whether a pause needs to be introduced before attempting the next activity or other aspects of the operation of a third - party application), available APIs and native OS functions related to RPA workflow activities, telemetry data on how users are using the RPA designer application, and may include but is not limited to any combination of these. For example, the Cognitive AI layer can learn that users tend to notify in a particular channel of the RPA workflow (e.g., general announcements in Slack (registered trademark)), that users tend to input data into a particular spreadsheet or database table, that users tend to prefer one communication method over another, etc. The Cognitive AI layer may be part of the Smart Analyzer, a separate process running on the same computing system, or remotely located from the computing system on which the RPA workflow is developed (e.g., hosted by a cloud service provider). In some embodiments, the Cognitive AI layer is trained based on automations built by other users, finds patterns in the RPA workflow code from the automations, and makes suggestions for the RPA workflow based on the learned patterns.
[0152] In some embodiments, a proposal can be made regarding when a change occurs to an application that an RPA robot performing an RPA workflow is intended to interact with. For example, if the automation interacts with a Salesforce (registered trademark) screen and a new field is added to the screen in a new version, it may propose an adaptation to the RPA workflow or automatically modify the form. Thus, in some embodiments, the RPA designer application can be a presentation layer creator where forms, applications, workflows, other designs, etc. can be utilized in one application.
[0153] Next, the smart analyzer receives the output from the cognitive AI layer at 930. The output can include, but is not limited to, indicating that no change should be made, proposed error corrections and / or improvement proposals regarding the RPA workflow, proposals for changes to the RPA workflow or automation to improve efficiency, proposals for deletion and / or addition of activities, proposals for what should be changed in future versions of the RPA workflow to avoid the same problem, proposals for security enhancements, blocks of auto-generated code, respective trust scores, any combination of these, etc. If there are no changes proposed by the cognitive AI layer at 940, the process returns to step 910 and the smart analyzer continues to monitor the RPA workflow development. However, if changes are proposed by the cognitive AI layer at 940, at 950 the smart analyzer provides the proposals to the user or automatically makes these changes. For example, the smart analyzer can change the RPA workflow itself or send instructions to the RPA designer application to do so.
[0154] Figure 10 is a flowchart showing process 1000 of a smart handler according to an embodiment of the present invention. The smart handler is an RPA robot or another software process that runs on a computing system on which automation is being performed. The process begins at 1010 with the smart handler analyzing and evaluating RPA automation code (or test automation or application code for another purpose) at runtime. This can be done by invoking the cognitive AI layer and / or deterministic logic can be used (e.g., logic to find and separate nested loops, logic to remove assignments of values to uninitialized variables, or logic to initialize these variables before assignment of values, etc.). Step 1010 can include reading one or more logs containing information on how the automation was executed at runtime. The log information can include timestamps of the execution of each activity of the RPA workflow, values of variables within the RPA workflow, or combinations thereof. Thereafter, the smart handler warns the user at 1020 if such problems are found in the RPA automation or automatically corrects the problems.
[0155] Next, the smart handler monitors the RPA automation at runtime at 1030. If an error and / or performance problem(s) occur in the automation at 1040, information for dealing with the error and / or performance problem(s) is provided to the cognitive AI layer at 1050. The information can include, but is not limited to, the automation code, screenshots (if any), the RPA workflow related to the automation, execution logs from the computing system, a list of processes currently running on the computing system, current internet connection speed information, the initial definition of the automation, the process automation document, design-time information, the RPA automation language, the screen ontology, the technical representation of the current screen, boundaries and / or rules that prevent the RPA robot from performing certain actions and / or accessing certain information, etc.
[0156] The smart handler receives, at 1060, outputs from the cognitive AI layer, such as suggestions on how to handle errors and / or performance issues (if any). Next, at 1070, the smart handler automatically attempts to self-heal the automation, or provides suggestions to the user, receives a selection from the user, and then attempts self-healing. For example, the smart handler may attempt to revert the actions performed by the automation, change the code of the automation to avoid an obstacle, or change the code to improve performance. If successful at 1080, the automation continues its execution (e.g., by an RPA robot), and the smart handler continues to monitor the automation at 1030. If not successful at 1080, the smart handler fails the automation (e.g., terminates the RPA process and instructs the RPA robot to stop executing the automation), and notifies the user at 1090. Next, the RPA workflow associated with the automation can be provided to the RPA design application, and the RPA workflow can be repaired (e.g., by a smart analyzer).
[0157] The process steps performed in FIGS. 7, 9, and 10 may be performed by a computer program that encodes instructions to a processor (s) to perform at least a portion of the process (es) described in FIGS. 7, 9, and 10, in accordance with embodiments of the present invention. The computer program may be stored on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, a tape, and / or any other such medium or combination of media used to store data. The computer program may include encoded instructions for controlling a processor (s) of a computing system (e.g., the processor (s) 510 of the computing system 500 of FIG. 5) to implement all or a portion of the process steps described in FIGS. 7, 9, and 10, which may also be stored on a computer-readable medium.
[0158] The computer program may be implemented in hardware, software, or a hybrid implementation. The computer program may be composed of modules that communicate operably with each other and are designed to send information or instructions to a display. The computer program may be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0159] It will be readily understood that the components of the various embodiments of the present invention may be arranged and designed in a variety of different configurations, as generally described and illustrated herein. Accordingly, the detailed description of the embodiments of the present invention as represented in the accompanying figures is not intended to limit the scope of the present invention as claimed, but rather represents only selected embodiments of the present invention.
[0160] The features, structures, or characteristics of the invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments", "some embodiments", or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment of the invention. Thus, the appearances of "in certain embodiments", "in some embodiments", "in other embodiments", or similar language throughout this specification are not necessarily referring to the same group of all embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0161] It should be noted that references throughout this specification to features, advantages, or similar language do not mean that all features and advantages achievable in the invention should be in any single embodiment of the invention or in any embodiment of the invention. Rather, language referring to features and advantages is understood to mean that the particular features, advantages, or characteristics described in connection with the embodiments are included in at least one embodiment of the invention. Thus, discussions of features and advantages throughout this specification, as well as similar language, can refer to the same embodiment, but not necessarily.
[0162] Furthermore, the described features, advantages, and characteristics of the invention may be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the invention may be practiced without the specific features or advantages of one or more of the particular embodiments. In other instances, additional features and advantages may be recognized in certain embodiments but may not exist in all embodiments of the invention.
[0163] Those having ordinary skill in the art will readily understand that the present invention as described above can be implemented using steps in a different order and / or using hardware elements of a different configuration than those disclosed. Accordingly, although the present invention has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations may become apparent while remaining within the spirit and scope of the present invention. Therefore, reference should be made to the appended claims to determine the scope of the present invention.
Claims
1. 1. A non-transitory computer readable medium having stored thereon a computer program for a smart analyzer, the computer program comprising: Oversee the development of Robotic Process Automation (RPA) workflows in the RPA Designer application; providing information regarding the RPA workflow development to a cognitive artificial intelligence (AI) layer; receiving output from the cognitive AI layer including one or more suggestions for repairing the RPA workflow, improving performance of the RPA workflow, or both; A non-transitory computer-readable medium configured to provide the one or more suggestions from a cognitive AI model to a user of the RPA designer application, automatically modify the RPA workflow using the one or more suggestions from the cognitive AI model, or both.
2. 2. The non-transitory computer readable medium of claim 1, wherein the information regarding the RPA workflow development provided to the cognitive AI layer includes activities within the RPA workflow and their parameters, relationships between activities of the RPA workflow, most frequently used RPA workflows and activities and how users configure the most frequently used RPA workflows and activities, logs containing information about how each automation was performed during production, available application programming interface (API) and / or native operating system (OS) functions related to RPA workflow activities, telemetry data regarding how users are using the RPA designer application, or combinations thereof.
3. 2. The non-transitory computer readable medium of claim 1, wherein the output from the cognitive AI layer includes one or more proposed error corrections and / or improvements for the RPA workflow, one or more proposed modifications to the RPA workflow or respective automation to improve the efficiency of the RPA workflow, execute faster, consume less processing resources and / or consume less memory when performing the RPA workflow, one or more suggestions for removing activities and / or adding activities, one or more suggestions on what to change in future versions of the RPA workflow to avoid the same problem, one or more suggestions for improving security, one or more automatically generated code blocks, one or more respective confidence scores, or any combination thereof.
4. The computer program further comprises:
2. The non-transitory computer-readable medium of claim 1, configured to continue monitoring further development of the RPA workflow by the user after providing the one or more suggestions from the cognitive AI model to the user of the RPA designer application, after automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both.
5. 2. The non-transitory computer-readable medium of claim 1, wherein providing the one or more suggestions from the cognitive AI model to the user of the RPA designer application, automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both, comprises sending instructions to the RPA designer application to modify the RPA workflow.
6. 10. The non-transitory computer readable medium of claim 1, wherein the smart analyzer is part of an RPA designer application.
7. 10. The non-transitory computer readable medium of claim 1, wherein the smart analyzer is an RPA robot.
8. The cognitive AI layer is a generative AI model configured to facilitate understanding of the intent of the RPA workflow, how previous activities in the RPA workflow and / or the logical flow of the RPA workflow affect a particular activity, one or more best courses of action to take to repair or improve the RPA workflow, or any combination thereof; and one or more other AI / ML models configured to use output from the generative AI model to provide intelligent analysis capabilities to a smart analyzer.
9. 9. The non-transitory computer-readable medium of claim 8, wherein the generative AI model is configured to generate code, provide semantic associations between on-screen text, determine actions to address problems in an RPA workflow or runtime automation, or combinations thereof.
10. 2. The non-transitory computer-readable medium of claim 1 , wherein the one or more suggestions provided by the cognitive AI layer include suggesting to further sub-divide a workflow instead of using loops and / or suggesting to separate nested loops.
11. 10. The non-transitory computer-readable medium of claim 1, wherein the RPA workflow is a previously created workflow.
12. 2. The non-transitory computer-readable medium of claim 1, wherein automatically modifying the RPA workflow includes replacing one or more outdated dependencies and code portions that pose potential threats and / or blocking one or more malicious, unsafe, or unauthorized websites.
13. The RPA workflow is associated with an existing automation, and the computer program further comprises:
2. The non-transitory computer-readable medium of claim 1, configured to read one or more logs at runtime containing information about how the automation was performed, where the log information includes a timestamp of the completion of each activity of the RPA workflow, values of variables in the RPA workflow, or a combination thereof.
14. 2. The non-transitory computer-readable medium of claim 1, wherein the cognitive AI layer is trained based on automations built by other users to find patterns in RPA workflow code from the automations and to make suggestions about the RPA workflow based on the trained patterns.
15. a memory for storing computer program instructions; and at least one processor configured to execute the computer program instructions, the computer program instructions comprising: Oversee the development of Robotic Process Automation (RPA) workflows in the RPA Designer application; providing information regarding the RPA workflow development to a cognitive artificial intelligence (AI) layer; receiving output from the cognitive AI layer including one or more suggestions for repairing the RPA workflow, improving performance of the RPA workflow, or both; configured to provide the one or more suggestions from a cognitive AI model to a user of the RPA designer application, or automatically modify the RPA workflow using the one or more suggestions from the cognitive AI model, or both; the cognitive AI layer comprises a generative AI model configured to facilitate understanding of the intent of the RPA workflow, how previous activities within the RPA workflow and / or the logical flow of the RPA workflow affect a particular activity, one or more best courses of action to take to repair or improve the RPA workflow, or any combination thereof; The generative AI layer comprises one or more computing systems, the one or more other AI / ML models configured to use output from the generative AI models to provide intelligent analytical capabilities to the smart analyzer.
16. 16. The one or more computing systems of claim 15, wherein the information regarding the RPA workflow development provided to the cognitive AI layer includes activities within the RPA workflow and their parameters, relationships between activities of the RPA workflow, most frequently used RPA workflows and activities and how users configure the most frequently used RPA workflows and activities, logs containing information about how each automation was performed during production, available application programming interface (API) and / or native operating system (OS) functions related to RPA workflow activities, telemetry data regarding how users are using the RPA designer application, or combinations thereof.
17. 16. The one or more computing systems of claim 15, wherein the output from the cognitive AI layer includes one or more proposed error corrections and / or improvements for the RPA workflow, one or more proposed modifications to the RPA workflow or respective automation to improve the efficiency of the RPA workflow, execute faster, consume less processing resources and / or consume less memory when performing the RPA workflow, one or more suggestions for removing activities and / or adding activities, one or more suggestions on what to change in future versions of the RPA workflow to avoid the same problem, one or more suggestions for improving security, one or more automatically generated code blocks, one or more respective confidence scores, or any combination thereof.
18. The computer program instructions further include causing the at least one processor to:
16. The one or more computing systems of claim 15, configured to continue monitoring further development of the RPA workflow by the user after providing the one or more suggestions from the cognitive AI model to the user of the RPA designer application, after automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both.
19. 16. The one or more computing systems of claim 15, wherein providing the one or more suggestions from the cognitive AI model to the user of the RPA designer application, automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both, comprises sending instructions to the RPA designer application to modify the RPA workflow.
20. 16. The one or more computing systems of claim 15, wherein the generative AI model is configured to generate code, provide semantic associations between on-screen text, determine actions to address problems in an RPA workflow or runtime automation, or any combination thereof.
21. 16. The one or more computing systems of claim 15, wherein the one or more suggestions provided by the cognitive AI layer include suggesting to further sub-divide a workflow instead of using loops and / or suggesting to separate nested loops.
22. 16. The one or more computing systems of claim 15, wherein automatically modifying the RPA workflow includes replacing one or more outdated dependencies and code portions that pose potential threats and / or blocking one or more malicious, unsafe, or unauthorized websites.
23. The RPA workflow is associated with an existing automation, and the computer program instructions further include causing the at least one processor to:
16. The one or more computing systems of claim 15, configured to read one or more logs at runtime containing information about how the automation was performed, where the log information includes a timestamp of the completion of each activity in the RPA workflow, values of variables in the RPA workflow, or a combination thereof.
24. 16. The one or more computing systems of claim 15, wherein the cognitive AI layer is trained based on automations built by other users to find patterns in RPA workflow code from the automations and to make suggestions about the RPA workflow based on the trained patterns.
25. monitoring, by a computing system, the development of a robotic process automation (RPA) workflow in an RPA designer application; providing, by the computing system, information regarding the RPA workflow development to a cognitive artificial intelligence (AI) layer; receiving, by the computing system, output from the cognitive AI layer including one or more suggestions for repairing the RPA workflow, improving performance of the RPA workflow, or both; A computer-implemented method for a smart analyzer comprising: providing, by the computing system, the one or more suggestions from a cognitive AI model to a user of the RPA designer application, automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both.
26. 26. The computer-implemented method of claim 25, wherein the information regarding the RPA workflow development provided to the cognitive AI layer includes activities within the RPA workflow and their parameters, relationships between activities of the RPA workflow, most frequently used RPA workflows and activities and how users configure the most frequently used RPA workflows and activities, logs containing information about how each automation was performed during production, available application programming interface (API) and / or native operating system (OS) capabilities related to RPA workflow activities, telemetry data regarding how users are using the RPA designer application, or combinations thereof.
27. 26. The computer-implemented method of claim 25, wherein the output from the cognitive AI layer includes one or more suggested error corrections and / or improvements for the RPA workflow, one or more proposed modifications to the RPA workflow or respective automation to improve the efficiency of the RPA workflow, execute faster, consume less processing resources and / or consume less memory when performing the RPA workflow, one or more suggestions for removing activities and / or adding activities, one or more suggestions on what to change in future versions of the RPA workflow to avoid the same problem, one or more suggestions for improving security, one or more automatically generated code blocks, one or more respective confidence scores, or any combination thereof.
28. moreover, 26. The computer-implemented method of claim 25, further comprising: after providing, by the computing system, the one or more suggestions from the cognitive AI model to the user of the RPA designer application, automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both, continuing to monitor further development of the RPA workflow by the user.
29. 26. The computer-implemented method of claim 25, wherein providing the one or more suggestions from the cognitive AI model to the user of the RPA designer application, automatically modifying the RPA workflow using the one or more suggestions from the cognitive AI model, or both, comprises sending instructions to the RPA designer application to modify the RPA workflow.
30. The cognitive AI layer is a generative AI model configured to facilitate understanding of the intent of the RPA workflow, how previous activities in the RPA workflow and / or the logical flow of the RPA workflow affect a particular activity, one or more best courses of action to take to repair or improve the RPA workflow, or any combination thereof; and one or more other AI / ML models configured to use output from the generative AI model to provide intelligent analysis capabilities to the smart analyzer; 26. The computer-implemented method of claim 25, wherein the generative AI model is configured to generate code, provide semantic associations between on-screen text, determine actions to address problems in an RPA workflow or runtime automation, or any combination thereof.
31. 26. The computer-implemented method of claim 25, wherein the one or more suggestions provided by the cognitive AI layer include suggesting to further sub-divide a workflow instead of using loops and / or suggesting to separate nested loops.
32. 26. The computer-implemented method of claim 25, wherein automatically modifying the RPA workflow includes replacing one or more outdated dependencies and code portions that pose potential threats and / or blocking one or more malicious, unsafe, or unauthorized websites.
33. The RPA workflow is related to an existing automation and further comprises:
26. The computer-implemented method of claim 25, further comprising reading, by the computing system, one or more logs containing information about how the automation was performed at runtime, wherein log information includes timestamps of performance of each activity in the RPA workflow, values of variables in the RPA workflow, or a combination thereof.
34. 26. The computer-implemented method of claim 25, wherein the cognitive AI layer is trained based on automations built by other users to find patterns in RPA workflow code from the automations and to make suggestions about the RPA workflow based on the trained patterns.
Citation Information
Cited By
Information processing device, information processing method, and program
JP7926815B1