Automatic code generation for robotic process automation
By automatically generating code for RPA workflows using the cognitive AI layer, the problem of developers needing programming knowledge and specific programming language knowledge is solved, and more efficient and accurate RPA workflow development is achieved.
Patent Information
- Application Number
- CN202410795145.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-06-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the development of robotic process automation (RPA) workflows, developers need to have programming knowledge and knowledge of specific programming languages, resulting in the introduction of language barriers and errors.
The cognitive AI layer is used to automatically generate computer program code, and by converting the input source into computer program code, it reduces dependence on programming languages.
Eliminate language barriers due to insufficient programming experience or unfamiliarity with programming languages, improving the development efficiency and accuracy of RPA workflows.
Smart Images

Figure CN120179252A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to artificial intelligence (AI), and more particularly to automatic code generation for robotic process automation (RPA) that uses a cognitive AI layer to convert an input source into computer program code. Background Art
[0002] During the development of an RPA workflow, a developer may need to build expressions or create and call code snippets written in Python, Visual Basic (VB), C#, C++, Java, C, etc. However, this requires programming knowledge and knowledge of the specific programming language to be used. As a result, only relatively advanced users can take advantage of these features, and even experienced developers may inadvertently introduce errors into the RPA workflow logic. Thus, improvements and / or alternative approaches to RPA workflow development can be beneficial. Summary of the Invention
[0003] Certain embodiments of the present invention can provide solutions to problems and needs in the art that have not been fully identified, understood, or solved by current RPA technologies. For example, some embodiments of the present invention relate to automatic code generation for RPA that uses a cognitive AI layer to convert an input source into computer program code.
[0004] In an embodiment, a non-transitory computer-readable medium stores a computer program. The computer program is configured to cause at least one processor to provide source data as input to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer program is also configured to cause at least one processor to receive an output including automatically generated code from the cognitive AI layer and implement the automatically generated code in an RPA workflow.
[0005] In another embodiment, one or more computing systems include a memory storing computer program instructions and at least one processor configured to execute the stored computer program instructions. The computer program instructions are configured to cause at least one processor to provide source data as input to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer program instructions are also configured to cause at least one processor to receive an output including automatically generated code from the cognitive AI layer and implement the automatically generated code in an RPA workflow. The source data includes natural language text, a recording of a user's voice, a video recording of a user's actions on a computing system, a document, pseudocode for a desired task, a drawing in a desired user interface or form, a drawing of a process, a process document, or a process description, an RPA workflow, or any combination thereof.
[0006] In another embodiment, a computer-implemented method for performing automated code generation includes providing, by a computing system, source data as input to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer-implemented method also includes receiving, by the computing system, an output from the cognitive AI layer that includes the automatically generated code. The computer-implemented method further includes implementing, by the computing system, the automatically generated code in an RPA workflow. The source data includes natural language text, recordings of user speech, video recordings of user actions on the computing system, documents, pseudocode for a desired task, drawings in a desired user interface or form, drawings of processes, process documents or descriptions, RPA workflows, or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] For the advantages of certain embodiments of the present invention to be readily understood, the present invention, briefly described above, will be described more specifically with reference to specific embodiments shown in the accompanying drawings. While it should be understood that these drawings only depict typical embodiments of the present invention and are not to be considered limiting of its scope, the present invention will be described and explained with additional specificity and detail by use of the accompanying drawings, in which:
[0008] Figure 1 is an architectural diagram showing a hyper-automation system according to an embodiment of the present invention;
[0009] Figure 2 is an architectural diagram showing an RPA system according to an embodiment of the present invention;
[0010] Figure 3 is an architectural diagram showing a deployed RPA system according to an embodiment of the present invention;
[0011] Figure 4 is an architectural diagram showing the relationship between a designer, activities, and a driver according to an embodiment of the present invention;
[0012] Figure 5 is an architectural diagram showing a computing system configured to perform automated code generation for RPA according to an embodiment of the present invention;
[0013] Figure 6A shows an example of a neural network that has been trained to supplement automated code generation for RPA according to an embodiment of the present invention;
[0014] Figure 6B shows an example of a neuron according to an embodiment of the present invention;
[0015] Figure 7 is a flowchart showing a process for training an AI / ML model(s) according to an embodiment of the present invention;
[0016] Figures 8A to 8C is a screenshot showing an automated code generation interface configured to analyze natural language input from a user and provide prompts according to an embodiment of the present invention;
[0017] Figure 9 shows an AI / ML model of a cognitive AI layer according to an embodiment of the present invention;
[0018] Figure 10 is a flowchart showing a process for automated code generation for RPA according to an embodiment of the present invention;
[0019] Figure 11 is a flowchart showing a process for automated document generation using a cognitive AI layer according to an embodiment of the present invention; and
[0020] Figure 12 is a flowchart showing a process for generating automation from natural language text according to an embodiment of the present invention.
[0021] Unless otherwise indicated, like reference numerals consistently represent corresponding features in the drawings. Detailed Description
[0022] Some embodiments relate to automated code generation for RPA that uses a cognitive AI layer to automatically generate computer program code, text, or other items based on an input source. Such embodiments can eliminate language barriers caused by RPA developers having no programming experience or programmers not knowing one or more programming languages. To eliminate these language barriers, some embodiments convert text or another input source into expressions or code snippets in one or more programming languages. An expression is a syntactic entity of a programming language that can be evaluated to determine its value. An expression is a combination of one or more constants, variables, functions, and operators that the programming language interprets according to its specific precedence and associativity rules and computes to produce another value. A code snippet is a relatively small region of reusable source code, machine code, or text.
[0023] Regarding expressions and code snippets, sometimes RPA developers need to adapt RPA workflow code for various purposes. Although RPA workflow development can be a relatively low-code process compared to other types of programming, there is a certain degree of customization achieved through code in expressions and snippets. Generally, this code does not fundamentally change the scope of the RPA workflow but enriches it.
[0024] The RPA designer application can support various programming languages. For example, UiPath Studio TMCurrently, Python, VB, and C# are supported. To utilize these features, RPA developers typically have to understand the corresponding programming languages. However, in some embodiments, generative AI models such as large language models (LLMs) can be used to create expressions or code snippets that perform specific actions in response to user requests or in response to interpreting input sources. Such embodiments can further enrich the encoded workflows or encoded test cases based on the user requests or input sources. The encoded workflows or test cases are the ways of writing code for RPA robots. Developers can use C# code or code in some other language instead of the low-code visual RPA workflow designer. This may put the full power of.NET in the hands of RPA developers, with the platform advantages of the RPA designer application (RPA robots, orchestrators, insights, test managers, etc.).
[0025] In response to user requests or input sources, some embodiments generate expressions that are part of a workflow activity or activity configuration (e.g., an expression can be assigned to export a certain activity parameter value). Certain embodiments can configure code in different call parts or segments. For example, some embodiments can generate activities from a UI object library or some other source with a certain configuration, or can generate entirely new code. Thus, the cognitive AI layer can generate expressions and code snippets that are incorporated into RPA workflows.
[0026] The input source can take various forms. For example, a user can input a natural language statement, provide a source document (e.g., a spreadsheet file, a JavaScript Object Notation (JSON) file, an Extensible Application Markup Language (XAML) file, an Extensible Markup Language (XML) file, a Hypertext Markup Language (HTML) file, a document, a Portable Document Format (PDF) file, a scanned image of a physical document, etc.), write pseudocode for a desired task, and so on. In some embodiments, the cognitive AI layer is capable of suggesting RPA automation solely based on this input.
[0027] Consider the case where the PDF is an invoice. The cognitive AI layer can determine that the type of the PDF is an invoice and understand that for such an invoice, the user typically wants to extract data from it whenever the invoice is received. The cognitive AI layer can then suggest automation to the user and, if needed, automatically generate the automation. This can involve creating an RPA workflow with appropriate activities, such as opening the invoice, performing optical character recognition (OCR) on it if the invoice is scanned, e.g., extracting text and numbers from it, and saving the information to the billing system.
[0028] Any suitable information can be used as a source for such automation. For example, the cognitive AI layer can use process definition documents (PDDs), information technology (IT) strategies, approval interfaces, etc. as sources. Then, if needed, the source information can be used to generate high-level automation that the user can further refine. In some embodiments, the cognitive AI layer can add security and / or compliance rules to the generated RPA workflow to ensure that the workflow complies with regulations and / or policies. In some embodiments, task mining outputs (e.g., what applications the user is using, what information is being input into those applications, keystrokes, mouse clicks and locations, in graphical elements in the user interface (UI), which graphical elements are active elements at different times, etc.) can be used to help generate an RPA workflow. For example, see U.S. Patent Application Publication No. 2022 / 0113991 and U.S. Patent No. 11,301,269, the entire contents of which are incorporated herein by reference.
[0029] In some embodiments, the cognitive AI layer can determine what the user is doing based on task mining outputs and record the process. For example, the LLM of the cognitive AI layer can be used to generate a human-like description of the actions the user is performing, potentially including screenshots. This information can then be provided as a manual to other users who want to perform the corresponding task.
[0030] In some embodiments, automation can be created from such records. By understanding what the user is doing and mapping the actions to RPA workflow activities, the cognitive AI layer can automatically create an RPA workflow based on the actions taken in the record. This can allow for the automation of user tasks without even a cursory understanding of RPA on the part of the user.
[0031] In some embodiments, a diagram representing a screen design can be used to automatically generate an associated application. Consider the case of a diagram of a UI for a screen design, such as a form for collecting certain data and submitting it to a customer relationship management (CRM) system. The cognitive AI layer can use a CV model to understand the graphical elements and text present in the screen, and a generative AI model can create a software application that includes these graphical elements and the ability to submit to the CRM system, such as capturing customer information and submitting this information to Thus, the user can provide the form, and the cognitive AI layer can build a UI for the form. These embodiments can thus essentially convert a description of graphical elements and text into code. In some embodiments, the user can submit a diagram of the desired process to automate. Then the high-level steps of the process can be generated as an RPA workflow. For example, a generative AI model can be trained to understand the text in the steps and the associations between them, as indicated by connectors.
[0032]
[0033] When developing automation using RPA tools (such as UiPath's Form Builder TM ), importing data sources may require manual customization and a lengthy process. Additionally, setting up dynamic form elements (such as dropdown lists) is a time-consuming process. Moreover, developers with professional coding skills and knowledge need to incorporate complex functions when developing automation. For example, writing code using activities such as <invoke code> or combining complex functions requires in-depth knowledge and professional coding skills.
[0034] Accordingly, some embodiments utilize natural language processing (NLP) to automate form building and code generation in automation development. A natural language description is provided by the user (e.g., using the interface 800 shown in Figures 8A to 8C ). Then the LLM can be used to process the description and understand the user's intent. As a result, a form is automatically built or other code is automatically generated based on the user description.
[0035] The natural language description can be used to define the appearance of the user interface. For example, the user can be prompted in the RPA designer application with a natural language query description to generate a form for an insurance claim with the policyholder's name, policy number, claim details, and accident date. Next, the RPA designer application automatically generates the form, populates the form with the requested fields, and configures the required validations and correct input types for the fields.
[0036] In some embodiments, the user is also able to add new fields to the generated form. For example, using the example in the previous paragraph, the user can provide a natural language description to add a field for the estimated claim amount and supporting documents. Accordingly, the tool adds or updates the requested information in the form. Additionally, a fully functional form can be used to build various desired types of automation. For example, the input for the automation can be taken from the form, and an email with policy details can be sent to the user.
[0037] The LLM(s) of some embodiments convert the user description into machine-executable code. For example, an invoke code activity can be used, and a user description in natural language can be provided to "calculate the total number of days since an event occurred based on the date of the event occurrence, and if it exceeds 10 days, a review group is required." Moreover, in some embodiments, code can be generated in any suitable programming language based on the user description. The generated code can be adjusted by the LLM(s) to conform to the given programming language.
[0038] The embodiments disclosed herein can provide various advantages over existing automation technologies. The manual work involved in form building and code generation for other purposes in RPA can be reduced, thus accelerating process implementation and reducing the need for coding expertise. This can increase the number of RPA projects that an organization can create and manage, and the development time per project can be reduced. Developers with coding expertise can be freed up to focus on designing efficient workflows rather than creating forms or writing auxiliary code.
[0039] Figure 1 is an architectural diagram of a hyper-automation system 100 according to an embodiment of the present invention. As used herein, "hyper-automation" refers to an automation system that combines components of process automation, integration tools, and technologies that enhance work automation capabilities. For example, in some embodiments, robotic process automation (RPA) can be used at the core of the hyper-automation system, and in certain embodiments, the automation capabilities can be extended using AI / machine learning (ML), process mining, analytics, and / or other advanced tools. As the hyper-automation system learns processes, trains AI / ML models, and adopts analytics, for example, more and more knowledge work can be automated, and computing systems within an organization (e.g., computing systems used by individuals and computing systems that operate autonomously) can all participate in the hyper-automation process. The hyper-automation system of some embodiments allows users and organizations to efficiently and effectively discover, understand, and scale automation.
[0040] The hyper-automation system 100 includes user computing systems such as a desktop computer 102, a tablet computer 104, and a smartphone 106. However, any desired user computing system can be used without departing from the scope of the present invention, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc. Additionally, although Figure 1 three user computing systems are shown, any suitable number of user computing systems can be used without departing from the scope of the present invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of user computing systems can be used. The user computing systems can be actively used by users or can operate automatically with little or no user input.
[0041] Each user computing system 102, 104, 106 has (a) corresponding automation process(es) 110, 112, 114 running thereon. In some embodiments, the automation processes are stored remotely (e.g., on a server 130 or in a database 140 and accessed via a network 120) and are loaded by an RPA robot to achieve automation. The automation can exist as a script (e.g., XML, XAML, etc.) or can be compiled into machine-readable code (e.g., as a digital link library).
[0042] Without departing from the scope of the present invention, the automated processes 110, 112, 114 may include, but are not limited to, RPA robots, a part of an operating system, downloadable applications for the corresponding computing system, any other suitable software and / or hardware, or any combination thereof. In some embodiments, one or more of the processes 110, 112, 114 may be a listener. Without departing from the scope of the present invention, the listener may be an RPA robot, a part of an operating system, a downloadable application for the corresponding computing system, or any other software and / or hardware. In fact, in some embodiments, the logical part of the listener(s) is implemented partially or fully via physical hardware.
[0043] The listener monitors and records data related to the user's interaction with the corresponding computing system and / or the operation of an unattended computing system, and sends the data to the core hyper-automation system 120 via a network (e.g., local area network (LAN), mobile communication network, satellite communication network, Internet, any combination thereof, etc.). The data may include, but are not limited to, which buttons are clicked, where the mouse is moved, text entered in a field, one window is minimized while another window is opened, the application associated with the window, etc. In certain embodiments, the data from the listener may be sent periodically as part of a heartbeat message. In some embodiments, the data may be sent to the core hyper-automation system 120 when a predetermined amount of data has been collected, after a predetermined period of time, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0044] The automated process may execute the logic developed in the workflow during design time. In the case of RPA, the workflow may include a set of steps (defined herein as "activities") that are executed in a sequence or some other logical flow. Each activity may include actions, such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, the workflow may be nested or embedded.
[0045] In some embodiments, a long-running workflow for RPA is the main project that supports service orchestration, human intervention, and long-running transactions in an unattended environment. For example, see U.S. Patent No. 10,860,905, which is incorporated herein by reference in its entirety. Human intervention comes into play when certain processes require human input to handle exceptions, approvals, or validations before proceeding to the next step in the activity. In such cases, the process execution is paused, thereby releasing the RPA robot until the human task is completed.
[0046] Long-running workflows can support workflow fragmentation via persistent activities and can be combined with call processes and non-user interaction activities to orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or many computing systems can participate in executing the logic of a long-running workflow. Long-running workflows can run in a session to facilitate rapid execution. In some embodiments, long-running workflows can orchestrate background processes that can include activities that execute API calls and run in a long-running workflow session. In some embodiments, these activities can be invoked by call process activities. A process having user interaction activities running in a user session can be invoked by starting a job from a bootstrap activity (the bootstrap will be described in more detail later in this document). In some embodiments, the user can interact through tasks of a form that need to be completed in the bootstrap. Activities can be included that cause the RPA robot to wait for form tasks to be completed and then resume the long-running workflow.
[0047] (One or more of) the automation processes 110, 112, 114 are in communication with a core hyperautomation system 120. In some embodiments, the core hyperautomation system 120 can run a bootstrap application on one or more servers, such as server 130. Although one server 130 is shown for illustrative purposes, multiple or many servers that are close to each other or in a distributed architecture can be employed without departing from the scope of the present invention. For example, one or more servers can be provided for bootstrap functionality, AI / ML model serving, authentication, governance, and / or any other suitable functionality without departing from the scope of the present invention. In some embodiments, the core hyperautomation system 120 can incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyperautomation system 120 can host multiple software-based servers, such as server 130, on one or more computing systems. In some embodiments, one or more servers of the core hyperautomation system 120, such as server 130, can be implemented via one or more virtual machines (VMs).
[0048] In some embodiments, one or more of the automation processes 110, 112, 114 may invoke one or more AI / ML models 132 that are deployed on or accessible by the core hyper-automation system 120 and are trained to perform various tasks. For example, the AI / ML models 132 may include models trained to find various application versions, perform CV, perform OCR, generate UI descriptors, provide suggestions for the next activity or sequence of activities in an RPA workflow, etc. The AI / ML models may be trained using labeled data that includes, but is not limited to, elements from data sources (e.g., web pages, forms, scanned documents, application interfaces, screens, etc.), previously created RPA workflows, screenshots of various application screens for various versions and their corresponding UI elements, libraries of UI objects, etc. The AI / ML models 132 may be trained to achieve a desired confidence threshold without overfitting to a given training dataset.
[0049] Without departing from the scope of the present invention, the AI / ML models 132 may be trained for any suitable purpose, as will be discussed in more detail below. In some embodiments, two or more of the AI / ML models 132 may be linked (e.g., in series, in parallel, or a combination thereof) such that they jointly provide a collaborative output. The AI / ML models 132 may perform or assist with CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated RPA workflow generation, sequence extraction, cluster detection, audio-to-text conversion, any combination thereof, etc. However, without departing from the scope of the present invention, any desired number and / or type of AI / ML models may be used. For example, using multiple AI / ML models may allow the system to develop a global picture of what is happening on a given computing system. For example, one AI / ML model may perform OCR, another AI / ML model may detect buttons, another AI / ML model may compare sequences, etc. Patterns may be determined by an AI / ML model alone or jointly by multiple AI / ML models. In certain embodiments, one or more AI / ML models are locally deployed on at least one of the computing systems 102, 104, 106.
[0050] In some embodiments, multiple AI / ML models 132 can be used. Each AI / ML model 132 is an algorithm (or model) that runs on data, and for example, the AI / ML model itself can be a deep learning neural network (DLNN) of trained artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 can have multiple layers that perform various functions, such as statistical modeling (e.g., Hidden Markov Model (HMM)), and utilize deep learning techniques (e.g., Long Short-Term Memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform the desired functions.
[0051] In some embodiments, the hyper-automation system 100 can provide four main sets of functions: (1) discovery; (2) build automation; (3) management; and (4) orchestration. In some embodiments, automation (e.g., running on a user computing system, server, etc.) can be run by a software robot such as an RPA robot. For example, attended robots, unattended robots, and / or test robots can be used. Attended robots work with users to assist them in completing tasks (e.g., via UiPath Assistant TM ). Unattended robots work independently of users and can run in the background without the users' knowledge. Test robots are unattended robots that run test cases against applications or RPA workflows. In some embodiments, test robots can run in parallel on multiple computing systems.
[0052] The discovery function can discover and provide automated recommendations for different opportunities for business process automation. This function can be implemented by one or more servers such as server 130. In some embodiments, the discovery function can include providing an automation hub, process mining, task mining, and / or task capture. The automation hub (e.g., UiPath Automation Hub TM ) can provide a mechanism for managing the rollout of automation with visibility and control. For example, ideas for automation can be crowdsourced from employees via a submission form. Feasibility and ROI calculations for automating these ideas can be provided, documentation for future automation can be collected, and collaboration can be provided to build faster from automation discoveries.
[0053] Process mining (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TMrefers to the process of collecting and analyzing data from applications (such as enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization and how to effectively automate those processes, and to indicate what the impact of the automation will be. For example, this data can be collected by a listener from user computing systems 102, 104, 106 and processed by a server (such as server 130). In some embodiments, one or more AI / ML models 132 can be used for this purpose. The information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of process mining goals include, but are not limited to, increasing profit, improving customer satisfaction, complying with regulations and / or contracts, improving employee efficiency, etc.
[0054] Task mining (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TM ) identifies and aggregates workflows (such as employee workflows), and then applies AI to reveal patterns and variations in daily tasks, scores these tasks to facilitate automation and potential savings (such as time and / or cost savings). One or more AI / ML models 132 can be employed to reveal repetitive task patterns in the data. Then, repetitive tasks ripe for automation can be identified. In some embodiments, this information can initially be provided by a listener and analyzed on a server (such as server 130) of the core hyperautomation system 120. Findings from task mining (such as XAML process data) can be exported to a process documentation or designer application (such as UiPath Studio TM ) to create and deploy automation more quickly. Task mining in some embodiments can include taking screenshots of user actions (such as mouse click locations, keyboard inputs, application windows and graphical elements the user is interacting with, timestamps of interactions, etc.), collecting statistics (such as execution times, number of actions, text entries, etc.), editing and annotating screenshots, specifying the types of actions to record, and so on.
[0055] Task capture (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TM)Automatically record the processes involved when the user is working, or provide a framework for unattended processes. Such documentation can include the desired tasks to be automated in the following forms: Process Definition Document (PDD), framework workflows, capture actions for each part of the process, record user actions, and automatically generate a comprehensive workflow diagram, including details about each step, Microsoft documents, XAML files, etc. In some embodiments, the built-ready workflows can be directly exported to designer applications, such as UiPath Studio TM . Task capture can simplify the requirements gathering process for subject matter experts interpreting the process and members of the Center of Excellence (CoE) providing production-level automation.
[0056] Build automation can be achieved via designer applications (e.g., UiPath Studio TM , UiPath StudioX TM or UiPath Studio Web TM ). For example, an RPA developer of the RPA development facility 150 can use the RPA designer application 154 of the computing system 152 to build and test automations for various applications and environments, such as network, mobile, and virtual desktops. API integrations can be provided for various applications, technologies, and platforms. Predefined activities, drag-and-drop modeling, and workflow recorders can make automation easier with minimal coding. Document understanding capabilities can be provided via drag-and-drop AI skills for data extraction and interpretation for calling one or more AI / ML models 132. Such automation can handle almost any document type and format, including tables, checkboxes, signatures, and handwriting. When data is verified or exceptions are handled, this information can be used to retrain the corresponding AI / ML models to improve their accuracy over time.
[0057] The RPA designer application 152 can be designed to call one or more of the trained AI / ML models 132 on the server 130 and / or the generative AI model 172 in the cloud environment via the network 120 (e.g., local area network (LAN), mobile communication network, satellite communication network, Internet, any combination thereof, etc.) to assist in the RPA automation development process. In some embodiments, one or more of the AI / ML models can be packaged with the RPA designer application 152 or otherwise stored locally on the computing system 150.
[0058] In some embodiments, one or more of the RPA Designer application 152 and the AI / ML model 132 may be configured to use an object library stored in the database 140. For example, see U.S. Patent No. 11,748,069, the entire content of which is incorporated herein by reference. The object library may include a UI object library that can be used to develop RPA workflows via the RPA Designer application 152. The object library can be used to add UI descriptors to activities in the workflows of the RPA Designer application 152 for UI automation. In some embodiments, one or more of the AI / ML models 132 may generate new UI descriptors and add them to the object library in the database 140. When the automation is completed in the Designer application 152, they can be published on the server 130, rolled out to the computing systems 102, 104, 106, etc.
[0059] For example, an integration service may allow developers to seamlessly combine UI automation with API automation. Automations that require APIs or traverse both API and non-API applications and systems can be built. A repository (e.g., UiPath Object Repository TM ) or a marketplace (e.g., UiPathMarketplace TM ) may be provided to allow developers to automate various processes more quickly. Thus, when building automations, the Hyper-Automation System 100 can provide a user interface, a development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. In some embodiments, the Hyper-Automation System 100 implements the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the Hyper-Automation System 100.
[0060] In some embodiments, components of the hyper-automation system 100, such as the (multiple) designer applications and / or external rule engines, provide support for managing and implementing governance policies that control the various functions provided by the hyper-automation system 100. Governance is about an organization formulating policies to prevent users from developing the ability to automate (e.g., RPA robots) that can take actions that may harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the US Health Insurance Portability and Accountability Act (HIPAA), the terms of service of third-party applications, etc. Since developers may otherwise create automations that violate privacy laws, terms of service, etc. when executing their automations, some embodiments implement access control and governance restrictions at the robot and / or robot design application levels. In some embodiments, this can be done by preventing developers from relying on unapproved software libraries to provide an additional level of security and compliance to the automation process development pipeline, which may introduce security risks or operate in a way that violates policies, regulations, privacy laws, and / or privacy policies. For example, see U.S. Patent No. 11,733,668, the entire content of which is incorporated herein by reference.
[0061] The management function can provide management, deployment, and optimization of automation across the organization. In some embodiments, the management function can include orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point for third-party solutions and applications for automation applications and / or RPA robots. The management capabilities of the hyper-automation system 100 can include, but are not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots, etc.
[0062] A bootstrap application, such as UiPath Orchestrator TM (in some embodiments, it can be provided as part of UiPathAutomation Cloud TM or provided locally, in a virtual machine, in a private or public cloud, in a Linux TM VM, or provided as a cloud-native single-container suite via UiPath Automation Suite TM ) provides orchestration capabilities for deploying, monitoring, optimizing, scaling, and securing the deployment of RPA robots. A test suite (e.g., UiPathTest Suite TM ) can provide test management to monitor the quality of the deployed automation. The test suite can facilitate test planning and execution, requirements fulfillment, and defect traceability. The test suite can include comprehensive test reports.
[0063] Analytics software (e.g., UiPath InsightsTM ) It can track, measure, and manage the performance of deployed automation. Analytics software can align automated operations with specific key performance indicators (KPIs) and strategic outcomes for an organization. The analytics software can present results in a dashboard format for better understanding by human users.
[0064] For example, data services (e.g., UiPath Data Service TM ) can be stored in database 140 and bring data to a single scalable and secure place using a drag-and-drop storage interface. Some embodiments can provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-level security, and scalability. AI capabilities can be provided by an AI hub (e.g., UiPath AI Center TM ) which facilitates incorporating AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can make these capabilities accessible even to non-data scientists. Deployed automation (e.g., RPA robots) can call AI / ML models, such as AI / ML model 132, from the AI hub. The performance of AI / ML models can be monitored, trained, and improved using human-verified data such as that provided by data review center 160. Human reviewers can provide labeled data to core hyperautomation system 120 via review application 152 on computing system 154. For example, human reviewers can verify that predictions made by AI / ML model 132 and / or generative AI model 172 are accurate or otherwise provide corrections. For example, then this dynamic input can be saved as training data for retraining AI / ML model 132 and / or generative AI model 172 and can be stored in a database such as database 140. The AI hub can then schedule and execute training jobs to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining AI / ML model 132 and / or generative AI model 172.
[0065] The participation function teams humans and automation to seamlessly collaborate on desired processes. Low-code applications (e.g., via UiPath Apps TM ) can be built to connect browser tabs and legacy software even in some embodiments where APIs are lacking. For example, applications can be quickly created using a web browser with a rich library of drag-and-drop controls. Applications can connect to a single automation or multiple automations.
[0066] The Action Center (e.g., UiPath Action Center TM) provides a simple and efficient mechanism to hand off processes from automated to human and vice versa. A human can provide approvals or escalations, make exceptions, etc. Then automation can perform the automated functions of a given workflow.
[0067] A local assistant can be provided as a launchpad for users to initiate automation (e.g., UiPath Assistant TM ). This functionality can be provided in a tray, for example, provided by the operating system, and can allow users to interact with RPA robots and RPA robot-driven applications on their computing systems. The interface can list the automations approved for a given user and allow the user to run them. These can include off-the-shelf automations from an automation marketplace, an internal automation store in an automation hub, etc. When the automations run, they can run in parallel as local instances with other processes on the computing system, so the user can use the computing system while the automation performs its actions. In some embodiments, the assistant is integrated with a task capture feature so that users can record the processes they are about to automate from the assistant launchpad.
[0068] Chatbots (e.g., UiPath Chatbots TM ), social messaging applications, and / or voice commands can enable users to run automations. This can simplify access to the information, tools, and resources that users need to interact with customers or perform other activities. As with other processes, conversations between people can be easily automated. Trigger RPA robots launched in this way can perform operations such as checking order status, posting data in a CRM, etc., possibly using natural language commands.
[0069] In some embodiments, end-to-end measurement and governance of automation programs of any scale can be provided by the hyperautomation system 100. According to the above, analytics can be employed to understand the performance of automation (e.g., via UiPath Insights TM ). Data modeling and analysis using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards allow visualizing data across desired metrics, discovering new analytical insights, tracking performance metrics, finding ROI for automation, performing telemetry monitoring on user computing systems, detecting errors and anomalies, and debugging automations. An automation management console (e.g., UiPath Automation Ops TM ) can be provided to manage automations throughout their lifecycle. Organizations can govern how automations are built, what users can do with them, and which automations users can access.
[0070] In some embodiments, the hyper-automation system 100 provides an iterative platform. Processes can be discovered, automation can be built, tested, and deployed, performance can be measured, usage of the automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the process can repeat itself. This facilitates a more robust and efficient automation suite.
[0071] In some embodiments, generative AI models are used. Generative AI can generate various types of content, such as text, images, audio, and synthetic data. Various types of generative AI models can be used, including but not limited to large language models (LLMs), generative adversarial networks (GANs), variational autoencoders (VAEs), transformers, etc. These models can be part of the AI / ML models 132 hosted on the server 130. For example, generative AI models can be trained on large text information corpora to perform semantic understanding, to understand the essence of the content presented on a screen from text, to automatically generate code, and so on. In certain embodiments, generative AI models 172 provided by existing cloud ML service providers (such as etc.) can be adopted and trained to provide such functionality. In generative AI embodiments where the generative AI models 172 are remotely hosted, the server 130 can be configured to integrate with a third-party API that allows the server 130 to send requests including the necessary input information to the generative AI models 172 and receive responses in return (e.g., semantic matching of fields between application versions, classification of the type of application on a screen, etc.). Such embodiments can provide a more advanced and complex user experience, as well as provide access to the existing NLP and other ML capabilities of these companies.
[0072] In some embodiments, one aspect of generative AI models is the use of transfer learning. In transfer learning, a pre-trained generative AI model (such as an LLM) is fine-tuned on a specific task or domain. This allows the LLM to fully utilize the knowledge it has learned during its initial training and adapt it to a specific application. In the case of an LLM, the pre-training phase involves training the LLM on a large text corpus typically consisting of billions of words. During this phase, the LLM learns the relationships between words and phrases, which enables the LLM to generate coherent human-like responses to text-based inputs. The output of this pre-training phase is an LLM with a very high level of understanding of the basic patterns in natural language.
[0073] In the fine-tuning phase, a pre-trained LLM is adapted to a specific task or domain by training the LLM on a smaller dataset specific to the task. For example, in some embodiments, the LLM may be trained to analyze a certain type or multiple types of data sources to improve its accuracy about their contents. This information may be provided as part of the training data, and the LLM may learn to focus on these areas and more accurately identify the data elements therein. Fine-tuning allows the LLM to learn the nuances of a task or domain, such as the specific vocabulary and grammar used in that domain, without requiring as much data as would be required if the LLM were trained from scratch. By leveraging the knowledge learned in the pre-training phase, a fine-tuned LLM can achieve the performance of the state of the art on a specific task using a relatively small amount of training data.
[0074] LLM can be trained using a vector database. A vector database indexes, stores, and provides access to structured or unstructured data (e.g., text, images, time series data, etc.) and its vector embeddings. Data such as text can be tokenized, where individual letters, words, or sequences of words are parsed from the text into tokens. These tokens are then "embedded" into vector embeddings, which are digital representations of this data. Vector databases allow users to find and retrieve similar objects quickly and at scale in a production environment.
[0075] AI and ML allow unstructured data to be represented digitally without losing its semantics in vector embeddings. A vector embedding is a long string of numbers, each of which describes the characteristics of the data object represented by the vector embedding. Similar objects are grouped together in the vector space. In other words, the more similar the objects are, the closer the vector embeddings representing the objects will be to each other. Similar objects can be found using vector search, similarity search, or semantic search. The distance between vector embeddings can be calculated using various techniques, including but not limited to squared Euclidean or L2 squared distance, Manhattan or L1 distance, cosine similarity, dot product, Hamming distance, etc. It can be beneficial to choose the same metric used to train the AI / ML model.
[0076] Vector indexes can be used to organize vector embeddings, enabling efficient data retrieval. If there are a large number of data points, calculating the distance between a vector embedding and all other vector embeddings in a vector database using the k-nearest neighbor (kNN) algorithm can be computationally expensive because the required computations increase linearly (O(n)) with the dimensionality and number of data points. Using the approximate nearest neighbor (ANN) method to find similar objects is more efficient. The distances between vector embeddings are pre-computed, and similar vectors are organized and stored close to each other (e.g., in clusters or graphs), allowing similar objects to be found more quickly. This process is called "vector indexing". ANN algorithms that can be used in some embodiments include, but are not limited to, clustering-based indexes, proximity graph-based indexes, tree-based indexes, hash-based indexes, compression-based indexes, etc.
[0077] Figure 2 is an architectural diagram showing the RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 is Figure 1 part of the hyper-automation system 100. The RPA system 200 includes a designer 210 that allows developers to design and implement workflows. The designer 210 can provide solutions for application integration and automating third-party applications, managing information technology (IT) tasks, and business IT processes. The designer 210 can facilitate the development of an automation project, which is a graphical representation of a business process. Simply put, the designer 210 facilitates the development and deployment of workflows and robots. In some embodiments, the designer 210 can be an application running on a user's desktop, an application running remotely in a VM, a web application, etc.
[0078] Automation projects automate rule-based processes by giving developers control over the execution order and the relationships between the set of custom steps (defined herein as the above "activities") developed in a workflow. A commercial example of an embodiment of the designer 210 is UiPath Studio TM . Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows can be nested or embedded.
[0079] Some types of workflows can include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. Sequences can be particularly suitable for linear processes, implementing a flow from one activity to another without cluttering the workflow. Flowcharts can be particularly suitable for more complex business logic, implementing decision integration and activity connection in more diverse ways through multiple branching logical operators. FSMs can be particularly suitable for large workflows. FSMs can use a finite number of states in their execution, which are triggered by conditions (i.e., transitions) or activities. Global exception handlers can be particularly suitable for determining workflow behavior when execution errors are encountered and for the debugging process.
[0080] When a workflow is developed in the designer 210, the execution of the business process is orchestrated by the orchestrator 220, which orchestrates the execution of one or more robots 230 of the workflow developed in the designer 210. A commercial example of an embodiment of the orchestrator 220 is UiPath Orchestrator TM . The orchestrator 220 facilitates the management of the creation, monitoring, and deployment of resources in the environment. The orchestrator 220 can act as an integration point with third - party solutions and applications. As described above, in some embodiments, the orchestrator 220 can be Figure 1 part of the core hyper - automation system 120.
[0081] The orchestrator 220 can manage a fleet of robots 230, connecting and executing the robots 230 from a central point. The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but for development and testing purposes), and non - production robots (similar to attended robots 232 but for development and testing purposes). Attended robots 232 are triggered by user events and operate with a human on the same computing system. Attended robots 232 can be used by the orchestrator 220 for centralized process deployment and logging medium. Attended robots 232 can assist human users in completing various tasks and can be triggered by user events. In some embodiments, a process cannot start from the orchestrator 220 on this type of robot, and / or they cannot run under a locked screen. In certain embodiments, attended robots 232 can only start from the robot tray or from the command prompt. In some embodiments, attended robots 232 should run under human supervision.
[0082] The unattended robot 234 operates unattended in a virtual environment and can automate many processes. The unattended robot 234 can be responsible for remote execution, monitoring, scheduling, and providing support for work queues. In some embodiments, debugging for all robot types can be run in the designer 210. Both attended and unattended robots can automate various systems and applications, including but not limited to mainframes, web applications, VMs, enterprise applications (e.g., applications produced by etc.) and computing system applications (e.g., desktop and laptop computer applications, mobile device applications, wearable computer applications, etc.).
[0083] The orchestrator 220 can have various capabilities, including but not limited to provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning can include creating and maintaining a connection (e.g., a web application) between the robot 230 and the orchestrator 220. Deployment can include ensuring the correct delivery of the package version to the assigned robot 230 for execution. Configuration can include the maintenance and delivery of the robot environment and process configuration. Queuing can include providing management of queues and queue items. Monitoring can include tracking robot identification data and maintaining user permissions. Logging can include storing logs into a database (e.g., a Structured Query Language (SQL) database or a "Not Only SQL" (NoSQL) database) and / or other storage mechanisms (e.g., those that provide the ability to store and quickly query large datasets of ) and indexing. The orchestrator 220 can provide interconnectivity by acting as a centralized communication point for third-party solutions and / or applications.
[0084] The robot 230 is an execution agent that implements the workflows built in the designer 210. A commercial example of some embodiments of the (multiple) robots 230 is UiPath Robots TM . In some embodiments, the robot 230 defaults to installing a service managed by the Microsoft Service Control Manager (SCM). Thus, such a robot 230 can open an interactive session under the local system account and has the permissions of the service.
[0085] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, this means they have the same rights as the user under whom the given robot 230 is installed. This feature can also be used for high-density (HD) robots, which ensures the full utilization of each machine to its maximum potential. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0086] In some embodiments, the robot 230 is split into several components, each component dedicated to a specific automation task. In some embodiments, the robot components include, but are not limited to, the SCM-managed robot service, the user-mode robot service, the actuator, the agent, and the command line. The SCM-managed robot service manages and monitors sessions and acts as an agent between the bootstrapper 220 and the execution host (i.e., the computing system on which the robot 230 executes). These services are trusted and manage certificates for the robot 230. The console application is launched by the SCM under the local system.
[0087] In some embodiments, the user-mode robot service manages and monitors sessions and acts as an agent between the bootstrapper 220 and the execution host. The user-mode robot service can be trusted and manage certificates for the robot 230. If the SCM-managed robot service is not installed, the application can be automatically launched.
[0088] The actuator can run a given job (i.e., they can execute a workflow) under a session. The actuator can know the dots per inch (DPI) settings of each monitor. The agent can be a Windows Presentation Foundation (WPF) application that displays available jobs in the system tray window. The agent can be a client of the service. The agent can request to start or stop a job and change settings. The command line is a client of the service. The command line is a console application that can request to start a job and wait for its output. As described above, splitting the components of the robot 230 helps developers, support users, and computing systems more easily run, identify, and track what each component is doing. Special behaviors can be configured for each component in this way, such as setting different firewall rules for the actuator and the service. In some embodiments, the actuator can always know the DPI settings of each monitor. Thus, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. In some embodiments, the projects from the designer 210 can also be independent of the browser zoom level. In some embodiments, for applications where the DPI is unknown or deliberately marked as unknown, the DPI can be disabled.
[0089]
[0090] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA bots that utilize AI / ML models (e.g., as part of its AI hub) deployed in the core hyper-automation system 240. Such RPA bots can send inputs for executing the AI / ML models and receive outputs therefrom via the core hyper-automation system 240.
[0091] As described above, one or more of the bots 230 can be listeners. These listeners can provide information to the core hyper-automation system 240 about what the user is doing while using their computing system. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0092] An assistant / chatbot 250 can be provided on the user computing system to allow the user to initiate an RPA local bot. For example, the assistant can be located in the system tray. The chatbot can have a user interface such that the user can see text in the chatbot. Alternatively, the chatbot can lack a user interface and run in the background, listening to the user's voice using the computing system's microphone.
[0093] In some embodiments, data tagging can be performed by the user of the computing system on which the bot is executing or on another computing system to which the bot provides information. For example, if a bot invokes an AI / ML model that performs CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user can draw a rectangle around the mis-identified or un-identified component and potentially provide text with the correct identification. This information can be provided to the core hyper-automation system 240 and then later used to train a new version of the AI / ML model.
[0094] Figure 3 is an architectural diagram showing the deployment of an RPA system 300 according to an embodiment of the present invention. In some embodiments, the RPA system 300 can be Figure 2 part of the RPA system 200 and / or Figure 1 part of the hyper-automation system 100. The deployed RPA system 300 can be a cloud-based system, an on-premises system, a desktop-based system, etc., that provides enterprise-level, user-level, or device-level automation solutions for automating different computing processes.
[0095] It should be noted that, without departing from the scope of the present invention, the client side, the server side, or both may include any desired number of computing systems. On the client side, the robotic application 310 includes an actuator 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not run on the same computing system as the actuator 312 and the agent 314. The actuator 312 is in the running process. A number of business items may run simultaneously, as shown in Figure 3 . The agent 314 (e.g., service) is a single point of contact for all the actuators 312 in this embodiment. All messages in this embodiment are logged into the bootstrapper 340, which further processes these messages via the database server 350, the AI / ML server 360, the indexer server 370, or any combination thereof. As discussed above with respect to Figure 2 , the actuator 312 may be a robotic component.
[0096] In some embodiments, a robot represents an association between a machine name and a user name. A robot can manage multiple actuators simultaneously. On a computing system that supports multiple interactive sessions running simultaneously (e.g., server 2012), multiple robots can run simultaneously, each using a unique user name in a separate session. This is the HD robot mentioned above.
[0097] The agent 314 is also responsible for sending the status of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still running), and downloading the desired version of the packet to be executed. In some embodiments, the communication between the agent 314 and the bootstrapper 340 is always initiated by the agent 314. In a notification scenario, the agent 314 can open a WebSocket channel, which is later used by the bootstrapper 340 to send commands to the robot (e.g., start, stop, etc.).
[0098] The listener 330 monitors and records data related to the interaction of the user with the attended computing system and / or the operation of the unattended computing system where the listener 330 is located. Without departing from the scope of the present invention, the listener 330 can be an RPA robot, a part of the operating system, a downloadable application for the corresponding computing system, or any other software and / or hardware. In fact, in some embodiments, the logical part of the listener is implemented partially or fully via physical hardware.
[0099] On the server side, it includes a presentation layer (web application 342, Open Data Protocol (OData) Representational State Transfer (REST) Application Programming Interface (API) endpoint 344, and notification and monitoring 346), a service layer (API implementation / business logic 348), and a persistence layer (database server 350, AI / ML server 360, and indexer server 370). The bootstrapper 340 includes the web application 342, the OData REST API endpoint 344, the notification and monitoring 346, and the API implementation / business logic 348. In some embodiments, most of the actions performed by the user in the interface of the bootstrapper 340 (e.g., via the browser 320) are performed by calling various APIs. Without departing from the scope of the present invention, such actions may include, but are not limited to, starting a job on a robot, adding / removing data in a queue, scheduling a job to run unattended, etc. The web application 342 is the visual layer of the server platform. In this embodiment, the web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, without departing from the scope of the present invention, any desired markup language, scripting language, or any other format may be used. In this embodiment, the user interacts with the web page from the web application 342 via the browser 320 to perform various actions to control the bootstrapper 340. For example, the user can create a group of robots, assign groups to robots, analyze logs for each robot and / or each process, start and stop robots, and so on.
[0100] In addition to the web application 342, the bootstrapper 340 also includes a service layer that exposes the OData REST API endpoint 344. However, without departing from the scope of the present invention, other endpoints may be included. The REST API is consumed by both the web application 342 and the proxy 314. In this embodiment, the proxy 314 is a supervisor for one or more robots on the client computer.
[0101] The REST API in this embodiment covers configuration, logging, monitoring, and queuing functions. In some embodiments, the configuration endpoint can be used to define and configure application users, permissions, robots, assets, releases, and environments. For example, the logging REST endpoint can be used to log different information, such as errors, explicit messages sent by robots, and other environment-specific information. If the start job command is used in the bootstrapper 340, the deployment REST endpoint can be used by the robot to query the version of the group that should be executed. The queuing REST endpoint can be responsible for queue and queue item management, such as adding data to the queue, retrieving transactions from the queue, setting the status of transactions, etc.
[0102] The monitoring REST endpoints can monitor the web application 342 and the agent 314. The notification and monitoring API 346 can be a REST endpoint for registering the agent 314, delivering configuration settings to the agent 314, and for sending / receiving notifications from the server and the agent 314. In some embodiments, the notification and monitoring API 346 can also use WebSocket communication.
[0103] In some embodiments, the APIs in the service layer can be accessed through the configuration of appropriate API access paths, e.g., based on whether the bootstrapper 340 and the overall hyper-automation system have a local deployment type or a cloud-based deployment type. The API for the bootstrapper 340 can provide customized methods for querying statistics about various entities registered in the bootstrapper 340. In some embodiments, each logical resource can be an OData entity. In such an entity, components such as robots, processes, queues, etc. can have properties, relationships, and operations. In some embodiments, the web application 342 and / or the agent 314 can consume the API of the bootstrapper 340 in two ways: by obtaining API access information from the bootstrapper 340, or by registering an external application to use the Oauth flow.
[0104] In the present embodiment, the persistence layer includes three servers in the server - a database server 350 (e.g., an SQL server), an AI / ML server 360 (e.g., a server providing AI / ML model services, such as an AI hub function), and an indexer server 370. The database server 350 in this embodiment stores the configurations of robots, robot groups, associated processes, users, roles, schedules, etc. In some embodiments, this information is managed by the web application 342. The database server 350 can manage queues and queue items. In some embodiments, the database server 350 can store the messages logged by robots (in addition to or instead of the indexer server 370). For example, the database server 350 can also store process mining, task mining, and / or task capture related data received from the listener 330 installed on the client side. Although no arrow is shown between the listener 330 and the database 350, it should be understood that in some embodiments, the listener 330 is capable of communicating with the database 350 and vice versa. This data can be stored in the form of PDD, images, XAML files, etc. The listener 330 can be configured to intercept user actions, processes, tasks, and performance metrics on the corresponding computing system where the listener 330 is located. For example, the listener 330 can record user actions (e.g., clicks, typed characters, location, application, active elements, time, etc.) on its corresponding computing system and then convert these actions into a suitable format to be provided to and stored in the database server 350.
[0105] The AI / ML Server 360 facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can make these capabilities accessible to those who are not even data scientists. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML Server 360. The performance of the AI / ML models can be monitored and they can be trained and improved using human-verified data. The AI / ML Server 360 can schedule and execute training jobs to train new versions of the AI / ML models.
[0106] The AI / ML Server 360 can store data related to AI / ML models and ML groupings used to configure various ML skills for users at development time. As used herein, an ML skill is a pre-built and trained ML model for a process, e.g., which can be used by automation. The AI / ML Server 360 can also store data related to document understanding techniques and frameworks, algorithms, and software groupings for various AI / ML capabilities, which include but are not limited to intent analysis, NLP, speech analysis, different types of AI / ML models, etc.
[0107] The Indexer Server 370 (which is optional in some embodiments) stores and indexes information logged by robots. In certain embodiments, the Indexer Server 370 can be disabled via configuration settings. In some embodiments, the Indexer Server 370 uses which is an open-source project full-text search engine. Messages logged by robots (e.g., using activities such as log messages or write lines) can be sent to the Indexer Server 370 via the logging REST endpoint(s), where they are indexed for future use.
[0108] Figure 4It is an architecture diagram showing the relationship 400 among the designer 410, activities 420, 430, 440, 450, driver 460, API 470, and AI / ML model 480 according to an embodiment of the present invention. According to the above, developers use the designer 410 to develop workflows executed by robots. In some embodiments, various types of activities can be presented to the developers. The designer 410 can be local or remote to the user's computing system (e.g., accessed via a VM or a local web browser interacting with a remote web server). The workflow can include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. The user-defined activities 420 and API-driven activities 440 interact with applications via their APIs. In some embodiments, the user-defined activities 420 and / or AI / ML activities 440 can call one or more AI / ML models 480, which can be located locally and / or remotely on the computing system where the robot operates.
[0109] Some embodiments are capable of identifying non-text visual components in an image, which are referred to herein as CV. However, it should be noted that in some embodiments, CV incorporates OCR. CV can be performed at least in part by the (one or more) AI / ML models 480. Some CV activities associated with such components can include, but are not limited to, extracting text from segmented tagged data using OCR, fuzzy text matching, cropping segmented tagged data using ML, comparing the text extracted from the tagged data with ground truth data, etc. In some embodiments, there can be hundreds or even thousands of activities that can be implemented in the user-defined activities 420. However, any number and / or type of activities can be used without departing from the scope of the present invention.
[0110] The UI automation activities 450 are a subset of special lower-level activities written in lower-level code and facilitating screen interaction. The UI automation activities 450 facilitate these interactions via the driver 460, which allows the robot to interact with the desired software. For example, the driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, etc. In some embodiments, one or more of the AI / ML models 480 can be used by the UI automation activities 450 to perform interactions with the computing system. In certain embodiments, the AI / ML models 480 can enhance the driver 460 or replace them entirely. In fact, in certain embodiments, the driver 460 is not included.
[0111] The driver 460 can interact with the OS at a lower level via the OS driver 462 to find hooks, monitor keys, etc. The driver 460 can facilitate interactions with Integrations such as. For example, the "click" activity performs the same role in these different applications via the driver 460.
[0112] Figure 5 is an architectural diagram of a computing system 500 configured to perform automatic code generation for RPA according to an embodiment of the present invention. In some embodiments, the computing system 500 may be one or more of the computing systems depicted and / or described herein. In certain embodiments, the computing system 500 may be part of a hyper-automation system such as shown in Figure 1 and Figure 2 . The computing system 500 includes a bus 505 or other communication mechanism for conveying information, and (a) processor(s) 510 coupled to the bus 505 for processing the information. The (a) processor(s) 510 may be any type of general-purpose or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The (a) processor(s) 510 may also have multiple processing cores, and at least some of these cores may be configured to perform specific functions. Multiprocessing in parallel may be used in some embodiments. In certain embodiments, at least one of the (a) processor(s) 510 may be a neuromorphic circuit including processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require the typical components of a von Neumann computing architecture.
[0113] The computing system 500 further includes a memory 515 for storing information and instructions to be executed by the (a) processor(s) 510. The memory 515 may include any combination of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage such as a magnetic or optical disk, or any other type of non-transitory computer-readable medium or combination thereof. The non-transitory computer-readable medium may be any available medium that can be accessed by the (a) processor(s) 510, and may include volatile media, non-volatile media, or both. The medium may also be removable, non-removable, or both. The computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, without departing from the scope of the present invention, the communication device 520 may include one or more antennas, which are singular, arrayed, phased, switched, beamforming, beam steering, combinations thereof, and / or any other antenna configuration.
[0114] The (multiple) processors 510 are also coupled to a display 525 via a bus 505. Without departing from the scope of the present invention, any suitable display device and haptic I / O can be used. A keyboard 530 and a cursor control device 535 (such as a computer mouse, touchpad, etc.) are further coupled to the bus 505 to enable a user to interface with the computing system 500. However, in some embodiments, there may be no physical keyboard and mouse, and the user may interact with the device only through the display 525 and / or a touchpad (not shown). Any type and combination of input devices can be used as a matter of design choice. In some embodiments, there is no physical input device and / or display. For example, a user may interact remotely with the computing system 500 via another computing system communicating with it, or the computing system 500 may operate automatically.
[0115] The memory 515 stores software modules that provide functionality when executed by the (multiple) processors 510. The module includes an operating system 540 for the computing system 500. The module also includes an automatic code generation module 545 that is configured to execute all or part of the AI / ML processes or derivatives thereof described herein. The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0116] One of ordinary skill in the art will understand that, without departing from the scope of the present invention, a "system" may be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices. Presenting the above functions as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. In fact, the methods, systems, and devices disclosed herein may be implemented in a localized and distributed form consistent with computing technology, including cloud computing systems. The computing system may be a part of or otherwise accessible from: a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public or private cloud, a hybrid cloud, a server farm, any combination thereof, etc. Without departing from the scope of the present invention, any localized or distributed architecture may be used.
[0117] It should be noted that some of the system features described in this specification have been presented as modules to more particularly emphasize their implementation independence. For example, a module may be implemented as a hardware circuit that includes a custom very large scale integration (VLSI) circuit or gate array, off-the-shelf semiconductors, such as logic chips, transistors, or other discrete components. A module may also be implemented in a programmable hardware device, such as a field programmable gate array, a programmable array logic, a programmable logic device, a graphics processing unit, etc.
[0118] The module can also be implemented at least partially in software for execution by various types of processors. The identified executable code units can include, for example, one or more physical or logical blocks of computer instructions, which can be organized, for example, as objects, procedures, or functions. However, the executable programs of the identified modules do not need to be physically located together, but can include different instructions stored in different locations, which, when logically linked together, include the module and achieve the above-mentioned purpose for the module. In addition, without departing from the scope of the present invention, the module can be stored on a computer-readable medium, which can be, for example, a hard disk drive, a flash device, a RAM, a magnetic tape, and / or any other such non-transitory computer-readable medium for storing data.
[0119] In fact, the modules of executable code can be a single instruction or many instructions, and can even be distributed across several different code segments, different programs, and across several memory devices. Similarly, the operational data can be identified and shown within the module, and can be embodied in any suitable form and organized within any suitable type of data structure. The operational data can be collected as a single data set, or can be distributed across different locations, including on different storage devices, and can exist at least partially only as electronic signals on a system or network.
[0120] Without departing from the scope of the present invention, various types of AI / ML models can be trained and deployed. For example, Figure 6A An example of a neural network 600 trained to supplement automated code generation for RPA according to an embodiment of the present invention is shown. The neural network 600 includes a number of hidden layers. Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) generally have multiple layers, although in some cases, an SLNN can have only one or two layers and generally fewer than a DLNN. Generally, a neural network architecture includes an input layer, multiple intermediate layers, and an output layer, as is the case with the neural network 600.
[0121] DLNNs often have many layers (e.g., 10, 50, 200, etc.), and subsequent layers generally reuse features from previous layers to compute more complex general functions. On the other hand, SLNNs tend to have only a few layers and are relatively fast to train because expert features are created in advance from the raw data samples. However, feature extraction is laborious. On the other hand, DLNNs generally do not require expert features but tend to take longer to train and have more layers.
[0122] For these two methods, this layer is trained simultaneously on the training set, and overfitting is typically checked on an isolated cross-validation set. Both of these techniques can produce excellent results, and there is a fair amount of enthusiasm for both methods. The optimal size, shape, and number of individual layers vary depending on the problem solved by the corresponding neural network.
[0123] Return Figure 6A , providing an RPA workflow, source files (e.g., PDF documents, PDDs, policies, XAML files, etc.), UI drawings and screenshots, task mining information, etc. as the input layer and providing it as input to the J neurons of hidden layer 1. Various other inputs are possible, including but not limited to computing system state information, published automations, business rules, information about what the RPA workflow and / or tasks are related to, initial definitions of automations, process automation documents, etc. Although all of these inputs are fed to each neuron in this example, various architectures can be used alone or in combination without departing from the scope of the present invention, including but not limited to feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks.
[0124] Hidden layer 2 receives input from hidden layer 1, hidden layer 3 receives input from hidden layer 2, and so on for all hidden layers until the last hidden layer provides its output as input to the output layer. Even though multiple suggestions are shown as outputs here, in some embodiments, only a single output suggestion is provided. In certain embodiments, the suggestions are ranked based on confidence scores.
[0125] It should be noted that the numbers of neurons I, J, K, and L are not necessarily equal. Thus, any desired number of layers can be used for a given layer of neural network 600 without departing from the scope of the present invention. In fact, in some embodiments, the types of neurons in a given layer may not all be the same.
[0126] The neural network 600 is trained to assign (a) confidence score(s) to the appropriate output. To reduce inaccurate predictions, in some embodiments, only those results with confidence scores that meet or exceed a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs with confidence scores above that amount may be used, and the rest may be ignored.
[0127] A neural network is a probabilistic structure that typically has (a) confidence score(s). This can be a score learned by an AI / ML model based on the frequency of correctly identifying similar inputs during training. Some common types of confidence scores include decimal numbers between 0 and 1 (which can also be interpreted as a confidence percentage), numbers between negative infinity and positive infinity, a set of expressions (e.g., "low", "medium", and "high"), etc. Various post - processing calibration techniques can also be employed to obtain more accurate confidence scores, such as temperature scaling, batch normalization, weight decay, negative log - likelihood (NLL), etc.
[0128] A "neuron" in a neural network is algorithmically implemented as a mathematical function typically based on the functionality of a biological neuron. A neuron receives weighted inputs and has a summation and an activation function that governs whether it passes an output to the next layer. The activation function can be a non - linear threshold activity function where nothing happens if the value is below the threshold, but then the function responds linearly above the threshold (i.e., the rectified linear unit (ReLU) non - linearity). The summation function and the ReLU function are used in deep learning because real neurons can have approximately similar activity functions. Via linear transformations, information can be subtracted, added, etc. Essentially, a neuron acts as a gating function that passes an output to the next layer, which is governed by its underlying mathematical function. In some embodiments, different functions can be used for at least some neurons.
[0129] In Figure 6B an example of a neuron 610 is shown. Inputs x1, x2, ……, x n from the previous layer are assigned corresponding weights w1, w2, ……, w n . Thus, the collective input from the previous neuron 1 is w1x1. These weighted inputs are used in the summation function of the neuron modified by a bias, such as:
[0130]
[0131] This sum is compared with an activation function f(x) to determine whether the neuron "fires". For example, f(x) can be given by:
[0132]
[0133] Therefore, the output y of neuron 610 can be given by the following equation:
[0134]
[0135] In this case, neuron 610 is a single-layer perceptron. However, without departing from the scope of the present invention, any suitable neuron type or combination of neuron types can be used. It should also be noted that in some embodiments, without departing from the scope of the present invention, the range of values of the weights of the activation function and / or the output value(s) can be different.
[0136] A target or "reward function" is often employed. The reward function explores intermediate transitions and steps with short-term and long-term rewards to guide the search of the state space and attempt to achieve the goal (e.g., finding the most accurate answer to a user query based on an associated metric). During training, various labeled data are fed through the neural network 600. Successful identifications enhance the weights for the inputs to the neurons, while unsuccessful identifications weaken the weights. A cost function (such as mean squared error (MSE) or gradient descent) can be used to penalize slightly incorrect predictions, much less than very incorrect predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, the data scientist can modify the reward function, provide corrections for incorrect predictions, and so on.
[0137] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "uncover" the hidden layers of a neural network to see how much loss each node is to bear, and then update the weights in such a way as to minimize the loss by giving lower weights to nodes with a higher error rate and vice versa. In other words, backpropagation allows the data scientist to repeatedly adjust the weights in order to minimize the difference between the actual output and the desired output.
[0138] The backpropagation algorithm is built on the mathematical basis in optimization theory. In supervised learning, training data with known outputs are passed through the neural network, and the error is calculated using a cost function from the known target outputs, which gives the error for backpropagation. The error is calculated at the output and transformed into a correction for the network weights, which will minimize the error.
[0139] In the case of supervised learning, an example of backpropagation is provided below. A column vector input x is processed through a series of N non-linear activation functions f between each layer i = 1,..., N of the network i where the output at a given layer is first multiplied by the synaptic matrix W i , and the bias vector b is added i . The network output o is given by the following equation
[0140] o = f N (W N f N-1 (W N-1 f N-2 (...f1(W1x + b1)...)+b N-1 )+b N ) (4)
[0141] In some embodiments, comparing o to the target output t results in an error that is minimized
[0142] Optimization in the form of a gradient descent process can be used to minimize the error by modifying the synaptic weights W for each layer i The gradient descent process requires computing the output o given the input x corresponding to the known target output t, and producing the error o - t. The global error is then propagated backward to give a local error for weight updates, which is computed similarly to but not exactly the same as the computation used for forward propagation. In particular, the backpropagation step typically requires an activation function of the form p j (n j ) = f j ′(n j ) where n j is the network activity at layer j (i.e., n j = W j o j-1 + b j ) where o j = f j (n j ), and the prime ′ denotes the derivative of the activation function f.
[0143] The weight updates can be computed via the formula:
[0144]
[0145] where denotes the Hadamard product (i.e., the element-wise product of two vectors), T denotes matrix transpose, and o j denotes f j (W j o j-1 + b j ), where o0 = x. Here, the learning rate η is chosen with respect to machine learning considerations. Below, η is related to the neural Hebbian learning mechanism used in the neural implementation. Note that the synapses W and b can be combined into a large synaptic matrix where it is assumed that the input vector has 1 appended, and an additional column representing b synapses is included in W.
[0146] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., using the F2 or F4 threshold for detection and approximating 2000 epochs to reach 97% or better). Without departing from the scope of the present invention, in some embodiments, the F1 score, F2 score, F4 score, or any other suitable technique may be used to determine this level of accuracy. When training on training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not encountered before. This helps ensure that the AI / ML model does not "overfit", such that it performs well on the training data but poorly on other data.
[0147] In some embodiments, it may not be known what level of accuracy the AI / ML model can achieve. Thus, if the accuracy of the AI / ML model starts to decline when analyzing the evaluation data (i.e., the model performs well on the training data but starts to perform poorly on the evaluation data), the AI / ML model can undergo more training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is only deployed when the accuracy reaches a certain level or if the accuracy of the trained AI / ML model is better than an existing deployed AI / ML model. In certain embodiments, a collection of trained AI / ML models can be used to complete a task. For example, one model can be trained to recognize images, another model can recognize text, another model can recognize semantic and / or ontological associations, and so on.
[0148] Some embodiments may use a transformer network, such as SentenceTransformers TM , which is a Python TM framework for state-of-the-art sentence, text, and image embeddings. Such a transformer network learns associations of words and phrases with both high and low scores. This trains the AI / ML model to determine which are close to the input and which are not close to the input, respectively. The transformer network can also use field lengths and field types, rather than just using pairs of words / phrases.
[0149] As described above, in some embodiments, NLP techniques (such as word2vec, BERT, GPT-3, ChatGPT, other LLMs, etc.) can be used to facilitate semantic understanding and provide more accurate and more user-friendly answers. Other techniques (such as clustering algorithms) can be used to discover similarities between groups of elements. Clustering algorithms can include, but are not limited to, density-based algorithms, distribution-based algorithms, centroid-based algorithms, hierarchical-based algorithms, K-means clustering algorithms, DBSCAN clustering algorithms, Gaussian mixture model (GMM) algorithms, balanced iterative reducing and clustering using hierarchies (BIRCH) algorithms, etc. Such techniques can also assist in classification.
[0150] Figure 7 FIG. is a flowchart showing a process 700 for training an AI / ML model(s) according to an embodiment of the present invention. In some embodiments, as described above, the AI / ML model(s) can be generative AI models. The neural network architecture of an AI / ML model generally includes multiple layers of neurons, including input, output, and hidden layers. For example, see Figure 6A and Figure 6B . The hidden layers are between processing the input data and generating an intermediate representation of the input for generating the output. These hidden layers can include various types of neurons, such as convolutional neurons, recurrent neurons, and / or transform neurons.
[0151] At 710, the training process begins by providing an RPA workflow, source files, UI drawings and screenshots, task mining information, etc., whether labeled or unlabeled. Then at 720, the AI / ML model is trained over multiple epochs, and at 730, the results are reviewed. Although various types of AI / ML models can be used, LLMs and other generative AI models are typically trained using a process called "supervised learning", which was also discussed above. Supervised learning involves providing the model with a large dataset that the model uses to learn the relationship between the input and the output. During the training process, the model adjusts the weights and biases of the neurons in the neural network to minimize the difference between the predicted output and the actual output in the training dataset.
[0152] In some embodiments, one aspect of the model is the use of transfer learning. For example, transfer learning can utilize a pre-trained model, such as ChatGPT that is fine-tuned for a specific task or domain in step 720. This allows the model to leverage the knowledge already learned from the pre-training phase and adapt it to a specific application via the training phase in step 720.
[0153] The pre-training phase involves training the model on an initial training dataset that can be more general. During this phase, the model learns the relationships in the data. In the fine-tuning phase (e.g., in some embodiments, if the pre-trained model is used as the initial basis for the final model, it is performed during step 720 in addition to or instead of the initial training phase), the pre-trained model is adapted to a specific task or domain by training the model on a smaller task-specific dataset. For example, in some embodiments, the model can focus on a certain type(s) of data source. This can help the model identify the data elements therein more accurately than a generative AI model pre-trained alone. Fine-tuning allows the model to learn the nuances of the source, such as specific vocabulary and grammar, certain graphical features, certain data formats, etc., without requiring as much data as would be necessary to train the model from scratch. By leveraging the knowledge learned in the pre-training phase, the fine-tuned model can achieve state-of-the-art performance on a specific task with relatively little additional training data.
[0154] If the AI / ML model fails to meet the desired confidence threshold at 740, the training data is supplemented and / or the reward function is modified at 750 to help the AI / ML model better achieve its goal, and the process returns to step 720. If the AI / ML model meets the confidence threshold at 740, at 760, the AI / ML model is tested on the evaluation data to ensure that the AI / ML model generalizes well and the AI / ML model does not overfit with respect to the training data. The evaluation data includes information that the AI / ML model has not processed before. If the confidence threshold of the evaluation data at 770 is met, at 780, the AI / ML model is deployed. If not, the process returns to step 750 and the AI / ML model is further trained.
[0155] Figures 8A to 8C is a screenshot showing an automated code generation interface 800 configured to analyze natural language input from a user and prompt automation according to an embodiment of the present invention. The automated code generation interface includes a text box 810 in which the user can enter text related to the desired task to be automated. The automated code generation interface 800 also includes a cancel button 820 and a confirmation button 822, and the confirmation button 822 is Figure 8A and Figure 8B is inactive.
[0156] In Figure 8B, the user enters text in text box 810 to automate the order review process by the financial manager. Specifically, the user types "Orders greater than $10k are reviewed by the financial manager; if approved, the order is updated in Salesforce; if rejected, the order is canceled." When the user enters text in text box 810 and presses the enter key, the automation generation status indicator 830 indicates that the automation is generating prompts based on the user input.
[0157] Steering Figure 8C , the automatic code generation interface 800 prompts for the automation of the requested task. A rule 840 is created for the case where the order exceeds $10,000.00. When this condition is met, the financial manager review user task 850 is executed. The (multiple) assigned reviewers are able to perform review options 860 - that is, the (multiple) reviewers can approve the order or cancel the order. If the prompted automation is correct, the user can confirm it using the confirmation button 822, and then generate the automation (e.g., performed by (multiple) RPA robots). When the process has been generated, in some embodiments, the user can use the RPA designer application to drill into the RPA workflow and make any desired edits to the workflow.
[0158] Figure 9 The AI / ML model of the cognitive AI layer 900 according to an embodiment of the present invention is shown. In some embodiments, the AI / ML model of the cognitive AI layer 900 according to an embodiment of the present invention can be used. Figure 7 The process 700 is used to train the model of the cognitive AI layer 900. The generative AI model 910 provides the results to other cognitive AI layer models in a series configuration 940, a parallel configuration 942, or a combination 944 of (multiple) series and parallel configurations. In some embodiments, the cognitive layer can be a single generative AI model. For example, the generative AI model can generate code, provide semantic associations between text on the screen, suggest expressions and code snippets for RPA workflow activities, generate applications based on UI drawings, provide descriptions based on records of user actions, and so on. The CV model 920 and the OCR model also provide detected graphic elements and recognized text to the generative AI model 910 and other cognitive AI layer models in configurations 940, 942, or 944, respectively.
[0159] Other cognitive AI layer models in configurations 940, 942, or 944 use the outputs from generative AI model 910, CV model 920, and / or OCR model 930 to provide the automated code generation functionality discussed in detail herein. The cognitive AI layer 900 can facilitate an understanding of what code should be added in an RPA workflow, what code should be converted to another language in the workflow, which RPA workflows are recommended based on user actions, what descriptions should be generated based on a record of a user performing tasks on a computing system, what applications, UI drawings, combinations thereof, etc. should be generated based on those actions. The output (i.e., result 950) from other cognitive AI layer models in configurations 940, 942, or 944 can include RPA workflows with expressions and / or code in the relevant programming languages included therein, newly generated RPA workflows, human-like descriptions of tasks that a user is performing on a computing system, applications, etc.
[0160] Figure 10 is a flowchart showing a process 1000 for automated code generation for RPA according to an embodiment of the present invention. In some embodiments, the process begins at 1020 by receiving source information from a user. However, in other embodiments, this can be provided to the system automatically. At 1020, the source data is provided to the cognitive AI layer. The cognitive AI layer is one or more AI / ML models. In some embodiments, the cognitive AI layer is part of a computer program that performs the automated code generation process, such as an RPA designer application, an RPA robot, etc. The source data can include, but is not limited to, RPA workflows, natural language statements, recordings of user speech, video recordings of user actions on a computing system, source documents (e.g., spreadsheet files, JSON files, XAML files, XML files, HTML files, documents, PDF files, scanned images of physical documents, etc.), pseudocode for a desired task, drawings of a desired user interface or form, drawings of a process, any combination thereof, etc.
[0161] Then, at 1030, the source data is processed by the cognitive AI layer (i.e., the cognitive AI layer runs the input source information through its AI / ML model(s)), and the cognitive AI layer provides an output at 1040. For example, the output from the cognitive AI layer can be provided to an RPA designer application, an RPA robot, or some other process that can usefully implement the automatically generated code. The output from the cognitive AI layer can include, but is not limited to, existing RPA workflows with expressions and / or code snippets in the relevant programming language(s) included therein, newly generated RPA workflows, human-like descriptions of tasks that a user is performing on a computing system, applications, expressions and / or code snippets, etc.
[0162] If the user accepts the automatically generated code from the cognitive AI layer at 1050, the automatically generated code is implemented at 1060. For example, the RPA workflow is updated in the RPA designer application by adding expressions and / or code snippets to the activities of the RPA designer application and / or its configuration, the RPA workflow generated by the cognitive AI layer is compiled into an automation for execution by the RPA robot at runtime, the description of the computer-implemented task is saved to a database and provided to other users, the application generated by the cognitive AI layer is deployed to the user computing system, and so on. In some embodiments, this can be done automatically without manual user acceptance. In certain embodiments, the implementation of the automatically generated code at 1050 includes adding security and / or compliance rules to the RPA workflow for compliance with regulations and / or policies. In some embodiments, the implementation of the automatically generated code includes compiling the RPA workflow into an automation configured to be executed by the RPA robot at runtime.
[0163] If the user does not accept the output of the cognitive AI layer at 1050 and the modifications made by the user at 1070 can fix the automatically generated code (e.g., the user can change an expression, correct the syntax in the automatically generated document, etc.), then the user can make the changes at 1080 and the automatically generated code can be implemented at 1060. However, in some cases, this is not possible. For example, the user may not know how to modify the code of the automatically generated application to fix the errors in it, the output from the cognitive AI layer may be incomprehensible to the user, the proposed RPA workflow / automation may be incorrect, and so on. In such cases, at 1090, the user can fail the automatic code generation process and can send data related to the failure for retraining the AI / ML model(s) of the cognitive AI layer. For example, the user can provide a text description of which elements of the automatically generated code are incorrect and why.
[0164] Figure 11 is a flowchart of a process 1100 for automatic document generation using a cognitive AI layer according to an embodiment of the present invention. The process begins with receiving task mining output related to the user's actions at 1110. The task mining output can include, but is not limited to, what applications the user is using, what information is being input into those applications, keystrokes, mouse clicks and positions, graphical elements in the UI, which graphical elements are active elements at different times (e.g., the current element the user is typing in, the position of the cursor, etc.), and so on.
[0165] The task mining output is provided as input to the cognitive AI layer and processed at 1120. The cognitive AI layer includes an LLM that is configured to analyze the task mining input and generate text (and in some embodiments, generate images) based on the task mining output. Thus, at 1130, the cognitive AI layer determines what the user is doing based on the task mining output and records the process. For example, the LLM of the cognitive AI layer can be used to generate a human-like description of the actions the user is performing, potentially including screenshots. Then, at 1140, this information can be provided as a manual to other users who want to perform the corresponding task.
[0166] Figure 12 is a flowchart showing a process 1200 for generating automation from natural language text according to an embodiment of the present invention. The process begins at 1210 by receiving natural language text or an audio recording from a user. In the latter case, the audio is converted to text using a speech-to-text model. In some embodiments, the natural language text can be input by the user in an application having an automatic code generation interface. For example, see Figures 8A to 8C .
[0167] At 1220, a cognitive AI layer having an LLM receives the natural language text as input and processes the text. An application having an automatic code generation interface receives, at 1230, the output from the cognitive AI layer, which includes information related to the automation of the prompt, and displays, at 1240, the automated steps of the prompt. For example, rules, tasks, steps, etc. associated with the automation of the prompt can be displayed.
[0168] If the user accepts the automation at 1250, the automation is generated at 1260. The automation can be a stand-alone application or an automation executed by an RPA robot. If the user does not accept the automation of the prompt at 1250, then at 1270, the automatic code generation interface requests user input regarding what is incorrect about the automation of the prompt and submits this information for retraining the LLM.
[0169] According to an embodiment of the present invention, the process steps performed in Figure 7 and Figures 10 to 12 can be performed by a computer program encoding instructions for one or more processors to perform at least a portion of the one or more processes described in Figure 7 and Figures 10 to 12 The computer program can be embodied on a non-transitory computer-readable medium. The computer-readable medium can be, but is not limited to, a hard drive, a flash device, RAM, magnetic tape, and / or any other such medium or combination of media for storing data. The computer program can include instructions for controlling one or more processors of a computing system (e.g., Figure 5The (multiple) processors 510 of computing system 500 are implemented to Figure 7 and Figures 10 to 12 encode instructions that implement all or part of the processing steps described in
[0170] These encoded instructions may also be stored on a computer-readable medium. A computer program may be implemented in hardware, software, or a hybrid implementation. A computer program may consist of modules that communicate operationally with each other and are designed to transfer information or instructions for display. A computer program may be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0171] It will be readily understood that, as generally described and illustrated in the accompanying drawings herein, the components of the various embodiments of the present invention may be arranged and designed in a variety of different configurations. Accordingly, the detailed description of the embodiments of the present invention shown in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.
[0172] The features, structures, or characteristics of the present invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, throughout the specification, references to "certain embodiments", "some embodiments", or similar language mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in certain embodiments", "in some embodiments", "in other embodiments", or similar language throughout this specification do not necessarily all refer to the same set of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0173] It should be noted that throughout this specification, references to features, advantages, or similar language do not mean that all features and advantages that can be achieved using the present invention should be or are present in any single embodiment of the present invention. Instead, language referring to features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the present invention. Thus, throughout this specification, discussions of features and advantages and similar language may, but do not necessarily, refer to the same embodiments.
[0174] In addition, the described features, advantages, and characteristics of the present invention may be combined in any suitable manner in one or more embodiments. A person skilled in the relevant art will recognize that the present invention may be practiced without one or more of the specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present invention.
[0175] A person having ordinary skill in the art will readily understand that the present invention as discussed above can be practiced using steps in a different order and / or using hardware elements in a configuration different from that disclosed. Accordingly, although the present invention has been described based on these preferred embodiments, certain modifications, variations and alternative constructions will be apparent to those skilled in the art while remaining within the spirit and scope of the present invention. Therefore, in order to determine the scope and bounds of the present invention, reference should be made to the appended claims.
Claims
1. A non-transitory computer readable medium storing a computer program, the computer program being configured to cause at least one processor to: providing source data as input to a cognitive AI layer, the cognitive AI layer being configured to process the input; receiving output from the cognitive AI layer including automatically generated code; and The automatic code generation is implemented in a Robotic Process Automation (RPA) workflow.
2. The non-transitory computer-readable medium of claim 1, wherein the source data comprises natural language text, a recording of a user's speech, a video recording of a user's actions on a computing system, a document, pseudocode for a desired task, a drawing of a desired user interface or form, a drawing of a process, a process document or process description, an RPA workflow, or any combination thereof.
3. The non-transitory computer-readable medium of claim 1, wherein the computer program comprises the cognitive AI layer.
4. The non-transitory computer-readable medium of claim 1, wherein the computer program is an RPA designer application.
5. The non-transitory computer-readable medium of claim 1, wherein the output from the cognitive AI layer comprises the RPA workflow.
6. The non-transitory computer-readable medium of claim 1 , wherein the output from the cognitive AI layer comprises one or more expressions and / or code snippets, and the implementation of the automatic code generation comprises updating the RPA workflow by adding the one or more expressions and / or code snippets to one or more corresponding activities of the RPA workflow.
7. The non-transitory computer-readable medium of claim 1 , wherein the output from the cognitive AI layer comprises one or more configuration changes to the RPA workflow, and the implementation of the automatically generated code comprises updating the RPA workflow by changing one or more corresponding activities of the RPA workflow to include the one or more configuration changes.
8. The non-transitory computer-readable medium of claim 1, wherein the implementing of the automatically generating code comprises compiling the RPA workflow into automation configured to be executed by an RPA robot at runtime.
9. The non-transitory computer readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: prompting a user regarding said output from said cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not acceptable; receiving one or more modifications to the RPA workflow from the user; and The RPA workflow is modified to include the output from the cognitive AI layer and the modifications from the user.
10. The non-transitory computer readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: prompting a user regarding said output from said cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is unacceptable; and Failing the automatic code generation and sending data related to the failure for retraining one or more AI / ML models of the cognitive AI layer.
11. The non-transitory computer readable medium of claim 1 , wherein the cognitive AI layer comprises: a generative AI model configured to generate code, provide semantic associations between text on the screen, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate procedures from drawings, or any combination thereof; as well as One or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation functionality.
12. The non-transitory computer readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: Security and / or compliance rules are added to the RPA workflow for compliance with regulations and / or policies.
13. One or more computing systems comprising: Memory, which stores computer program instructions; as well as at least one processor configured to execute the stored computer program instructions, wherein the computer program instructions are configured to cause the at least one processor to: providing source data as input to a cognitive AI layer, the cognitive AI layer being configured to process the input; receiving output from the cognitive AI layer including automatically generated code; as well as The automatic code generation is implemented in a Robotic Process Automation (RPA) workflow, wherein The source data includes natural language text, a recording of a user's speech, a video recording of a user's actions on a computing system, a document, pseudocode for a desired task, a drawing of a desired user interface or form, a drawing of a process, a process document or process description, an RPA workflow, or any combination thereof.
14. The one or more computing systems of claim 13, wherein the computer program instructions comprise the cognitive AI layer.
15. The one or more computing systems of claim 13, wherein the output from the cognitive AI layer comprises the RPA workflow.
16. One or more computing systems according to claim 13, wherein the output from the cognitive AI layer includes one or more expressions and / or code snippets, and the implementation of the automatic code generation includes updating the RPA workflow by adding the one or more expressions and / or code snippets to one or more corresponding activities of the RPA workflow.
17. One or more computing systems according to claim 13, wherein the output from the cognitive AI layer includes one or more configuration changes to the RPA workflow, and the implementation of the automatic code generation includes updating the RPA workflow by changing one or more corresponding activities of the RPA workflow to include the one or more configuration changes.
18. The one or more computing systems of claim 13, wherein the implementing of the automatically generated code comprises compiling the RPA workflow into automation configured to be executed by an RPA robot at runtime.
19. The one or more computing systems of claim 13, wherein the computer program instructions are further configured to cause the at least one processor to: prompting the user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not acceptable; receiving one or more modifications to the RPA workflow from the user; and The RPA workflow is modified to include the output from the cognitive AI layer and the modifications from the user.
20. The one or more computing systems of claim 13, wherein the computer program instructions are further configured to cause the at least one processor to: prompting the user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is unacceptable; and Failing the automatic code generation and sending data related to the failure for retraining one or more AI / ML models of the cognitive AI layer.
21. The one or more computing systems of claim 13, wherein the cognitive AI layer comprises: a generative AI model configured to generate code, provide semantic associations between text on the screen, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate procedures from drawings, or any combination thereof; as well as One or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation functionality.
22. The one or more computing systems of claim 13, wherein the computer program instructions are further configured to cause the at least one processor to: Security and / or compliance rules are added to the RPA workflow for compliance with regulations and / or policies.
23. A computer-implemented method for performing automatic code generation, comprising: providing, by a computing system, source data as input to a cognitive AI layer, the cognitive AI layer being configured to process the input; receiving, by the computing system, output from the cognitive AI layer including automatically generated code; as well as The automatic code generation is implemented by the computing system in a robotic process automation (RPA) workflow, wherein The source data includes natural language text, a recording of a user's speech, a video recording of a user's actions on a computing system, a document, pseudocode for a desired task, a drawing of a desired user interface or form, a drawing of a process, a process document or process description, an RPA workflow, or any combination thereof.
24. A computer-implemented method according to claim 23, wherein the output from the cognitive AI layer includes one or more expressions and / or code snippets, and the implementation of the automatic code generation includes updating the RPA workflow by adding the one or more expressions and / or code snippets to one or more corresponding activities of the RPA workflow.
25. A computer-implemented method according to claim 23, wherein the output from the cognitive AI layer includes one or more configuration changes to the RPA workflow, and the implementation of the automatic code generation includes updating the RPA workflow by changing one or more corresponding activities of the RPA workflow to include the one or more configuration changes.
26. The computer-implemented method of claim 23, wherein the implementing of the automatically generating code comprises compiling the RPA workflow into an automation configured to be executed by an RPA robot at runtime.
27. The computer-implemented method of claim 23, further comprising: prompting the user, by the computing system, regarding the output from the cognitive AI layer; receiving, by the computing system, an indication by the user that the output from the cognitive AI layer is not acceptable; receiving, by the computing system, one or more modifications to the RPA workflow from the user; as well as The RPA workflow is modified by the computing system to include the output from the cognitive AI layer and the modifications from the user.
28. The computer-implemented method of claim 23, further comprising: prompting the user, by the computing system, regarding the output from the cognitive AI layer; receiving, by the computing system, an indication by the user that the output from the cognitive AI layer is not acceptable; as well as The automatic code generation is failed by the computing system and data related to the failure is sent for retraining one or more AI / ML models of the cognitive AI layer.
29. The computer-implemented method of claim 23, wherein the cognitive AI layer comprises: a generative AI model configured to generate code, provide semantic associations between text on the screen, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate procedures from drawings, or any combination thereof; as well as One or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation functionality.
30. The computer-implemented method of claim 23, further comprising: Security and / or compliance rules are added to the RPA workflow for compliance with regulations and / or policies.
Citation Information
Patent Citations
Long running workflows for document processing using robotic process automation
US10860905B1
Determining sequences of interactions, process extraction, and robot generation using artificial intelligence / machine learning models
US11301269B1
Robot access control and governance for robotic process automation
US11733668B2
User interface (UI) descriptors, UI object libraries, UI object repositories, and UI object browsers for robotic process automation
US11748069B2
Training an artificial intelligence / machine learning model to recognize applications, screens, and user interface elements using computer vision
US20220113991A1
Cited By
Telemetering general analysis real-time compiling method and system and computer readable storage medium
CN120743285A