Automatic code generation for robotic process automation
A cognitive AI layer in RPA technologies automatically generates code from various inputs, addressing the limitations of current RPA systems by simplifying workflow development and reducing errors.
Patent Information
- Application Number
- JP2024066696
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-04-17
- Publication Date
- 2025-06-30
AI Technical Summary
Current robotic process automation (RPA) technologies require advanced programming knowledge and specific language expertise, limiting accessibility and increasing the risk of errors during workflow development.
The implementation of a cognitive AI layer that automatically generates computer program code from various input sources, including natural language text, recordings, and diagrams, enabling RPA workflow development without extensive programming knowledge.
This solution simplifies RPA workflow development, reduces the need for coding expertise, and minimizes errors by converting diverse input sources into executable code, thereby enhancing accessibility and efficiency.
Smart Images

Figure 2025097252000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to artificial intelligence (AI), and more specifically, to automatic code generation for robotic process automation (RPA) that uses a cognitive AI layer to convert an input source into computer program code.
Background Art
[0002] During the development of RPA workflows, developers may need to construct formulas or create and call snippets of code written in Python, Visual Basic (VB), C#, C++, Java, C, etc. However, this requires both programming knowledge and knowledge of the specific programming language being used. Therefore, only relatively advanced users can utilize these functions, and even experienced developers may accidentally introduce errors into the RPA workflow logic. Thus, an improved and / or alternative approach to RPA workflow development may be beneficial.
Summary of the Invention
[0003] Certain embodiments of the present invention may provide solutions to problems and needs in the art that are not yet fully specified, evaluated, or solved by current RPA technologies. For example, some embodiments of the present invention relate to automatic code generation for RPA that uses a cognitive AI layer to convert an input source into computer program code.
[0004] In an embodiment, a non-transitory computer-readable medium stores a computer program. The computer program is configured such that at least one processor provides source data as an input to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer program is also configured such that at least one processor receives an output from the cognitive AI layer that includes automatically generated code and implements the automatically generated code in an RPA workflow.
[0005] In another embodiment, one or more computing systems include a memory storing computer program instructions and at least one processor configured to execute the stored computer program instructions. The computer program instructions are configured such that at least one processor provides source data as an input to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer program instructions are also configured such that at least one processor receives an output from the cognitive AI layer that includes automatically generated code and implements the automatically generated code in an RPA workflow. The source data includes natural language text, recordings of a user's voice, video recordings of a user's actions on a computing system, documents, pseudocode of a desired task, diagrams of a desired user interface or form, diagrams of a process, documents or descriptions of a process, RPA workflows, or any combination thereof.
[0006] In yet another embodiment, a computer-implemented method for performing automatic code generation includes providing source data as an input by a computing system to a cognitive AI layer. The cognitive AI layer is configured to process the input. The computer-implemented method also includes receiving, by the computing system, an output from the cognitive AI layer that includes code automatically generated by the computing system. The computer-implemented method further includes implementing, by the computing system, the code automatically generated in an RPA workflow. The source data includes natural language text, recordings of a user's voice, video recordings of a user's actions on a computing system, documents, pseudocode of a desired task, diagrams of a desired user interface or form, diagrams of a process, process documents or process descriptions, RPA workflows, or any combination thereof.
Brief Description of the Drawings
[0007] To facilitate an understanding of the advantages of particular embodiments of the present invention, a more detailed description of the invention briefly described above is depicted with reference to specific embodiments illustrated in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the invention and are not to be considered limiting of its scope, but that the invention will be described and explained with further particularity and detail by use of the following accompanying drawings.
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
[0012]
Figure 5
[0013]
Figure 6A
[0014]
Figure 6B
[0015]
Figure 7
[0016]
Figure 8A
Figure 8B
Figure 8C
[0017]
Figure 9
[0018]
Figure 10
[0019]
Figure 11
[0020]
Figure 12
[0021] Unless otherwise specified, similar reference characters denote corresponding features consistently throughout the accompanying drawings.
DETAILED DESCRIPTION OF THE INVENTION
[0022] (Detailed description of the embodiment) Some embodiments relate to automatic code generation for RPA that uses a cognitive AI layer to automatically generate computer program code, text, or other items based on an input source. Such embodiments can remove language barriers due to RPA developers having no programming experience or programmers not knowing one or more programming languages. To remove these language barriers, some embodiments convert text or another input source into expressions or code snippets of one or more programming languages. An expression is a syntactic entity of a programming language that can be evaluated to determine its value. An expression is a combination of one or more constants, variables, functions, and operators, and the programming language interprets them according to its specific rules of precedence and association and computes them to generate another value. A code snippet is a relatively small area of reusable source code, machine code, or text.
[0023] Regarding formulas and snippets, there may be cases where RPA developers need to adjust the RPA workflow code according to various purposes. Even if RPA workflow development can be a relatively less code-intensive process compared to other types of programming, some degree of customization can be achieved through the code of formulas and snippets. Usually, this code does not fundamentally change the scope of the RPA workflow but rather enhances it.
[0024] Various programming languages can be supported by RPA designer applications. For example, UiPath Studio (trademark) currently supports Python, VB, and C#. To utilize these features, RPA developers usually need to know each programming language. However, in some embodiments, generative AI models such as large language models (LLMs) can be used to create formulas or snippets of code that execute specific actions in response to user requests or in response to the interpretation of input sources. Such embodiments can be further enhanced by a coded workflow or a coded test case based on the user request or input source. A coded workflow or test case is a way to describe the code for an RPA robot. Developers can use C# code or the code of some other languages instead of the low-code visual RPA workflow designer. This way, all the capabilities of.NET are put in the hands of RPA developers, and the advantages of the platform of the RPA designer application (such as RPA robot, orchestrator, insights, test manager, etc.) can be obtained.
[0025] In response to user requests or input sources, some embodiments generate expressions that are part of a workflow activity or activity configuration (e.g., a specific expression can be assigned to derive specific activity parameter values). Certain embodiments may compose code in different call sections or snippets. For example, some embodiments can generate activities from a UI object repository or some other source with a specific configuration, or entirely new code can be generated. Thus, the cognitive AI layer can generate expressions and code snippets to be incorporated into the RPA workflow.
[0026] The input source can take various forms. For example, a user can input a natural language sentence and provide a source document (e.g., a spreadsheet file, a JavaScript Object Notation (JSON) file, an Extensible Application Markup Language (XAML) file, an Extensible Markup Language (XML) file, a Hypertext Markup Language (HTML) file, a Word® document, a PDF (Portable Document Format) file, a scanned image of a physical document, etc.) to describe the pseudo-code for the desired task. In some embodiments, the cognitive AI layer may be able to propose RPA automation based on this input alone.
[0027] Consider the case where the PDF is a claim document. The cognitive AI layer can learn that when the type of PDF is a claim document, the user usually wants to extract data from it every time they receive the claim document. The cognitive AI layer can propose automation to the user as needed and automatically generate the automation. This can include, for example, opening the claim document, performing optical character recognition (OCR) on the claim document if it has been scanned, extracting text and numerical values from it, and saving this information to a claims system or the like, including creating an RPA workflow with appropriate activities.
[0028] Any appropriate information can be used as a source for such automation. For example, the cognitive AI layer can use a process definition document (PDD), an information technology (IT) policy, an approval interface, etc. as a source. Next, the source information can be used to generate a high-level automation that the user can further adjust as needed. In some embodiments, the cognitive AI layer can add security and / or compliance rules to the generated RPA workflow to ensure that the workflow complies with laws and / or policies. In some embodiments, task mining outputs (e.g., the applications the user is using, the information input into these applications, key presses, mouse clicks and positions, graphical elements of the user interface (UI), which graphical elements are active elements at various times, etc.) can be used to assist in generating the RPA workflow. See, for example, U.S. Patent Application Publication No. 2022 / 0113991 and U.S. Patent No. 11,301,269, which are hereby incorporated by reference in their entireties.
[0029] In some embodiments, the cognitive AI layer can determine what the user was doing based on the task mining output and document the process. For example, using the LLM of the cognitive AI layer, a human - like description of the actions the user is performing can be generated, which may include screenshots in some cases. This information can then be provided as a manual to other users who want to perform each task.
[0030] In some embodiments, automation can be created from such records. By understanding what the user is doing and mapping the actions to RPA workflow activities, the cognitive AI layer can automatically create an RPA workflow based on the actions performed within the record. This can enable the automation of the user's tasks even without the user having general knowledge of RPA.
[0031] In some embodiments, diagrams representing the design of the screen can be used to automatically generate the relevant applications. Consider the case where the user designs a Visio (registered trademark) diagram of the UI of a screen, such as a form for collecting specific data and sending it to a customer relationship management (CRM) system. The cognitive AI layer can use a CV model to understand what graphical elements and text exist within the screen, and the generative AI model can create a software application that includes these graphical elements and a sending function to the CRM system, such as retrieving customer information and sending this information to Salesforce (registered trademark). Thus, the user provides the form, and the cognitive AI layer can construct the UI of the form. Such embodiments can essentially convert the description of graphical elements and text into code.
[0032] In some embodiments, the user can submit a diagram of the process to be automated. Then, the high - level steps of this process can be generated as an RPA workflow. For example, the generative AI model can be trained to understand the text within the steps indicated by the connectors and the relationships between the steps.
[0033] When developing automation using RPA tools such as UiPath's Form Builder (trademark), a manual customization and time-consuming process may be required to import data sources. Also, creating the setup of dynamic form elements such as dropdown lists is a time-consuming process. Furthermore, developers with specialized coding skills and knowledge need to incorporate complex functions during the development of automation. For example, deep knowledge and specialized coding skills are required to write code or incorporate complex functions using activities such as <code call>.
[0034] Accordingly, some embodiments utilize natural language processing (NLP) to automate form building and code generation in automation development. The natural language description is provided by the user (e.g., using the interface 800 shown in FIGS. 8A-8C). Then, an LLM is used to process this description and understand the user's intent. As a result, a form is automatically built or other code is automatically generated based on the user's description.
[0035] The natural language description can be used to define the appearance of the user interface. For example, the user can enter a description of the query in natural language in the RPA designer application and generate a form for insurance claims that includes the name of the insurance policyholder, the insurance number, the details of the claim, and the date of the accident. Next, the RPA designer application automatically generates the form, enters the required fields into the form, and configures the necessary validations and appropriate input types for the fields.
[0036] In some embodiments, the user can also add new fields to the generated form. For example, using the example from the previous paragraph, the user can provide a natural language description to add fields for the estimated claim amount and the support document. Thus, the tool adds or updates the requested information within the form. Further, using the fully functional form, various desired types of automation can be built. For example, the input for the automation can be obtained from the form and an email describing the policy details can be sent to the user.
[0037] The LLM of some embodiments converts user descriptions into machine-executable code. For example, a code call activity can be used and a user description in natural language such as "Calculate the total number of days since the accident occurred based on the accident date, and if it exceeds 10 days, a set of confirmations is required" can be provided. Further, in some embodiments, the code can be generated in any suitable programming language based on the user description. To accommodate a particular programming language, the code generated by the LLM(s) can be adjusted.
[0038] The embodiments disclosed herein can provide various advantages over existing automation technologies. Since the manual effort associated with form construction and code generation for other purposes in RPA is reduced, the implementation of the process is accelerated and the need for coding expertise can be reduced. This can increase the number of RPA projects that an organization can create and manage and can shorten the development time per project. Developers with coding expertise can focus on designing efficient workflows rather than creating forms or the accompanying code.
[0039] FIG. 1 is an architectural diagram showing a hyper-automation system 100 according to an embodiment of the present invention. As used herein, "hyper-automation" refers to an automation system that combines components of process automation, integration tools, and technologies that amplify the ability to automate work. For example, in some embodiments, robotic process automation (RPA) is used at the core of the hyper-automation system, and in certain embodiments, the automation capabilities can be extended by AI / machine learning (ML), process mining, analytics, and / or other advanced tools. When the hyper-automation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work can be automated, and both computing systems within an organization, such as those used by individuals and those that operate autonomously, can all participate as part of the hyper-automation process. The hyper-automation systems of some embodiments enable users and organizations to discover, understand, and expand automation efficiently and effectively.
[0040] The hyper-automation system 100 includes user computing systems such as desktop computer 102, tablet 104, and smartphone 106. However, any desired user computing system, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc., can be used without departing from the scope of the present invention. Also, although three user computing systems are shown in FIG. 1, any suitable number of user computing systems can be used without departing from the scope of the present invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of user computing systems can be used. The user computing systems may be actively used by users or may be automatically executed without much or any user input.
[0041] Each of the user computing systems 102, 104, 106 has its respective automation process(es) 110, 112, 114 running thereon. In some embodiments, the automation process is stored remotely (e.g., on server 130 or database 140) and accessed via network 120 and is loaded by an RPA robot to implement the automation. The automation can exist as a script (e.g., XML, XAML, etc.) or can be compiled into machine-readable code (e.g., as a digital link library).
[0042] The automation process(es) 110, 112, 114 can include, without limitation and without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application(s) for each computing system, any other suitable software and / or hardware, or any combination thereof. In some embodiments, one or more of the process(es) 110, 112, 114 can be a listener. The listener can be, without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. Indeed, in some embodiments, the logic of the listener(s) is implemented partially or fully via physical hardware.
[0043] The listener monitors and records data related to user interactions with each computing system and / or the operation of unattended computing systems, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data can include, but is not limited to, which buttons were clicked, where the mouse moved, the text entered in fields, that one window was minimized and another window was opened, the applications associated with the windows, etc. In certain embodiments, the data from the listener can be transmitted periodically as part of a heartbeat message. In some embodiments, the data can be transmitted to the core hyper-automation system 120 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0044] The automation process can execute logic developed in a workflow during design time. In the case of RPA, the workflow can include a set of steps performed in a sequence or some other logical flow, defined herein as an "activity". Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, the workflow can be nested or embedded.
[0045] The long-running workflows for RPA in some embodiments are master projects that support service orchestration, human intervention, and long-running transactions in an unattended environment. See, for example, U.S. Patent No. 10,860,905, which is hereby incorporated by reference in its entirety. Human intervention occurs when a particular process requires human input for exception handling, approval, or verification before proceeding to the next step of the activity. In this case, the execution of the process is paused and the RPA robot is released until the human task is completed.
[0046] The long-running workflow may support fragmentation of the workflow via persistence activities, combined with call processes and non-user interaction activities, and may orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or a large number of computing systems may participate in the execution of the logic of the long-running workflow. The long-running workflow may be executed in a session to facilitate rapid execution. In some embodiments, the long-running workflow may orchestrate a background process that executes API calls and may include activities that execute in the long-running workflow session. These activities may be called by call process activities in some embodiments. A process having a user interaction activity that executes in a user session may be called by starting a job from a conductor activity (the conductor is described in more detail later in this specification). The user may interact in some embodiments through a task that requires the user to complete a form in the conductor. An activity may be included that causes the RPA robot to wait for the form task to be completed and then resume the long-running workflow.
[0047] One or more automation processes 110, 112, 114 communicate with a core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may execute a conductor application on one or more servers such as server 130. Although one server 130 is shown for illustration purposes, a plurality or a number of servers in proximity to each other or in a distributed architecture may be employed without departing from the scope of the present invention. For example, one or more servers may be provided for conductor functionality, AI / ML model provision, authentication, governance, and / or any other suitable functionality without departing from the scope of the present invention. In some embodiments, the core hyper-automation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyper-automation system 120 may host multiple software-based servers on one or more computing systems such as server 130. In some embodiments, one or more servers of the core hyper-automation system 120 such as server 130 may be implemented via one or more virtual machines (VMs).
[0048] In some embodiments, one or more automation processes 110, 112, 114 may invoke one or more AI / ML models 132 that are deployed on or accessible by a core hyper-automation system 120 and trained to accomplish various tasks. For example, the AI / ML models 132 may include models trained to search for various application versions, perform CV, perform OCR, generate UI descriptors, and provide suggestions for the next activity or sequence of activities in an RPA workflow. The AI / ML models may be trained using labeled data including elements of data sources (e.g., web pages, forms, scanned documents, application interfaces, screens, etc.), previously created RPA workflows, screenshots of various application screens of various versions, including corresponding UI elements, libraries of UI objects, and the like. The AI / ML models 132 may be trained to achieve a desired confidence threshold without overfitting to a given set of training data.
[0049] The AI / ML model 132 can be trained for any suitable purpose without departing from the scope of the present invention, as will be discussed in more detail later in this specification. In some embodiments, two or more AI / ML models 132 may be chained (e.g., in series, parallel, or a combination thereof) such that they collectively provide a collaborative output(s). The AI / ML model 132 may perform or assist with CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automatic RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, any combination thereof, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of the present invention. By using multiple AI / ML models, for example, the system can develop an overall picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, etc. Patterns may be determined individually by an AI / ML model or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one of the computing systems 102, 104, 106.
[0050] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that executes on data, and the AI / ML model itself can be, for example, a deep learning neural network (DLNN) of artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions such as statistical modeling (e.g., hidden Markov model (HMM)), and may utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform the desired functions.
[0051] In some embodiments, the Hyper Automation System 100 may provide four main functional groups: (1) discovery, (2) automation construction, (3) management, and (4) engagement. Automation (e.g., executed on user computing systems, servers, etc.) may, in some embodiments, be performed by software robots such as RPA robots. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with users to assist them in tasks (e.g., via UiPath Assistant™). Unattended robots operate independently of users and may potentially execute in the background without the users' knowledge. Test robots are unattended robots that execute test cases against applications or RPA workflows. Test robots may, in some embodiments, be executed in parallel on multiple computing systems.
[0052] The discovery function may discover various opportunities for automating business processes and provide automated recommendations therefor. Such a function may be implemented by one or more servers such as server 130. The discovery function may, in some embodiments, include providing automation hubs, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath Automation Hub™) may provide a mechanism for managing the rollout of automation with visibility and control. Automation ideas may be crowdsourced from employees, for example, via a submission form. Feasibility and ROI calculations for automating these ideas are provided, documentation for future automation is collected, and collaboration for quickly moving from discovery to construction of automation may be provided.
[0053] Process mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Center™) refers to the process of collecting and analyzing data from applications (such as enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how they can be effectively automated, and the impact of automation. This data can be obtained, for example, by a listener from user computing systems 102, 104, 106 and processed by a server such as server 130. In some embodiments, one or more AI / ML models 132 may be employed for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of the goals of process mining include, but are not limited to, increased profitability, improved customer satisfaction, regulatory and / or compliance, and improved employee efficiency.
[0054] (For example, via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) Task mining identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in routine tasks, and scores such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to reveal repetitive task patterns in the data. Repetitive tasks ripe for automation can then be identified. This information can initially be provided by a listener and, in some embodiments, analyzed on a server of a core hyperautomation system 120 such as server 130. Discoveries from task mining (e.g., XAML process data) are exported to a process document or a designer application such as UiPath Studio (trademark) to enable faster creation and deployment of automation. Task mining in some embodiments may include taking screenshots with user actions (e.g., mouse click location, keyboard input, application windows and graphical elements the user interacted with, timestamps for the interaction, etc.), collecting statistical data (e.g., execution time, number of actions, text input, etc.), editing and annotating the screenshots, specifying the types of actions being recorded, and so on.
[0055] (Via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) Task capture automatically documents attended processes while the user is working or provides a framework for unattended processes. Such documentation may include process definition documents (PDDs), skeleton workflows, capture of actions for each part of the process, recording of user actions, and automatic generation of comprehensive workflow diagrams including details about each step, tasks that are desirably automated in formats such as Microsoft Word (registered trademark) documents, XAML files, etc. The constructible workflows can, in some embodiments, be directly exported to designer applications such as UiPath Studio (trademark). Task capture can simplify the requirements gathering process for both subject matter experts who explain the process and members of the Center of Excellence (CoE) who provide production-grade automation.
[0056] Automation can be achieved through designer applications (such as UiPath Studio (trademark), UiPath StudioX (trademark), UiPath Studio Web (trademark), etc.). For example, RPA developers in the RPA development facility 150 can use the RPA designer application 154 on the computing system 152 to build and test automation for various applications and environments such as web, mobile, SAP (registered trademark), and virtual desktops. API integration can be provided for various applications, technologies, and platforms. Pre-defined activities, drag-and-drop modeling, and workflow recorders can facilitate automation with minimal coding. The document understanding function can be provided through drag-and-drop AI skills for data extraction and interpretation that call one or more AI / ML models 132. Such automation can process virtually any document type and format, including tables, checkboxes, signatures, and handwritten. When data is verified or exceptions are processed, this information can be used to retrain the respective AI / ML models, and their accuracy is improved over time.
[0057] The RPA designer application 152 can be designed to call one or more of the trained AI / ML models 132 on the server 130 and / or the generative AI model 172 within the cloud environment via the network 120 (such as a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.) to assist in the RPA automation development process. In some embodiments, one or more of the AI / ML models can be packaged with the RPA designer application 152 or otherwise stored locally on the computing system 150.
[0058] In some embodiments, the RPA Designer applications 152 and the one or more AI / ML models 132 can be configured to use an object repository stored in the database 140. For example, see U.S. Patent No. 11,748,069, which is hereby incorporated by reference in its entirety. The object repository can include a library of UI objects that can be used to develop RPA workflows via the RPA Designer application 152. The object repository can be used to add UI descriptors to activities within the workflows of the RPA Designer application 152 for UI automation. In some embodiments, one or more of the AI / ML models 132 can generate new UI descriptors and add them to the object repository within the database 140. When automation is completed in the Designer application 152, the automation can be published on the server 130 and pushed out to computing systems 102, 104, 106, etc.
[0059] With the integrated service, developers can seamlessly combine, for example, UI automation and API automation. Automation that requires APIs or that crosses both API and non-API applications and systems can be built. A repository (e.g., UiPath Object Repository (trademark)) or marketplace (e.g., UiPath Marketplace (trademark)) for pre-built RPA and AI templates and solutions can be provided so that developers can automate a wide variety of processes more quickly. Thus, when building automation, the hyper-automation system 100 can provide a user interface, a development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. The hyper-automation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.
[0060] In some embodiments, components of the hyperautomation system 100, such as designer applications and / or external rule engines, provide support for managing and enforcing governance policies for controlling the various functions provided by the hyperautomation system 100. Governance is the ability of an organization to introduce policies to prevent users from developing automation (such as RPA robots) that can harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), the terms of use of third-party applications, and so on. Otherwise, developers could create automation that violates privacy laws, terms of use, etc. during the execution of their automation. Thus, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. This can provide an additional level of security and compliance in the automation process development pipeline in some embodiments by preventing developers from introducing security risks or relying on unapproved software libraries that could operate in a way that violates policies, regulations, privacy laws, and / or privacy policies. See, for example, U.S. Patent No. 11,733,668, which is incorporated herein by reference in its entirety.
[0061] The management function can provide the management, deployment, and optimization of automation across the entire organization. The management function may include, in some embodiments, orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management function of the hyper-automation system 100 can include, among other things, but not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots.
[0062] Conductor applications such as UiPath Orchestrator (trademark) (which may be provided as part of UiPath Automation Cloud (trademark) in some embodiments, or on-premises, VM, private or public cloud, on a Linux (trademark) VM, or as a cloud-native single-container suite via UiPath Automation Suite (trademark)) provide orchestration capabilities to deploy, monitor, optimize, scale, and secure RPA robot deployments. A test suite (e.g., UiPath Test Suite (trademark)) can provide test management for monitoring the quality of deployed automation. The test suite can facilitate test planning and execution, requirement fulfillment, and defect traceability. The test suite can include comprehensive test reports.
[0063] Analytics software (e.g., UiPath Insights (trademark)) can track, measure, and manage the performance of deployed automation. The analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes of the organization. The analytics software can present results in dashboard form for easier understanding by human users.
[0064] A data service (e.g., UiPath Data Service (trademark)) can, for example, be stored in a database 140 and bring data into a single, scalable, and secure place using a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities may be provided by an AI Center (e.g., UiPath AI Center (trademark)), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may enable non-data scientists to access such capabilities. Deployed automation (e.g., an RPA robot) can call an AI / ML model from an AI Center such as AI / ML model 132. The performance of the AI / ML model can be monitored and trained and improved using human-verified data such as that provided by a data review center 160. Human reviewers may provide labeled data to a core hyperautomation system 120 via a review application 152 on a computing system 154. For example, human reviewers may verify that predictions by an AI / ML model 132 and / or a generative AI model 172 are accurate or otherwise provide corrections. This dynamic input may then be saved as training data for retraining the AI / ML model 132 and / or the generative AI model 172 and stored, for example, in a database such as database 140. The AI Center can then schedule and perform a training job to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining the AI / ML model 132 and / or the generative AI model 172.
[0065] The engagement function involves humans and automation as one team for seamless collaboration regarding a desired process. Low-code applications can be built (e.g., via UiPath Apps™) even if they lack an API in some embodiments, to connect browser tabs and legacy software. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. An application can be connected to one automation or multiple automations.
[0066] The Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for passing a process from automation to human or vice versa. A human can provide approvals or escalations and perform exception handling, etc. The automation can then execute the automated functions of a given workflow.
[0067] The local assistant can be provided as a launch pad for the user to start an automation (e.g., UiPath Assistant (trademark)). This feature can be provided, for example, in the tray provided by the operating system, enabling the user to interact with RPA robots and RPA robot - enabled applications on their computing system. The interface can list the automations approved for a given user and allow the user to execute them. These can include off - the - shelf automations from an automation marketplace, an internal automation store in an automation hub, etc. While an automation is running, they can execute as a local instance in parallel with other processes on the computing system so that the user can use the computing system while the automation performs its actions. In certain embodiments, the assistant is integrated with a task capture function so that the user can document the processes that will soon be automated from the assistant's launch pad.
[0068] Chatbots (e.g., UiPath Chatbots (trademark)), social messaging applications, and / or voice commands can enable the user to execute an automation. This can simplify access to the information, tools, and resources necessary to conduct customer interactions or other activities. Human - to - human conversations can be automated as easily as other processes. The triggered RPA robots launched in this way may be able to perform actions such as order status checks and data posting to CRM using plain - language commands.
[0069] End-to-end measurement of automation programs at any scale and governance can be provided by the hyper-automation system 100 in some embodiments. As such, analytics (e.g., via UiPath Insights™) may be employed to understand the performance of the automation. Data modeling and analytics using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for automation, perform remote monitoring on the user's computing system, detect errors and anomalies, and debug the automation. An automation management console (e.g., UiPath Automation Ops™) may be provided to manage the automation throughout its lifecycle. The organization may govern how the automation is built, what users can do with them, and which automation users can access.
[0070] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes can be discovered, automation can be built, tested, and deployed, performance can be measured, use of the automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained, retrained, and the process itself can be repeated. This promotes a more robust and effective set of automation.
[0071] In some embodiments, a generative AI model is used. Generative AI can generate various types of content, such as text, images, audio, and synthetic data. Various types of generative AI models can be used, including but not limited to large language models (LLMs), adversarial generative networks (GANs), variational autoencoders (VAEs), transformers, etc. These models may be part of the AI / ML model 132 hosted on server 130. For example, the generative AI model can be trained on a large corpus of text information to perform semantic understanding, understand the nature of what exists on the screen from text, automatically generate code, etc. In certain embodiments, a generative AI model 172 provided by an existing cloud ML service provider such as OpenAI®, Google®, Amazon®, Microsoft®, IBM®, Nvidia®, Facebook® can be employed and trained to provide such functionality. In a generative AI embodiment where the generative AI model(s) 172 is remotely hosted, server 130 can be configured to integrate with a third-party API, whereby server 130 can send requests containing the necessary input information to the generative AI model(s) 172 and receive its responses (e.g., semantic matching of fields between versions of an application, classification of the type of application on the screen, etc.). Such embodiments can not only provide a more advanced and refined user experience but also provide access to state-of-the-art NLP and other ML capabilities offered by these companies.
[0072] One aspect of the generative AI model in some embodiments is the use of transfer learning. In transfer learning, a pre-trained generative AI model such as an LLM is fine-tuned for a specific task or domain. This allows the LLM to leverage the knowledge it has already learned during its initial training and adapt it to a specific application. In the case of an LLM, during the pre-training phase, the LLM is typically trained on a large text corpus consisting of billions of words. During this phase, the LLM learns the relationships between words and phrases, enabling it to generate consistent human-like responses to text-based inputs. The output of this pre-training phase is an LLM that highly understands the patterns underlying natural language.
[0073] In the fine-tuning phase, the pre-trained LLM is adapted to a specific task or domain by training the LLM on a smaller, task-specific dataset. For example, in some embodiments, the LLM can be trained to analyze a specific type or types of data sources to improve its accuracy regarding their content. Such information can be provided as part of the training data, and the LLM can learn to focus on these areas and more accurately identify data elements within them. Fine-tuning enables the LLM to learn the subtleties of the task or domain, such as the specific vocabulary and syntax used in that domain, without requiring as much data as would be needed to train the LLM from scratch. By leveraging the knowledge learned during the pre-training phase, the fine-tuned LLM can achieve state-of-the-art performance on a specific task with a relatively small amount of training data.
[0074] LLMs can be trained using vector databases. Vector databases index, store, and provide access to structured or unstructured data (e.g., text, images, time series data, etc.) along with their vector embeddings. Data such as text can be tokenized, where single characters, words, or sequences of words are parsed from the text into tokens. These tokens are then "embedded" into vector embeddings, which are numerical representations of this data. Vector databases allow users to search and retrieve similar objects quickly and at scale in production environments.
[0075] AI and ML allow unstructured data to be represented numerically without losing its semantic meaning in vector embeddings. A vector embedding is a long list of numbers, where each number represents a feature of the data object that the vector embedding represents. Similar objects are grouped together in the vector space. In other words, the more similar the objects are, the closer the vector embeddings that represent them are to each other. Similar objects may be found using vector search, similarity search, or semantic search. The distance between vector embeddings may be calculated using various techniques, including but not limited to Euclidean squared or L2 squared distance, Manhattan or L1 distance, cosine similarity, dot product, Hamming distance, etc. It may be beneficial to choose the same metric used to train the AI / ML model.
[0076] Vector indexing can be used to organize vector embeddings so that data can be retrieved efficiently. When the number of data points is large, using the k-nearest neighbor (kNN) algorithm to calculate the distance between a vector embedding in a vector database and all other vector embeddings can be computationally expensive because the necessary calculations increase linearly (O(n)) depending on the dimension and the number of data points. It is more efficient to use an approximate nearest neighbor (ANN) approach to find similar objects. Since the distances between vector embeddings are pre-computed, similar vectors are organized and stored close to each other (e.g., within a cluster or graph), similar objects can be found more quickly. This process is called "vector indexing". ANN algorithms that can be used in some embodiments include, but are not limited to, clustering-based indexing, proximity graph-based indexing, tree-based indexing, hash-based indexing, compression-based indexing, etc.
[0077] FIG. 2 is an architectural diagram showing an RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 may facilitate the development of an automation project that is a graphical representation of a business process. Briefly, the designer 210 facilitates the development and deployment of workflows and robots. In some embodiments, the designer 210 may be an application running on a user's desktop, an application running remotely on a VM, a web application, or the like.
[0078] Automation projects enable the automation of rule-based processes by giving developers control over the execution order and relationships between steps in a custom set of workflows defined as "activities" in this specification as described above. A commercial example of an embodiment of Designer 210 is UiPath Studio (trademark). Each activity may include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.
[0079] Some types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers, etc. Sequences may be particularly suitable for linear processes that enable the flow from one activity to another without cluttering the workflow. Flowcharts may be particularly suitable for more complex business logic, enabling the integration of decision-making and the connection of activities in more diverse ways through multiple branching logic operators. FSMs may be particularly suitable for large-scale workflows. FSMs may use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. Global exception handlers may be particularly suitable for determining the behavior of the workflow when encountering execution errors or for debugging the process.
[0080] Once a workflow is developed within the designer 210, the execution of the business process is coordinated by the conductor 220, which coordinates one or more robots 230 that execute the workflow developed within the designer 210. A commercial example of an embodiment of the conductor 220 is UiPath Orchestrator (trademark). The conductor 220 facilitates the management of the generation, monitoring, and deployment of resources in the environment. The conductor 220 can operate as an integration point with third-party solutions and applications. As described above, in some embodiments, the conductor 220 can be part of the core hyper-automation system 120 of FIG. 1.
[0081] The conductor 220 can manage all robots 230 and connect and execute the robots 230 from a central point. The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non-production robots (similar to attended robots 232 but used for development and testing purposes). Attended robots 232 are triggered by user events and operate in parallel with humans on the same computing system. Attended robots 232 can be used with the conductor 220 for centralized process deployment and logging medium. Attended robots 232 may assist human users in achieving various tasks and may be triggered by user events. In some embodiments, a process cannot start from the conductor 220 on this type of robot and / or they cannot execute under a locked screen. In certain embodiments, attended robots 232 can only be launched from a robot tray or a command prompt. Attended robots 232 preferably operate under human supervision in some embodiments.
[0082] The unattended robot 234 can operate unmanned in a virtual environment and automate many processes. The unattended robot 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can be performed by the designer 210 in some embodiments. Both attended and unattended robots can automate various systems and applications including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0083] The conductor 220 can have various capabilities including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning can include creating and maintaining a connection between the robot 230 and the conductor 220 (e.g., a web application). Deployment can include ensuring the correct delivery of package versions to the robots 230 assigned for execution. Configuration can include maintaining and delivering the robot environment and process configurations. Queuing can include providing management of queues and queue items. Monitoring can include tracking specific data of the robots and maintaining user permissions. Logging can include saving and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a “not only” SQL (NoSQL) database) and / or another storage mechanism (e.g., ElasticSearch (registered trademark) that provides the ability to store large datasets and execute queries quickly). The conductor 220 can provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0084] The robot 230 is an execution agent that implements the workflows built by the designer 210. One commercial example of some embodiments of the robot(s) 230 is UiPath Robots (trademark). In some embodiments, the robot 230, by default, installs the Microsoft Windows (registered trademark) Service Control Manager (SCM) management service. As a result, such a robot 230 can open an interactive Windows (registered trademark) session under a local system account and can have the rights of a Windows (registered trademark) service.
[0085] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, it means having the same rights as the user in which a given robot 230 is installed. This feature may also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0086] The robot 230 in some embodiments is divided into multiple components, each specialized for a specific automation task. The robot components in some embodiments include, but are not limited to, SCM management robot service, user mode robot service, executor, agent, and command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 is executed). These services are entrusted with managing the qualification information of the robot 230. The console application is launched by the SCM under the local system.
[0087] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host. The user mode robot service may be entrusted with managing the qualification information of the robot 230. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0088] The executor can execute a job given under a Windows® session (i.e., can execute a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows® Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0089] As described above, the fact that the components of the robot 230 are split helps the developer, support user, and computing system to more easily perform, identify, and track what each component is doing. In this way, special behaviors can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 can be made independent of the browser zoom level. In the case of an application marked as not recognizing or intentionally not recognizing DPI, DPI can be disabled in some embodiments.
[0090] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for the execution of the AI / ML model(s) and receive outputs therefrom via the core hyper-automation system 240.
[0091] One or more robots 230 may be listeners, as described above. These listeners can provide information to the core hyper-automation system 240 regarding what users are doing when they use their computing systems. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0092] The assistant / chatbot 250 can be provided on the user computing system to enable the user to launch an RPA local robot. The assistant can be placed, for example, in the system tray. The chatbot can have a user interface so that the user can view the text of the chatbot. Alternatively, the chatbot can run in the background without a user interface and can listen for the user's utterances using the microphone of the computing system.
[0093] In some embodiments, data labeling may be performed by a user of the computing system that the robot is executing on, or on another computing system that the robot provides information to. For example, if the robot calls an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user may draw a rectangle around the mis-identified or non-identified component and potentially provide text with the correct identification. This information can be provided to the core hyper-automation system 240 and can then be used later for training a new version of the AI / ML model.
[0094] FIG. 3 is an architecture diagram showing an expanded RPA system 300 according to an embodiment of the present invention. In some embodiments, the RPA system 300 may be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The expanded RPA system 300 can be a cloud-based system, an on-premises system, or a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes.
[0095] It should be noted that the client side, the server side, or both can include any desired number of computing systems without departing from the scope of the present invention. On the client side, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is executing a process. As shown in FIG. 3, a plurality of business projects can be executed simultaneously. The agent 314 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via the database server 350, the AI / ML server 360, the indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 can be a robot component.
[0096] In some embodiments, the robot represents an association between a machine name and a user name. The robot can manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots can be run simultaneously, each running in a separate Windows® session using a unique user name. This is referred to as the HD robot described above.
[0097] Agent 314 is also responsible for transmitting the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. In some embodiments, the communication between Agent 314 and Conducter 340 is always initiated by Agent 314. In a notification scenario, Agent 314 may open a WebSocket channel that is later used by Conducter 340 to send commands (e.g., start, stop, etc.) to the robot.
[0098] Listener 330 monitors and records data related to user interactions with the operation of the attended computing system and / or unattended computing system in which listener 330 resides. Listener 330 can be, without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the listener is implemented partially or fully via physical hardware.
[0099] On the server side, there are a presentation layer (web application 342, Open Data Protocol (OData) Representational State Transfer (REST) Application Programming Interface (API) endpoint 344, notification and monitoring 346), a service layer (API implementation / business logic 348), and a persistence layer (database server 350, AI / ML server 360, indexer server 370). The conductor 340 includes the web application 342, the Odata REST API endpoint 344, notification and monitoring 346, and the API implementation / business logic 348. In some embodiments, most actions performed by the user at the interface of the conductor 340 (e.g., via the browser 320) are executed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of the present invention. The web application 342 is the visual layer of the server platform. In this embodiment, the web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, script language, or any other format may be used without departing from the scope of the present invention. The user interacts with the web page from the web application 342 via the browser 320 in this embodiment to perform various operations for controlling the conductor 340. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or each process, start and stop robots, etc.
[0100] In addition to the web application 342, the conductor 340 also includes a service layer that exposes an Odata REST API endpoint 344. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 342 and the agent 314. The agent 314 is, in this embodiment, a supervisor of one or more robots on a client computer.
[0101] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functions. Configuration endpoints may be used in some embodiments to define and configure the users, permissions, robots, assets, releases, and environments of the application. The logging REST endpoint may be used to log various information, such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by robots to query the version of the package to be executed when a job start command is used in the conductor 340. The queuing REST endpoint may be responsible for the management of queues and queue items, such as adding data to the queue, retrieving transactions from the queue, and setting the status of transactions.
[0102] Monitoring of the REST endpoint may monitor the web application 342 and the agent 314. The notification and monitoring API 346 may be a REST endpoint used for the registration of the agent 314, the distribution of configuration settings to the agent 314, and the sending and receiving of notifications from the server and the agent 314. The notification and monitoring API 346 may use WebSocket communication in some embodiments.
[0103] In some embodiments, the API of the service layer can be accessed through the configuration of an appropriate API access path, for example, based on whether the conductor 340 and the overall hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The API for the conductor 340 can provide custom methods for querying statistics regarding various entities registered with the conductor 340. In some embodiments, each logical resource may be an Odata entity. In such an entity, components such as robots, processes, queues, etc. may have properties, relationships, and operations. In some embodiments, the API of the conductor 340 can be consumed by the web application 342 and / or the agent 314 in the following two ways: by obtaining API access information from the conductor 340 or by registering an external application for using the Oauth flow.
[0104] In this embodiment, the persistent layer includes three servers: a database server 350 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides AI / ML model - providing services such as an AI center function), and an indexer server 370. The database server 350 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 342. The database server 350 may manage queues and queue items. In some embodiments, the database server 350 may store (in addition to, or instead of, the indexer server 370) messages recorded by robots. The database server 350 may also store, for example, process mining, task mining, and / or task capture - related data received from a listener 330 installed on the client side. Although no arrow is shown between the listener 330 and the database 350, it should be understood that in some embodiments, the listener 330 can communicate with the database 350 and vice versa. This data can be stored in forms such as PDD, images, XAML files, etc. The listener 330 may be configured to eavesdrop on user actions, processes, tasks, and performance metrics on each computing system on which the listener 330 resides. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, time, etc.) on its respective computing system and then convert them into a form suitable for being provided to and stored in the database server 350.
[0105] The AI / ML server 360 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can enable non-data scientists to access such capabilities. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of the AI / ML models can be monitored and trained and improved using human-verified data. The AI / ML server 360 can schedule and execute training jobs to train new versions of the AI / ML models.
[0106] The AI / ML server 360 can store data related to AI / ML models and ML packages for configuring various ML skills for users during development. The ML skills used herein are, for example, pre-built and trained ML models for processes that can be used by automation. The AI / ML server 360 can also store data related to document understanding technologies, frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, NLP, speech analysis, different types of AI / ML models, etc.
[0107] Optionally in some embodiments, the indexer server 370 stores information recorded by robots and creates an index. In certain embodiments, the indexer server 370 may be disabled via configuration settings. In some embodiments, the indexer server 370 uses ElasticSearch®, an open-source project's full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 370 via logging REST endpoint(s), where they are indexed for future use.
[0108] FIG. 4 is an architecture diagram illustrating a relationship 400 between a designer 410, activities 420, 430, 440, 450, a driver 460, an API 470, and an AI / ML model 480 according to an embodiment of the present invention. As described above, a developer uses the designer 410 to develop a workflow to be performed by a robot. Various types of activities may be presented to the developer in some embodiments. The designer 410 may be local or remote to the user's computing system (e.g., accessed via a local web browser that interacts with a VM or a remote web server). The workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. The user-defined activities 420 and API-driven activities 440 interact with applications via their APIs. The user-defined activities 420 and / or AI / ML activities 440 may call one or more AI / ML models 480 that may be located locally and / or remotely to the computing system on which the robot operates in some embodiments.
[0109] In some embodiments, non-text visual components in an image can be identified, which is referred to herein as CV. However, it should be noted that in some embodiments, CV incorporates OCR. CV can be at least partially executed by an AI / ML model(s) 480. Some CV activities related to such components can include, but are not limited to, extraction of text from segmented label data using OCR, fuzzy text matching, cropping of segmented label data using ML, comparison of the extracted text in the label data with ground truth data, etc. In some embodiments, the number of activities that can be implemented in user-defined activity 420 can be hundreds or thousands. However, any number and / or type of activities can be used without departing from the scope of the present invention.
[0110] UI automation activity 450 is a subset of special low-level activities written in low-level code that facilitate interaction with the screen. UI automation activity 450 facilitates these interactions via a driver 460 that enables the robot to interact with the desired software. For example, driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, etc. In some embodiments, one or more AI / ML models 480 can be used by UI automation activity 450 to perform interactions with the computing system. In certain embodiments, the AI / ML model 480 can enhance or completely replace the drivers 460. In fact, in certain embodiments, the drivers 460 are not included.
[0111] Driver 460 can interact with the OS at a low level, such as searching for hooks and monitoring keys, via the OS driver 462. Driver 460 may facilitate integration with Chrome (registered trademark), IE (registered trademark), Citrix (registered trademark), SAP (registered trademark), etc. For example, a "click" activity plays the same role in these different applications via driver 460.
[0112] Figure 5 is an architectural diagram showing a computing system 500 configured to perform automatic code generation for RPA according to an embodiment of the present invention. In some embodiments, the computing system 500 may be one or more computing systems depicted and / or described herein. In certain embodiments, the computing system 500 may be part of a hyper-automation system as shown in FIGS. 1 and 2. The computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to the bus 505 for processing information. The processor(s) 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The processor(s) 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one processor(s) 510 can be a neuromorphic circuit that includes processing elements mimicking biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0113] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by one or more processors 510. The memory 515 can be composed of a random access memory (RAM), a read-only memory (ROM), a flash memory, a cache, a static storage device such as a magnetic disk or an optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media can be any available media accessible by one or more processors 510, and can include volatile media, non-volatile media, or both. Also, the media can be removable, non-removable, or both. The computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via a wireless and / or wired connection. In some embodiments, the communication device 520 can include one or more antennas that are a single antenna, an array of antennas, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration without departing from the scope of the present invention.
[0114] The processor(s) 510 is further coupled to a display 525 via a bus 505. Without departing from the scope of the present invention, any suitable display device and tactile I / O may be used. A keyboard 530 and a cursor control device 535, such as a computer mouse, touch pad, etc., are further coupled to the bus 505 to enable a user to interface with the computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user may interact with the device only via the display 525 and / or a touch pad (not shown). Any type and combination of input devices may be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, a user may interact with the computing system 500 remotely via another computing system that is communicating with the computing system 500, or the computing system 500 may operate autonomously.
[0115] Memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include an automatic code generation module 545 configured to execute all or a portion of the AI / ML processes described herein or derivatives thereof. The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0116] One of ordinary skill in the art will understand that the "system" can be embodied, without departing from the scope of the present invention, as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. In fact, the methods, systems, and apparatuses disclosed herein may be implemented in a localized and distributed form consistent with computing techniques including cloud computing systems. The computing system may be part of, or accessible by, a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public cloud or private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of the present invention.
[0117] It should be noted that some of the system features described herein are presented as modules to emphasize implementation independence more. For example, a module may be implemented as a hardware circuit including custom very large scale integration (VLSI) circuits or gate arrays, logic chips, transistors, or other off-the-shelf semiconductors such as individual components. Also, a module may be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, and the like.
[0118] The module can also be at least partially implemented in software for execution by various types of processors. For example, a specified unit of executable code can include one or more physical or logical blocks of computer instructions that may be organized, for example, as objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and can include modules when logically combined and can include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module can be stored in a non-transitory computer-readable medium such as, for example, a hard disk drive, flash device, RAM, tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of the present invention.
[0119] In fact, a module of executable code can be a single instruction, a number of instructions, or even be distributed across multiple different code segments, between different programs, and across multiple memory devices. Similarly, the operational data can be specified within the module, shown here, embodied in any suitable form, and organized within any suitable type of data structure. The operational data can be collected as a single data set or be distributed in different locations across different storage devices and can exist, at least in part, simply as electronic signals on a system or network.
[0120] Without departing from the scope of the present invention, various types of AI / ML models can be trained and deployed. For example, FIG. 6A shows an example of a neural network 600 trained to supplement the automatic code generation of RPA according to an embodiment of the present invention. The neural network 600 includes a number of hidden layers. Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) typically have multiple layers, but an SLNN may sometimes have only one or two layers and is usually less than a DLNN. Typically, the architecture of a neural network includes an input layer, multiple intermediate layers, and an output layer, as in the case of the neural network 600.
[0121] Often, DLNNs have many layers (such as 10, 50, 200), and subsequent layers typically reuse the features from the previous layer to compute more complex and general functions. On the other hand, SLNNs have few layers and tend to be trained relatively quickly because expert features are pre-created from raw data samples. However, feature extraction is cumbersome. On the other hand, DLNNs usually do not require expert features but take time to train and tend to have more layers.
[0122] In either approach, the layers are trained simultaneously on a training set and overfitting is usually checked on a separate cross-validation set. Excellent results are obtained with both techniques, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem addressed by each neural network.
[0123] Returning to FIG. 6A, the RPA workflow, source files (such as PDF documents, PDDs, policies, XAML files, etc.), UI drawings and screenshots, task mining information, etc. are provided as an input layer and fed as input to the J neurons of hidden layer 1. Various other inputs are possible, including but not limited to the state information of the computing system, published automation, business rules, information related to what the RPA workflow and / or tasks are related to, the initial definition of the automation, process automation documents, etc. In this example, all of these inputs are fed to each neuron, but not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, noise removal autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks that do not depart from the scope of the present invention. Various architectures can be used, either individually or in combination.
[0124] Hidden layer 2 receives input from hidden layer 1, hidden layer 3 receives input from hidden layer 2, and the same is done for all hidden layers until the last hidden layer provides its output as the input to the output layer. Although multiple proposals are shown as outputs in this specification, in some embodiments, only a single output proposal is provided. In certain embodiments, the proposals are ranked based on a confidence score.
[0125] It should be noted that the numbers of neurons I, J, K, and L are not necessarily equal. Therefore, any desired number of layers can be used in a given layer of the neural network 600 without departing from the scope of the present invention. In fact, in certain embodiments, the types of neurons in a given layer may not all be the same.
[0126] The neural network 600 is trained to assign a confidence score(s) to the appropriate output. To reduce inaccurate predictions, in some embodiments, only those results with a confidence score above a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs with a confidence score exceeding this amount may be used and the rest may be ignored.
[0127] A neural network is typically a probabilistic construct that has a confidence score(s). This can be a score learned by the AI / ML model based on the frequency with which similar inputs were correctly identified during training. Some common types of confidence scores include decimal numbers between 0 and 1 (which can also be interpreted as a confidence percentage), numerical values between negative infinity and positive infinity, a series of expressions (e.g., "low", "medium", and "high"), etc. Various post-processing calibration techniques such as temperature scaling, batch normalization, weight decay, negative log-likelihood (NLL), etc. can also be used to obtain a more accurate confidence score.
[0128] The "neurons" of a neural network are usually implemented algorithmically as mathematical functions based on the functions of biological neurons. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass the output to the next layer. This activation function can be a non-linear thresholded activity function that does nothing if the value is below the threshold, and responds linearly when the function exceeds the threshold (i.e., the rectified linear unit (ReLU) non-linearity). Since actual neurons can have a nearly identical activity function, the sum function and the ReLU function are used in deep learning. Through linear transformation, information can be subtracted, added, etc. Essentially, neurons function as gating functions that pass the output to the next layer governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some of the neurons.
[0129] JPEG2025097252000002.jpg84144
[0130] JPEG2025097252000003.jpg46144
[0131] JPEG2025097252000004.jpg29132
[0132] In this case, neuron 610 is a single-layer perceptron. However, without departing from the scope of the present invention, any suitable neuron type or combination of neuron types can be used. It should also be noted that the weights of the activation function and / or the range of values of the output value(s) can be different in some embodiments without departing from the scope of the present invention.
[0133] A target, i.e., a "reward function", is often adopted. The reward function guides the search in the state space and tries to achieve the target (e.g., finding the most accurate answer to the user's query based on relevant indicators) by using both short-term and long-term rewards to explore intermediate transitions and steps. During training, various labeled data are supplied through the neural network 600. When a particular one succeeds, the weights of the inputs to the neurons are strengthened, while when a particular one fails, those weights are weakened. A cost function such as the mean squared error (MSE) or gradient descent can be used to make slightly incorrect predictions cost much less than greatly incorrect predictions. If the performance of the AI / ML model is not improved after a certain number of training iterations, the data scientist can change the reward function or correct incorrect predictions, etc.
[0134] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "pop up" the hidden layers of the neural network to check how much loss each node is bearing, and then assign low weights to the nodes with high error rates and vice versa to update the weights to minimize the loss. That is, backpropagation enables the data scientist to repeatedly adjust the weights so as to minimize the difference between the actual output and the desired output.
[0135] The algorithm of backpropagation is mathematically based on the optimization theory. In supervised learning, training data with known outputs are passed through the neural network, the error is calculated using the cost function from the known target outputs, and this gives the error for backpropagation. The error is calculated at the output, and this error is converted into a correction of the network's weights to minimize the error.
[0136] JPEG2025097252000005.jpg71144
[0137] JPEG2025097252000006.jpg22144
[0138] JPEG2025097252000007.jpg123144
[0139] JPEG2025097252000008.jpg106116
[0140] JPEG2025097252000009.jpg83144
[0141] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., above 97% using an F2 or F4 threshold for detection, about 2000 epochs). This accuracy level can be determined in some embodiments using an F1 score, an F2 score, an F4 score, or any other suitable technique that does not depart from the scope of the present invention. Once trained on the training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps prevent the AI / ML model from becoming "overfitted" such that it performs well on the training data but not on other data.
[0142] In some embodiments, it may not be known what accuracy levels an AI / ML model can achieve. Thus, when the accuracy of an AI / ML model begins to decline when analyzing evaluation data (i.e., the model performs well on training data but its performance begins to degrade on evaluation data), the AI / ML model can undergo additional training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when the accuracy reaches a certain level or when the accuracy of the trained AI / ML model is better than that of an existing deployed AI / ML model. In certain embodiments, a set of trained AI / ML models can be used to accomplish a task. For example, one model can be trained to recognize images, another model can be trained to recognize text, and yet another model can be trained to recognize semantic and / or ontology-related aspects, etc.
[0143] In some embodiments, transformer neural networks such as SentenceTransformers™, a Python™ framework for state-of-the-art sentence, text, and image embedding, can be used. Such transformer neural networks learn associations of words and phrases with both high and low scores. This trains the AI / ML model to determine what is close to the input and what is not, respectively. Instead of using only word / phrase pairs, the transformer neural network may also use field length and field type.
[0144] In some embodiments, NLP technologies such as word2vec, BERT, GPT-3, ChatGPT, and other LLMs can be used to facilitate semantic understanding and provide more accurate and human-like answers as described above. Other technologies such as clustering algorithms can be used to find similarities between groups of elements. Clustering algorithms can include, but are not limited to, density-based algorithms, distribution-based algorithms, centroid-based algorithms, and hierarchical-based algorithms. Such as the K-means clustering algorithm, DBSCAN clustering algorithm, Gaussian mixture model (GMM) algorithm, and balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm. Such techniques can also be useful for classification.
[0145] FIG. 7 is a flowchart showing a process 700 for training an AI / ML model(s) according to an embodiment of the present invention. In some embodiments, the AI / ML model(s) may be a generative AI model as described above. The neural network architecture of the AI / ML model typically includes multiple layers of neurons including an input layer, an output layer, and hidden layers. For example, refer to FIGS. 6A and 6B. The intermediate hidden layers process the input data and generate an intermediate representation of the input that is used for generating the output. These hidden layers can include various types of neurons such as convolutional neurons, recurrent neurons, and / or transformer neurons.
[0146] The training process begins at 710 by providing, regardless of the presence or absence of labels, RPA workflows, source files, UI images and screenshots, task mining information, etc. The AI / ML model is then trained at 720 over multiple epochs, and the results are reviewed at 730. Various types of AI / ML models can be used, but LLM and other generative AI models are typically trained using a process called "supervised learning" as described above. Supervised learning involves providing the model with a large dataset, which the model uses to learn the relationship between inputs and outputs. During the training process, the model adjusts the weights and biases of the neurons within the neural network to minimize the difference between the predicted output and the actual output within the training dataset.
[0147] One aspect of the model in some embodiments is the use of transfer learning. For example, transfer learning can utilize a pre-trained model such as ChatGPT that is fine-tuned for a specific task or domain at step 720. This allows the model to leverage the knowledge already learned from the pre-training phase and adapt it to a specific application through the training phase at step 720.
[0148] The pre-training phase involves training the model based on an initial set of training data that may be more general. In this phase, the model learns the relationships within the data. In the fine-tuning phase (e.g., in some embodiments, when a pre-trained model is used as the initial basis for the final model, in addition to or instead of the initial training phase, and executed during step 720), the pre-trained model is adapted to a specific task or domain by training the model with a smaller dataset specific to the task. For example, in some embodiments, the model may focus on a specific type(s) of data source. This can help the model more accurately identify data elements within it than a generative AI model pre-trained alone. Through fine-tuning, the model can learn nuances of the source such as specific vocabulary and syntax, specific graphical characteristics, specific data formats, etc., without requiring as much data as would be needed to train the model from scratch. By leveraging the knowledge learned in the pre-training phase, the fine-tuned model can achieve state-of-the-art performance on a specific task even with relatively little additional training data.
[0149] If the AI / ML model does not meet the desired confidence threshold at 740, the training data is supplemented and / or the reward function is modified at 750 to help the AI / ML model better achieve its purpose, and the process returns to step 720. If the AI / ML model meets the confidence threshold at 740, the AI / ML model is tested against the evaluation data at 760 to confirm that the AI / ML model generalizes well and does not overfit to the training data. The evaluation data includes information that the AI / ML model has not processed before. If the confidence threshold is met for the evaluation data at 770, the AI / ML model is deployed at 780. Otherwise, the process returns to step 750 and the AI / ML model is further trained.
[0150] Figures 8A - C are screenshots showing an automated code generation interface 800 configured to analyze natural language input from a user and propose automation. The automated code generation interface includes a text box 810 where the user can enter text regarding the task to be automated. The automated code generation interface 800 also includes a cancel button 820 and a confirm button 822, the latter being inactive in FIGS. 8A and 8B.
[0151] In FIG. 8B, the user enters text into the text box 810 to automate the order review process by a financial manager. Specifically, the user types "Orders over $10,000 are reviewed by the finance manager. If approved, the order is updated in Salesforce. If rejected, the order is cancelled." When the user enters text in the text box 810 and presses enter, the automation generation status indicator 830 indicates that the proposed automation is being generated based on the user input.
[0152] Referring to FIG. 8C, the automated code generation interface 800 proposes an automation for the requested task. A rule 840 is created for the condition that the order exceeds $10,000.00. When this condition is met, the financial manager review user task 850 is executed. The assigned reviewer(s) can execute review options 860. That is, the reviewer can approve the order or cancel the order. If the proposed automation is correct, the user can use the confirm button 822 to confirm this, and then the automation is generated (e.g., executed by an RPA robot(s)). Once the process is generated, in some embodiments, the user can use an RPA designer application to drill down into the RPA workflow and make any desired edits to the workflow.
[0153] Figure 9 shows the AI / ML model 900 of the cognitive AI layer according to an embodiment of the present invention. In some embodiments, the model of the cognitive AI layer 900 can be trained using the process 700 of FIG. 7. The generative AI model 910 provides results to other cognitive AI layer models in a serial configuration 940, a parallel configuration 942, or a combination 944 of serial and parallel configurations (s). In some embodiments, the cognitive layer can be a single generative AI model. For example, the generative AI model can perform processes such as generating code, providing semantic associations between texts on a screen, proposing formulas and code snippets for RPA workflow activities, generating applications based on UI images, and providing explanations based on records of user actions. The CV model 920 and the OCR model also provide the detected graphical elements and the recognized text to the generative AI model 910 and other cognitive AI layer models within the configurations 940, 942, or 944, respectively.
[0154] Other cognitive AI models of the configurations 940, 942, or 944 use the outputs from the generative AI model 910, the CV model 920, and / or the OCR model 930 to provide the automatic code generation functionality described in detail herein. The cognitive AI layer 900 can facilitate understanding of what code needs to be added to an RPA workflow, what code needs to be converted to another language within the workflow, which RPA workflows to propose based on user actions, what kind of explanations need to be generated based on records of tasks executed by users on a computing system, which applications to generate based on these actions, UI images, combinations thereof, etc. The output (i.e., result 950) from other cognitive AI layer models of the configurations 940, 942, or 944 can include RPA workflows that contain formulas and / or code in the relevant programming language, newly generated RPA workflows, human-like descriptions of tasks that users are executing on a computing system, applications, etc.
[0155] Figure 10 is a flowchart 1000 showing the process of automatic code generation for RPA according to an embodiment of the present invention. In some embodiments, the process begins at 1020 by receiving source information from a user. However, in other embodiments, this may be provided automatically to the system. The source data is provided to the cognitive AI layer at 1020. The cognitive AI layer is one or more AI / ML models. In some embodiments, the cognitive AI layer is part of a computer program that executes an automatic code generation process, such as an RPA designer application, an RPA robot, etc. The source data may include, but is not limited to, an RPA workflow, a natural language sentence, a recording of a user's voice, a video recording of a user's actions on a computing system, a source document (e.g., a spreadsheet file, a JSON file, a XAML file, an XML file, an HTML file, a Word® document, a PDF file, a scanned image of a physical document, etc.), pseudocode for a desired task, a diagram of a desired user interface or form, a diagram of a process, any combination thereof, etc.
[0156] Next, the source data is processed by the cognitive AI layer at 1030 (i.e., the cognitive AI layer executes the input source information through its AI / ML model(s)), and the cognitive AI layer provides an output at 1040. For example, the output of the cognitive AI layer may be provided to an RPA designer application, an RPA robot, or some other process that can implement the automatically generated code in a useful way. The output from the cognitive AI layer may include, but is not limited to, an existing RPA workflow that includes expressions and / or code snippets (plural) in a relevant programming language(s), a newly generated RPA workflow, a human-like explanation of the task that the user is performing on a computing system, an application, expressions (plural) and / or code snippets (plural), etc.
[0157] When the user receives the automatically generated code from the cognitive AI layer at 1050, the automatically generated code is implemented at 1060. For example, the RPA workflow is updated in the RPA designer application by adding expressions and / or code snippets to its activities and / or configurations, and the RPA workflow generated by the cognitive AI layer is compiled into an automation to be executed by the RPA robot at runtime, the description of the task to be executed by the computer is saved in a database and provided to other users, and the application generated by the cognitive AI layer is deployed to the user's computing system, etc. In some embodiments, this may be done automatically without manual user approval. In certain embodiments, implementing the automatically generated code at 1050 includes adding security and / or compliance rules to the RPA workflow to comply with laws and / or policies. In some embodiments, implementing the automatically generated code includes compiling the RPA workflow into an automation configured to be performed by the RPA robot at runtime.
[0158] If the user does not accept the cognitive AI layer output at 1050 and can modify the automatically generated code at 1070 due to user changes (for example, the user can change the representation of the automatically generated document or correct the grammar), the user can make the changes at 1080 and the automatically generated code can be implemented at 1060. However, this may not be possible in some cases. For example, the user may not know how to modify the code of the automatically generated application to correct the errors therein, the output from the cognitive AI layer may be incomprehensible to the user, or the proposed RPA workflow / automation may be incorrect. In this case, the user may fail in the automatic code generation process, and data related to the failure may be sent at 1090 to retrain the AI / ML model(s) of the cognitive AI layer. For example, the user may provide a textual explanation of which elements of the automatically generated code were incorrect and why.
[0159] Figure 11 is a flowchart showing a process 1100 for automatic document generation using a cognitive AI layer according to an embodiment of the present invention. The process begins at 1110 by receiving task mining output regarding a user's actions. The output of task mining can include, but is not limited to, the applications the user is using, the information input into these applications, key presses, mouse clicks and positions, graphical elements within the UI, and which graphical elements are active elements at various times (e.g., the current element the user is inputting into, cursor position, etc.).
[0160] The task mining output is provided as input to the cognitive AI layer and processed at 1120. The cognitive AI layer includes an LLM configured to analyze the task mining input and generate text (and in some embodiments, images) based on the task mining output. Thus, the cognitive AI layer determines what the user was doing based on the task mining output and documents the process at 1130. For example, using the LLM of the cognitive AI layer, a human - like description of the actions the user is performing, which may in some cases include a screenshot, can be generated. This information can then be provided as a manual to other users who want to perform each task at 1140.
[0161] Figure 12 is a flowchart showing a process 1200 for generating automation from natural language text according to an embodiment of the present invention. The process begins at 1210 by receiving natural language text or an audio recording from the user. In the latter case, the audio is converted to text using an audio - to - text conversion model. In some embodiments, the natural language text can be input by the user in an application equipped with an automatic code generation interface. See, for example, FIGS. 8A - 8C.
[0162] The cognitive AI layer with an LLM receives natural language text as input and processes the text at 1220. The application with an automatic code generation interface receives the output from the cognitive AI layer containing information about the proposed automation at 1230 and displays the proposed automation steps at 1240. For example, rules, tasks, steps, etc. related to the proposed automation may be displayed.
[0163] When the user accepts the automation at 1250, the automation is generated at 1260. The automation can be an automation performed by a stand-alone application or an RPA robot. If the proposed automation at 1250 is not accepted by the user, the automatic code generation interface requests user input on what was wrong with the proposed automation and sends information for retraining the LLM at 1270.
[0164] The process steps executed in FIGS. 7 and 10 - 12 may be executed by a computer program that encodes instructions to a processor(s) to execute at least a part of the process(es) described in FIGS. 7 and 10 - 12 according to an embodiment of the present invention. The computer program may be stored in a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, a tape, and / or any other such medium or combination of media used to store data. The computer program may include encoded instructions for controlling a processor(s) of a computing system (e.g., the processor(s) 510 of the computing system 500 in FIG. 5) to implement all or part of the process steps described in FIGS. 7 and 10 - 12, which may also be stored in a computer-readable medium.
[0165] A computer program can be implemented in hardware, software, or a hybrid implementation. The computer program can be composed of modules that communicate operably with each other and is designed to send information or instructions to a display. The computer program can be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0166] It will be readily understood that the components of the various embodiments of the present invention can be arranged and designed in a variety of different configurations as generally described and illustrated herein. Accordingly, the detailed description of the embodiments of the present invention as represented in the accompanying figures is not intended to limit the scope of the invention as claimed, but merely represents selected embodiments of the present invention.
[0167] The features, structures, or characteristics of the present invention described throughout this specification can be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments", "some embodiments", or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. Accordingly, the appearances throughout this specification of "in certain embodiments", "in some embodiments", "in other embodiments", or similar language are not necessarily all referring to the same group of embodiments, and the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0168] References throughout this specification to features, advantages, or similar language do not imply that all of the features and advantages realizable with the present invention should be in any single embodiment of the invention or in any one embodiment of the invention. Rather, the language referring to the features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the present invention. Thus, discussions of the features and advantages throughout this specification, as well as similar language, may refer to the same embodiment, but do not necessarily have to.
[0169] Furthermore, the described features, advantages, and characteristics of the present invention may be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the present invention can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments but not in all embodiments of the present invention.
[0170] Those of ordinary skill in the art will readily understand that the present invention, as described above, can be practiced using steps in a different order and / or using hardware elements in a configuration different from that disclosed. Accordingly, while the present invention has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations will become apparent while remaining within the spirit and scope of the invention. Therefore, reference should be made to the appended claims to determine the scope of the present invention.
Claims
1. A non-transitory computer-readable medium having stored thereon a computer program, the computer program being configured to cause at least one processor to: providing source data as input to a cognitive AI layer, the cognitive AI layer configured to process the input; receiving an output from the cognitive AI layer including an automatically generated code; A non-transitory computer readable medium configured to implement the automatically generated code into a robotic process automation (RPA) workflow.
2. 2. The non-transitory computer readable medium of claim 1, wherein the source data comprises natural language text, a recording of a user's voice, a video recording of a user's actions on a computing system, a document, pseudocode for a desired task, a diagram of a desired user interface or form, a process diagram, a process document or process description, an RPA workflow, or any combination thereof.
3. The non-transitory computer readable medium of claim 1 , wherein the computer program comprises the cognitive AI layer.
4. 10. The non-transitory computer readable medium of claim 1, wherein the computer program is an RPA designer application.
5. 2. The non-transitory computer readable medium of claim 1 , wherein the output from the cognitive AI layer comprises the RPA workflow.
6. 2. The non-transitory computer-readable medium of claim 1, wherein the output from the cognitive AI layer includes one or more formulas and / or code snippets, and implementing the automatically-generated code includes updating the RPA workflow by adding the one or more formulas and / or code snippets to one or more respective activities of the RPA workflow.
7. 2. The non-transitory computer-readable medium of claim 1, wherein the output from the cognitive AI layer includes one or more configuration changes for the RPA workflow, and wherein implementing the automatically-generated code includes updating the RPA workflow by modifying one or more respective activities of the RPA workflow to include the one or more configuration changes.
8. 2. The non-transitory computer-readable medium of claim 1, wherein implementing the automatically generated code includes compiling the RPA workflow into an automation configured to be performed by an RPA robot at run time.
9. The computer program further comprises: prompting a user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not accepted; receiving one or more changes to the RPA workflow from the user; 2. The non-transitory computer readable medium of claim 1 configured to modify the RPA workflow to include the output from the cognitive AI layer with the modification from the user.
10. The computer program further comprises: prompting a user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not accepted; 2. The non-transitory computer-readable medium of claim 1 , configured to: fail automatic code generation and send data associated with the failure to retrain one or more AI / ML models of the cognitive AI layer.
11. The cognitive AI layer is a generative AI model configured to generate code, provide semantic associations between on-screen text, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate processes from drawings, or any combination thereof; and one or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation capabilities.
12. The computer program further comprises:
10. The non-transitory computer-readable medium of claim 1, configured to add security and / or compliance rules to the RPA workflow to comply with laws and / or policies.
13. a memory for storing computer program instructions; and at least one processor configured to execute the stored computer program instructions, the computer program instructions comprising: providing source data as input to a cognitive AI layer, the cognitive AI layer configured to process the input; receiving an output from the cognitive AI layer including an automatically generated code; configured to implement the automatically generated code into a robotic process automation (RPA) workflow; The source data may be generated from one or more computing systems including natural language text, a recording of a user's voice, a video recording of a user's actions on a computing system, documentation, pseudocode of a desired task, a diagram of a desired user interface or form, a process diagram, a process document or process description, an RPA workflow, or any combination thereof.
14. 14. The one or more computing systems of claim 13, wherein the computer program instructions include the cognitive AI layer.
15. 14. The one or more computing systems of claim 13, wherein the output from the cognitive AI layer comprises the RPA workflow.
16. 14. The one or more computing systems of claim 13, wherein the output from the cognitive AI layer includes one or more formulas and / or code snippets, and implementing the automatically-generated code includes updating the RPA workflow by adding the one or more formulas and / or code snippets to one or more respective activities of the RPA workflow.
17. 14. The one or more computing systems of claim 13, wherein the output from the cognitive AI layer includes one or more configuration changes for the RPA workflow, and wherein implementing the automatically generated code includes updating the RPA workflow by modifying one or more respective activities of the RPA workflow to include the one or more configuration changes.
18. 14. The one or more computing systems of claim 13, wherein implementing the automatically generated code includes compiling the RPA workflow into an automation configured to be performed by an RPA robot at run time.
19. The computer program instructions further include causing the at least one processor to: prompting the user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not accepted; receiving one or more changes to the RPA workflow from the user; 14. The one or more computing systems of claim 13, configured to modify the RPA workflow to include the output from the cognitive AI layer with the modification from the user.
20. The computer program instructions further include causing the at least one processor to: prompting the user regarding the output from the cognitive AI layer; receiving an indication from the user that the output from the cognitive AI layer is not accepted; 14. The one or more computing systems of claim 13, configured to fail automatic code generation and send data associated with the failure to retrain one or more AI / ML models of the cognitive AI layer.
21. The cognitive AI layer is a generative AI model configured to generate code, provide semantic associations between on-screen text, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate processes from drawings, or any combination thereof; and one or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation capabilities.
22. The computer program instructions further include causing the at least one processor to:
14. The one or more computing systems of claim 13, configured to add security and / or compliance rules to the RPA workflow to comply with laws and / or policies.
23. providing source data as input to a cognitive AI layer by a computing system, the cognitive AI layer being configured to process the input; receiving, by the computing system, an output from the cognitive AI layer, the output including automatically generated code; implementing, by the computing system, the automatically generated code into a robotic process automation (RPA) workflow; A computer-implemented method for performing automatic code generation, wherein the source data includes natural language text, a recording of a user's voice, a video recording of a user's actions on a computing system, documentation, pseudocode for a desired task, a diagram of a desired user interface or form, a process diagram, a process document or process description, an RPA workflow, or any combination thereof.
24. 24. The computer-implemented method of claim 23, wherein the output from the cognitive AI layer includes one or more formulas and / or code snippets, and implementing the automatically-generated code includes updating the RPA workflow by adding the one or more formulas and / or code snippets to one or more respective activities of the RPA workflow.
25. 24. The computer-implemented method of claim 23, wherein the output from the cognitive AI layer includes one or more configuration changes for the RPA workflow, and implementing the automatically-generated code includes updating the RPA workflow by modifying one or more respective activities of the RPA workflow to include the one or more configuration changes.
26. 24. The computer-implemented method of claim 23, wherein implementing the automatically generated code includes compiling the RPA workflow into an automation configured to be performed by an RPA robot at run time.
27. moreover, prompting the user regarding the output from the cognitive AI layer, receiving, by the computing system, an indication from the computing system that the user does not accept the output from the cognitive AI layer; receiving, by the computing system, one or more changes to the RPA workflow from the user; 24. The computer-implemented method of claim 23, comprising modifying, by the computing system, the RPA workflow to include the output from the cognitive AI layer with the modification from the user.
28. moreover, prompting the user regarding the output from the cognitive AI layer, receiving, by the computing system, an indication from the computing system that the user does not accept the output from the cognitive AI layer; 24. The computer-implemented method of claim 23, further comprising: failing the automatic code generation, by the computing system, and sending data associated with the failure to retrain one or more AI / ML models of the cognitive AI layer.
29. The cognitive AI layer is a generative AI model configured to generate code, provide semantic associations between on-screen text, suggest expressions and code snippets for activities of the RPA workflow, generate forms from natural language descriptions or drawings, generate processes from drawings, or any combination thereof; and one or more other AI / ML models configured to use output from the generative AI model to provide automatic code generation capabilities.
30. moreover, 24. The computer-implemented method of claim 23, comprising adding security and / or compliance rules to the RPA workflow to comply with laws and / or policies.