Automatic annotation and technical specification generation for robotic process automation workflow using artificial intelligence (AI)
By employing AI for automatic annotation and technical specification generation, the challenges of lacking documentation in RPA workflows are addressed, enhancing understanding and troubleshooting, and facilitating workflow conversion between vendors.
Patent Information
- Application Number
- JP2024066697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-04-17
- Publication Date
- 2025-06-30
AI Technical Summary
Manual construction, testing, publishing, and deployment of RPA workflows often lack annotation and documentation, making it difficult to understand and troubleshoot the automation process.
The use of AI for automatic annotation and technical specification generation of RPA workflows, where a cognitive AI layer processes the RPA workflow code or process definition document to provide annotations and explanations of each activity and the overall process.
This approach enables smarter search and documentation of RPA workflows, improving understanding and troubleshooting capabilities, and allowing for the conversion of workflows between different RPA vendors.
Smart Images

Figure 2025097253000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to artificial intelligence (AI), and more specifically, to automatic annotation and technical specification generation for robotic process automation (RPA) workflows using AI.
Background Art
[0002] When RPA workflows are manually constructed, tested, published, and deployed, there is often a lack of annotation and other documentation, making it difficult to understand what is done at each stage of the automation / workflow during troubleshooting / debugging. Therefore, improved and / or alternative approaches may be beneficial.
Summary of the Invention
[0003] Certain embodiments of the present invention may provide solutions to problems and needs in the art that are not yet fully specified, evaluated, or solved by current RPA technologies. For example, some embodiments of the present invention relate to automatic annotation and technical specification generation for RPA workflows using AI.
[0004] In an embodiment, a non-transitory computer-readable medium stores a computer program for providing automatic annotation to an RPA workflow. The computer program is configured such that at least one processor provides the code of the RPA workflow or process definition document (PDD) to a cognitive AI layer. The computer program is also configured such that at least one processor processes the code of the RPA workflow or PDD by the cognitive AI layer. The computer program is further configured such that at least one processor provides, as output, an annotation of the RPA workflow or RPA workflow code including the annotation by a generative AI layer.
[0005] In another embodiment, one or more computing systems include a memory storing computer program instructions for providing automated annotation of RPAs, and at least one processor configured to execute the computer program instructions. The computer program instructions are configured such that the at least one processor provides semantic associations between on-screen text that generates an annotated RPA workflow from the code of the RPA workflow or PDD, logically groups classes of RPA workflow activities that provide semantic associations between on-screen text that generates an annotated RPA workflow from the code of the RPA workflow or PDD, infers subsequent activities to add based on the context of the RPA workflow, converts the RPA workflow from the format of one RPA vendor to the format of another RPA vendor, or any combination thereof, to a cognitive AI layer that includes a generative AI model. The computer program instructions are also configured such that the at least one processor processes the code of the RPA workflow or PDD by the cognitive AI layer. The computer program instructions are further configured such that the at least one processor is provided with, as output, an annotation of the RPA workflow or RPA workflow code including the annotation by the generative AI layer. The annotation provides an explanation of the entire process of the RPA workflow, an explanation of what each activity within the RPA workflow is doing, or both.
[0006] In yet another embodiment, a computer-implemented method for providing automatic annotations to an RPA workflow includes providing, by an RPA designer application executed on a computing system, the code of the RPA workflow to a cognitive AI layer or a PDD. The cognitive AI layer is configured to process the code of the RPA workflow and / or the PDD. The computer-implemented method also includes receiving, by the RPA designer application, an annotated RPA workflow or RPA workflow code including annotations from a generative AI layer. The computer-implemented method further includes displaying, by the RPA designer application, the annotated RPA workflow. The annotations provide an explanation of the overall process of the RPA workflow, an explanation of what each activity of the RPA workflow is doing, an explanation of the changes between the current version of the RPA workflow and one or more previous versions of the RPA workflow, or any combination thereof.
[0007] In yet another embodiment, a non-transitory computer-readable medium stores a computer program. The computer program configures at least one processor to provide a natural language description of a process to a cognitive AI layer. The computer program also configures at least one processor to process the natural language description by the cognitive AI layer. The computer program further configures at least one processor to generate an annotated RPA workflow by a generative AI layer.
[0008] In another embodiment, one or more computing systems include a memory storing computer program instructions and at least one processor configured to execute the computer program instructions. The computer program instructions configure the at least one processor to provide a natural language description of a process to a cognitive AI layer. The computer program instructions also configure the at least one processor to process the natural language description by the cognitive AI layer. The computer program instructions further configure the at least one processor to generate an annotated robotic process automation (RPA) workflow by a generative AI layer.
[0009] In yet another embodiment, a computer-implemented method includes providing a natural language description of a process to a cognitive AI layer on one or more computing systems. The computer-implemented method also includes processing the natural language description by the cognitive AI layer. The computer-implemented method further includes generating an annotated RPA workflow by a generative AI layer. The annotation includes an explanation of the overall process of the RPA workflow, an explanation of what each activity of the RPA workflow is doing, an explanation of the changes between the current version of the RPA workflow and one or more previous versions of the RPA workflow, or any combination thereof.
Brief Description of the Drawings
[0010] To facilitate an understanding of the advantages of particular embodiments of the present invention, a more particular description of the invention briefly described above is depicted with reference to the specific embodiments illustrated in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope, but the invention will be described and explained with further particularity and detail by use of the following accompanying drawings.
[0011]
Figure 1
[0012]
Figure 2
[0013]
Figure 3
[0014]
Figure 4
[0015]
Figure 5
[0016]
Figure 6A
[0017]
Figure 6B
[0018]
Figure 7
[0019]
Figure 8
[0020]
Figure 9
[0021]
Figure 10
[0022]
Figure 11
[0023]
Figure 12
[0024] Unless otherwise specified, like reference characters consistently denote corresponding features throughout the accompanying drawings.
Mode for Carrying Out the Invention
[0025] (Detailed Description of the Embodiment) Some embodiments relate to automated annotation and technical specification generation for RPA workflows using AI. Often, the annotations and other documentation that explain what the activities of an RPA workflow do are sparse or completely missing. Consider the case of an RPA workflow activity called "code call". There are 10 items of code within the activity. An automation tester or subsequent user may not know what that activity does. In another example, if an RPA workflow includes 30 or 40 activities that have "browser use" as a parent activity, it may not be easy for users other than technicians to understand what is happening or to memorize the automation flow.
[0026] Some embodiments enable smart search of workflows by an artificial intelligence / machine learning (AI / ML) model and automatically generate documentation of the workflow, including an explanation of each activity, input / output parameters, and an explanation of the overall process. These embodiments can be employed as part of a cognitive AI layer. Such embodiments can communicate to the user what the activity, sequence, workflow, process, and / or overall automation is doing. The text may be displayed within the RPA workflow itself and / or provided in another document (e.g., a process definition document (PDD)). In some embodiments, by automatically creating the workflow and its respective annotations and / or documentation, workflows from other RPA applications can be converted into workflows of a desired different RPA designer application. In other words, previously created automations can be automated.
[0027] In some embodiments, software (e.g., an RPA designer application such as UiPath Studio™, another process that monitors user actions during the development of an RPA workflow, etc.) monitors the workflow as it is being created. When a user develops an application, the software provides RPA workflow information as input to invoke an AI / ML model(s). For example, the AI / ML model(s) may be provided with a screenshot of the RPA designer application interface, the code of the RPA workflow (e.g., an Extensible Markup Language (XML), Extensible Application Markup Language (XAML), Hypertext Markup Language (HTML), other suitable format file or file stream, etc.) including activities and their parameters (e.g., input, output, settings, type of UI descriptor, etc.), business process documents, etc. The AI / ML model(s) processes this information and provides recommended annotations, PDDs, etc. For example, the AI / ML model(s) may provide a version of the code for the RPA workflow including the annotations. Subsequently, the RPA designer application may display this version of the RPA workflow to the user.
[0028] In some embodiments, annotation and documentation may be provided for the overall complex business automation that is the sum of multiple workflows and applications, such as the entire adoption process in an enterprise, the enterprise procurement process, financing applications, etc. In such embodiments, the PDD may be used as a starting point. The PDD shows an overview of the business process selected for automation. Specifically, the PDD aims to build a foundation to ensure that the process is a suitable candidate for automation. The PDD explains a series of steps that are executed as part of the business process, the conditions and rules of the process before automation, and how they are expected to function after the process is partially or fully automated. The PDD functions as a reference tool for both users other than developers and technicians and provides the details necessary to apply RPA technology to the process.
[0029] In some embodiments, if there is no PDD for a business process, it can be generated from the RPA workflow code itself. The AI / ML model(s) retrieves the code of the RPA workflow(s) including the activity and its parameters, and automatically generates the PDD as an output from this input. In some embodiments, the AI / ML model(s) can also output an annotated version(s) of the RPA workflow. In some embodiments, other documents such as audit documents, compliance documents required by laws or regulations can also be created.
[0030] In some embodiments, the process is iterative. The generative AI model is used as part of the cognitive AI layer and can automatically convert text into RPA workflow code. Runtime automation can be generated from this RPA workflow, and other documentation such as PDD can also be generated. In such embodiments, when a business user comes up with a business process for which automation is desired, automation can be created and documented without the need for manual programming.
[0031] Figure 1 is an architecture diagram showing a hyper-automation system 100 according to an embodiment of the present invention. As used herein, "hyper-automation" refers to an automation system that combines components of process automation, integration tools, and technologies that amplify the ability to automate work. For example, in some embodiments, RPA is used at the core of the hyper-automation system, and in certain embodiments, the automation capabilities can be extended by AI / ML, process mining, analytics, and / or other advanced tools. When the hyper-automation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work can be automated, and both computing systems within an organization, such as those used by individuals and those operating autonomously, can all participate as part of the hyper-automation process. The hyper-automation systems of some embodiments enable users and organizations to discover, understand, and scale automation efficiently and effectively.
[0032] The hyper-automation system 100 includes user computing systems such as desktop computer 102, tablet 104, and smartphone 106. However, any desired user computing system, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc., can be used without departing from the scope of the present invention. Also, although three user computing systems are shown in Figure 1, any suitable number of user computing systems can be used without departing from the scope of the present invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of user computing systems can be used. The user computing systems may be actively used by users or may be automatically executed with little or no user input.
[0033] Each user computing system 102, 104, 106 has its respective automation process(es) 110, 112, 114 running thereon. In some embodiments, the automation process is stored remotely (e.g., on server 130 or on database 140) (accessed via network 120) and loaded by an RPA robot to implement automation. The automation may exist as a script (e.g., XML, XAML, etc.) or may be compiled into machine-readable code (e.g., as a digital link library).
[0034] The automation process(es) 110, 112, 114 may include, without limitation and without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application(s) for each computing system, any other suitable software and / or hardware, or any combination thereof. In some embodiments, one or more of the process(es) 110, 112, 114 may be a listener. The listener may be, without departing from the scope of the present invention, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the listener(s) is implemented partially or fully via physical hardware.
[0035] The listener monitors and records data related to user interactions with each computing system and / or the operation of an unattended computing system, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data can include, but is not limited to, which buttons were clicked, where the mouse moved, the text entered in a field, that one window was minimized and another window was opened, the application associated with the window, etc. In certain embodiments, the data from the listener can be transmitted periodically as part of a heartbeat message. In some embodiments, the data can be transmitted to the core hyper-automation system 120 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0036] The automation process can perform logic developed in the workflow during design time. In the case of RPA, the workflow can include a set of steps performed in a sequence or some other logical flow, defined herein as an "activity". Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, the workflow can be nested or embedded.
[0037] The long-running workflows for RPA in some embodiments are master projects that support service orchestration, human intervention, and long-running transactions in an unattended environment. See, for example, U.S. Patent No. 10,860,905, which is incorporated herein by reference in its entirety. Human intervention occurs when a particular process requires human input for exception handling, approval, or verification before proceeding to the next step of the activity. In this case, the execution of the process is paused and the RPA robot is released until the human task is completed.
[0038] The long-running workflows may support fragmentation of the workflow via persistence activities, combined with call processes and non-user interaction activities, and may orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or a large number of computing systems may participate in the execution of the logic of the long-running workflow. The long-running workflow may be executed in a session to facilitate rapid execution. In some embodiments, the long-running workflow may orchestrate a background process that executes API calls and includes activities that execute in the long-running workflow session. These activities may, in some embodiments, be called by a call process activity. A process having a user interaction activity that executes in a user session may be called by starting a job from a conductor activity (the conductor is described in more detail later in this specification). The user may, in some embodiments, interact through a task that requires the user to complete a form in the conductor. An activity may be included that causes the RPA robot to wait for the form task to be completed and then resume the long-running workflow.
[0039] One or more automation processes 110, 112, 114 communicate with a core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may execute conductor applications on one or more servers such as server 130. Although one server 130 is shown for illustration purposes, multiple or numerous servers in close proximity to each other or in a distributed architecture may be employed without departing from the scope of the present invention. For example, one or more servers may be provided for conductor functions, AI / ML model provision, authentication, governance, and / or any other suitable functions without departing from the scope of the present invention. In some embodiments, the core hyper-automation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyper-automation system 120 may host multiple software-based servers on one or more computing systems such as server 130. In some embodiments, one or more servers of the core hyper-automation system 120, such as server 130, may be implemented via one or more virtual machines (VMs).
[0040] In some embodiments, one or more automation processes 110, 112, 114 may invoke one or more AI / ML models 132 that are deployed on or accessible by a core hyper-automation system 120 and trained to accomplish various tasks. For example, the AI / ML models 132 may include models trained to search for various application versions, perform CV, perform OCR, generate UI descriptors, and provide suggestions for the next activity or sequence of activities in an RPA workflow. The AI / ML models may be trained using labeled data including elements of data sources (e.g., web pages, forms, scanned documents, application interfaces, screens, etc.), previously created RPA workflows, screenshots of various application screens of various versions including corresponding UI elements, libraries of UI objects, and the like. The AI / ML models 132 may be trained to achieve a desired confidence threshold without overfitting to a given set of training data.
[0041] The AI / ML model 132 can be trained for any suitable purpose without departing from the scope of the present invention, as will be discussed in more detail later in this specification. Two or more AI / ML models 132 may be chained in some embodiments (e.g., in series, in parallel, or a combination thereof) such that they collectively provide collaborative output(s). The AI / ML model 132 may perform or assist with CV, OCR, document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automatic RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, any combination of these, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of the present invention. By using multiple AI / ML models, for example, the system can develop an overall picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, etc. Patterns may be determined individually by an AI / ML model or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one of the computing systems 102, 104, 106.
[0042] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that executes on data, and the AI / ML model itself may be, for example, a deep learning neural network (DLNN) of artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions such as statistical modeling (e.g., hidden Markov model (HMM)), and may utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform the desired functions.
[0043] In some embodiments, the Hyper Automation System 100 may provide four main functional groups: (1) Discovery, (2) Automation Construction, (3) Management, and (4) Engagement. Automation (e.g., executed on user computing systems, servers, etc.) may, in some embodiments, be performed by software robots such as RPA robots. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with users to assist them in tasks (e.g., via UiPath Assistant (trademark)). Unattended robots operate independently of users and may potentially execute in the background while the users are unaware. Test robots are unattended robots that execute test cases against applications or RPA workflows. Test robots may, in some embodiments, be executed in parallel on multiple computing systems.
[0044] The Discovery function may discover various opportunities for automating business processes and provide automated recommendations therefor. Such a function may be implemented by one or more servers such as server 130. The Discovery function may, in some embodiments, include providing an Automation Hub, Process Mining, Task Mining, and / or Task Capture. The Automation Hub (e.g., UiPath Automation Hub (trademark)) may provide a mechanism for managing the rollout of automation with visibility and control. Automation ideas may be crowdsourced from employees, for example, via a submission form. Feasibility and ROI calculations for automating these ideas are provided, documentation for future automation is collected, and collaboration for quickly going from discovery to construction of automation may be provided.
[0045] Process mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Center™) refers to the process of collecting and analyzing data from applications (such as enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how they can be effectively automated, and the impact of automation. This data can be obtained, for example, by a listener from user computing systems 102, 104, 106 and processed by a server such as server 130. In some embodiments, one or more AI / ML models 132 may be employed for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of the goals of process mining include, but are not limited to, increased profitability, improved customer satisfaction, regulatory and / or compliance, and improved employee efficiency.
[0046] Task mining (e.g., via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in daily tasks and scores such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to reveal repetitive task patterns in the data. Repetitive tasks ripe for automation can then be identified. This information can initially be provided by a listener and, in some embodiments, analyzed on a server of a core hyperautomation system 120 such as server 130. Discoveries from task mining (e.g., XAML process data) are exported to a process document or a designer application such as UiPath Studio (trademark) to enable faster creation and deployment of automation. Task mining in some embodiments may include taking screenshots with user actions (e.g., mouse click location, keyboard input, application windows and graphical elements the user interacted with, timestamps for interactions, etc.), collecting statistical data (e.g., execution time, number of actions, text input, etc.), editing and annotating screenshots, specifying the types of actions being recorded, etc.
[0047] Task capture (via UiPath Automation Cloud™ and / or UiPath AI Center™) automatically documents attended processes as the user works or provides a framework for unattended processes. Such documentation may include tasks that are desirable to automate in the form of PDDs, skeleton workflows, capture of actions for each part of the process, recording of user actions and automatic generation of comprehensive workflow diagrams including details for each step, Microsoft Word® documents, XAML files, etc. The constructible workflows can, in some embodiments, be directly exported to designer applications such as UiPath Studio™. Task capture can simplify the requirements gathering process for both subject matter experts who explain the process and Center of Excellence (CoE) members who provide production-grade automation.
[0048] Automation can be achieved through designer applications (such as UiPath Studio (trademark), UiPath StudioX (trademark), UiPath Studio Web (trademark), etc.). For example, RPA developers at the RPA development facility 150 can use the RPA designer application 154 on the computing system 152 to build and test automation for various applications and environments such as web, mobile, SAP (registered trademark), and virtual desktops. API integration can be provided for various applications, technologies, and platforms. Pre-defined activities, drag-and-drop modeling, and workflow recorders can facilitate automation with minimal coding. The document understanding function can be provided via drag-and-drop AI skills for data extraction and interpretation that call one or more AI / ML models 132. Such automation can process virtually any document type and format, including tables, checkboxes, signatures, and handwritten. When data is verified or exceptions are processed, this information may be used to retrain the respective AI / ML models, and their accuracy is improved over time.
[0049] The RPA designer application 152 can be designed to call one or more of the trained AI / ML models 132 on the server 130 and / or the generative AI model 172 within the cloud environment via a network 120 (such as a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.) to assist in the RPA automation development process. In some embodiments, one or more of the AI / ML models can be packaged with the RPA designer application 152 or otherwise stored locally on the computing system 150.
[0050] In some embodiments, the RPA designer applications 152 and one or more AI / ML models 132 can be configured to use an object repository stored in the database 140. For example, see U.S. Patent No. 11,748,069, which is hereby incorporated by reference in its entirety. The object repository can include a library of UI objects that can be used to develop RPA workflows via the RPA designer application 152. The object repository can be used to add UI descriptors to activities within the workflows of the RPA designer application 152 for UI automation. In some embodiments, one or more of the AI / ML models 132 can generate new UI descriptors and add them to the object repository within the database 140. When automation is completed in the designer application 152, the automation can be published on the server 130 and pushed out to computing systems 102, 104, 106, etc.
[0051] With the integrated service, developers can seamlessly combine, for example, UI automation and API automation. Automations that require APIs or that cross both API and non-API applications and systems can be built. A repository (e.g., UiPath Object Repository (trademark)) or marketplace (e.g., UiPath Marketplace (trademark)) for pre-built RPA and AI templates and solutions can be provided so that developers can automate a wide variety of processes more quickly. Thus, when building an automation, the hyper-automation system 100 can provide a user interface, a development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. The hyper-automation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.
[0052] In some embodiments, components of the hyper-automation system 100, such as designer applications and / or external rule engines, provide support for managing and enforcing governance policies for controlling the various functions provided by the hyper-automation system 100. Governance refers to the ability of an organization to introduce policies to prevent users from developing automations (such as RPA robots) that can harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), the terms of use of third-party applications, etc. Otherwise, developers could create automations that violate privacy laws, terms of use, etc. during the execution of their automations. Thus, in some embodiments, access control and governance restrictions are implemented at the robot and / or robot design application level. This can provide an additional level of security and compliance in the automation process development pipeline in some embodiments by preventing developers from introducing security risks or relying on unapproved software libraries that could operate in a way that violates policies, regulations, privacy laws, and / or privacy policies. See, for example, U.S. Patent No. 11,733,668, which is incorporated herein by reference in its entirety.
[0053] The management function can provide management, deployment, and optimization of automation across the entire organization. The management function may include, in some embodiments, orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management function of the hyper-automation system 100 can include, among other things, but not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots.
[0054] Conductor applications such as UiPath Orchestrator (trademark) (which may be provided as part of UiPath Automation Cloud (trademark) in some embodiments, or on-premises, VM, private or public cloud, on a Linux (trademark) VM, or as a cloud-native single-container suite via UiPath Automation Suite (trademark)) provide orchestration capabilities to deploy, monitor, optimize, scale, and secure RPA robot deployments. A test suite (e.g., UiPath Test Suite (trademark)) can provide test management for monitoring the quality of deployed automation. The test suite can facilitate test planning and execution, requirement fulfillment, and defect traceability. The test suite can include comprehensive test reports.
[0055] Analytics software (e.g., UiPath Insights (trademark)) can track, measure, and manage the performance of deployed automation. The analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes of the organization. The analytics software can present results in dashboard form for easier understanding by human users.
[0056] A data service (e.g., UiPath Data Service (trademark)) can, for example, be stored in a database 140 and bring data into a single, scalable, and secure location using a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities may be provided by an AI center (e.g., UiPath AI Center (trademark)), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may enable non-data scientists to access such capabilities. Deployed automation (e.g., an RPA robot) can call an AI / ML model from an AI center such as AI / ML model 132. The performance of the AI / ML model can be monitored and trained and improved using human-verified data such as that provided by a data review center 160. A human reviewer may provide labeled data to the core hyperautomation system 120 via a review application 152 on a computing system 154. For example, a human reviewer may verify that predictions by AI / ML model 132 and / or generative AI model 172 are accurate or, if not, provide corrections. This dynamic input may then be saved as training data for retraining AI / ML model 132 and / or generative AI model 172 and stored, for example, in a database such as database 140. The AI center can then schedule and execute a training job to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining AI / ML model 132 and / or generative AI model 172.
[0057] The engagement function involves humans and automation as one team for seamless collaboration regarding a desired process. Low-code applications can be built (e.g., via UiPath Apps™) even if they lack an API in some embodiments to connect browser tabs and legacy software. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. An application can be connected to one automation or multiple automations.
[0058] The Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for passing a process from automation to humans or vice versa. Humans can provide approvals or escalations and perform exception handling, etc. Then, the automation can execute the automated functions of a given workflow.
[0059] The local assistant may be provided as a launch pad for the user to launch an automation (e.g., UiPath Assistant (trademark)). This feature may be provided, for example, in the tray provided by the operating system, enabling the user to interact with RPA robots and RPA robot - enabled applications on their computing system. The interface may list the automations approved for a given user and allow the user to execute them. These may include off - the - shelf automations from an automation marketplace, an internal automation store of an automation hub, etc. When an automation is running, they may execute as a local instance in parallel with other processes on the computing system so that the user can use the computing system while the automation performs its actions. In certain embodiments, the assistant is integrated with a task capture function so that the user can document the processes that are about to be automated from the assistant's launch pad.
[0060] Chatbots (e.g., UiPath Chatbots (trademark)), social messaging applications, and / or voice commands may enable the user to execute an automation. This may simplify access to the information, tools, and resources necessary to conduct customer interactions or other activities. Human - to - human conversations can be easily automated just like other processes. Trigger RPA robots launched in this way may be able to perform actions such as order status checks, data posting to CRM, etc., using plain - language commands.
[0061] End-to-end measurement of automation programs at any scale and governance can be provided by the hyper-automation system 100 in some embodiments. As such, analytics (e.g., via UiPath Insights™) may be employed to understand the performance of the automation. Data modeling and analytics using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for automation, perform remote instrumentation monitoring on the user's computing system, detect errors and anomalies, and debug the automation. An automation management console (e.g., UiPath Automation Ops™) may be provided to manage the automation throughout its lifecycle. An organization may govern how the automations are built, what users can do with them, and which automations users can access.
[0062] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes can be discovered, automations can be built, tested, and deployed, performance can be measured, use of the automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the processes themselves can be repeated. This promotes a more robust and effective set of automations.
[0063] In some embodiments, a generative AI model is used. Generative AI can generate various types of content such as text, images, audio, and synthetic data. Various types of generative AI models can be used, including but not limited to large language models (LLMs), generative adversarial networks (GANs), variational autoencoders (VAEs), transformers, etc. These models may be part of the AI / ML model 132 hosted on the server 130. For example, the generative AI model can be trained on a large corpus of text information to perform semantic understanding, understand the nature of what exists on the screen from text, automatically generate code, etc. In certain embodiments, a generative AI model 172 provided by an existing cloud ML service provider such as OpenAI®, Google®, Amazon®, Microsoft®, IBM®, Nvidia®, Facebook® can be employed and trained to provide such functionality. In a generative AI embodiment where the generative AI model(s) 172 is hosted remotely, the server 130 can be configured to integrate with a third-party API, whereby the server 130 can send requests containing the required input information to the generative AI model(s) 172 and receive its responses (e.g., semantic matching of fields between versions of an application, classification of the type of application on the screen, etc.). Such embodiments can not only provide a more advanced and sophisticated user experience but also provide access to state-of-the-art natural language processing (NLP) and other ML capabilities provided by these companies.
[0064] One aspect of the generative AI model in some embodiments is the use of transfer learning. In transfer learning, a pre-trained generative AI model such as an LLM is fine-tuned for a specific task or domain. This allows the LLM to leverage the knowledge it has already learned during its initial training and adapt it to a specific application. In the case of an LLM, during the pre-training phase, the LLM is typically trained on a large text corpus consisting of billions of words. During this phase, the LLM learns the relationships between words and phrases, enabling it to generate consistent human-like responses to text-based inputs. The output of this pre-training phase is an LLM that highly understands the patterns underlying natural language.
[0065] In the fine-tuning phase, the pre-trained LLM is adapted to a specific task or domain by training the LLM on a smaller, task-specific dataset. For example, in some embodiments, the LLM can be trained to analyze a specific type or types of data sources to improve the accuracy regarding its content. Such information can be provided as part of the training data, and the LLM can learn to focus on these areas and more accurately identify the data elements within them. Through fine-tuning, without requiring as much data as would be necessary to train the LLM from scratch, the LLM can learn the subtle nuances of the task or domain, such as the specific vocabulary and syntax used in that domain. By leveraging the knowledge learned during the pre-training phase, the fine-tuned LLM can achieve state-of-the-art performance on a specific task with a relatively small amount of training data.
[0066] LLMs can be trained using vector databases. Vector databases index, store, and provide access to structured or unstructured data (e.g., text, images, time series data, etc.) along with their vector embeddings. Data such as text can be tokenized, where single characters, words, or sequences of words are parsed from the text into tokens. These tokens are then "embedded" into vector embeddings, which are numerical representations of this data. Vector databases allow users to search and retrieve similar objects quickly and at scale in production environments.
[0067] AI and ML allow unstructured data to be represented numerically without losing its semantic meaning in vector embeddings. A vector embedding is a long list of numbers, where each number represents a feature of the data object that the vector embedding represents. Similar objects are grouped together in the vector space. In other words, the more similar the objects are, the closer the vector embeddings that represent them are to each other. Similar objects may be found using vector search, similarity search, or semantic search. The distance between vector embeddings may be calculated using various techniques, including but not limited to Euclidean squared or L2 squared distance, Manhattan or L1 distance, cosine similarity, dot product, Hamming distance, etc. It may be beneficial to choose the same metric used to train the AI / ML model.
[0068] Vector indexing can be used to organize vector embeddings to enable efficient data retrieval. When the number of data points is large, using the k-nearest neighbor (kNN) algorithm to calculate the distance between a vector embedding in a vector database and every other vector embedding can be computationally expensive because the necessary computations increase linearly (O(n)) depending on the dimension and the number of data points. It is more efficient to use an approximate nearest neighbor (ANN) approach to find similar objects. The distances between vector embeddings are pre-computed, and similar vectors are organized and stored close to each other (e.g., within a cluster or graph), enabling faster discovery of similar objects. This process is called "vector indexing." ANN algorithms that can be used in some embodiments include, but are not limited to, clustering-based indexing, proximity graph-based indexing, tree-based indexing, hash-based indexing, compression-based indexing, etc.
[0069] Figure 2 is an architectural diagram showing an RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 may facilitate the development of automation projects that are graphical representations of business processes. Briefly, the designer 210 facilitates the development and deployment of workflows and robots. In some embodiments, the designer 210 may be an application that runs on a user's desktop, an application that runs remotely on a VM, a web application, or the like.
[0070] Automation projects enable the automation of rule-based processes by giving developers control over the execution order and relationships between steps in a custom set of workflows defined as "activities" in this specification as described above. A commercial example of an embodiment of Designer 210 is UiPath Studio (trademark). Each activity may include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.
[0071] Some types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers, etc. Sequences may be particularly suitable for linear processes that enable the flow from one activity to another without cluttering the workflow. Flowcharts may be particularly suitable for more complex business logic, enabling the integration of decision-making and the connection of activities in more diverse ways through multiple branching logic operators. FSMs may be particularly suitable for large-scale workflows. FSMs may use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. Global exception handlers may be particularly suitable for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0072] When a workflow is developed within the designer 210, the execution of the business process is coordinated by the conductor 220, which coordinates one or more robots 230 that execute the workflow developed within the designer 210. A commercial example of an embodiment of the conductor 220 is UiPath Orchestrator™. The conductor 220 facilitates the management of the generation, monitoring, and deployment of resources in an environment. The conductor 220 can operate as an integration point with third-party solutions and applications. As such, in some embodiments, the conductor 220 can be part of the core hyperautomation system 120 of FIG. 1.
[0073] The conductor 220 can manage all robots 230 and connect and execute the robots 230 from a central point. The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non-production robots (similar to attended robots 232 but used for development and testing purposes). Attended robots 232 are triggered by user events and operate alongside humans on the same computing system. Attended robots 232 can be used with the conductor 220 for centralized process deployment and logging medium. Attended robots 232 may assist human users in achieving various tasks and may be triggered by user events. In some embodiments, processes cannot start from the conductor 220 on this type of robot and / or they cannot execute under locked screens. In certain embodiments, attended robots 232 can only be launched from a robot tray or a command prompt. Attended robots 232 preferably operate under human supervision in some embodiments.
[0074] The unattended robot 234 can operate unmanned in a virtual environment and automate many processes. The unattended robot 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can, in some embodiments, be performed by the designer 210. Both attended and unattended robots can automate a variety of systems and applications including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0075] The conductor 220 may have various capabilities including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining a connection between the robot 230 and the conductor 220 (e.g., a web application). Deployment may include ensuring the correct delivery of the package version to the assigned robot 230 for execution. Configuration may include maintaining and delivering the robot environment and process configuration. Queuing may include providing management of queues and queue items. Monitoring may include tracking specific data of the robot and maintaining user permissions. Logging may include saving and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a “not only” SQL (NoSQL) database) and / or another storage mechanism (e.g., ElasticSearch® which provides the ability to store large datasets and execute queries quickly). The conductor 220 may provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0076] The robot 230 is an execution agent that implements the workflow constructed by the designer 210. One commercial example of some embodiments of the robot(s) 230 is UiPath Robots™. In some embodiments, the robot 230, by default, installs the Microsoft Windows® Service Control Manager (SCM) management service. As a result, such a robot 230 can open an interactive Windows® session under a local system account and may have the rights of a Windows® service.
[0077] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, it means having the same rights as the user in which a given robot 230 is installed. This feature can also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0078] The robot 230 in some embodiments is divided into multiple components, each specialized for a specific automation task. The robot components in some embodiments include, but are not limited to, SCM management robot services, user mode robot services, an executor, an agent, and a command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 is executed). These services are entrusted with managing the qualification information of the robot 230. The console application is launched by the SCM under the local system.
[0079] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host. The user mode robot service can be entrusted with managing the qualification information of the robot 230. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0080] The executor can perform a job given under a Windows® session (i.e., can execute a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows® Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0081] As described above, the fact that the components of the robot 230 are divided helps developers, support users, and computing systems to more easily execute, identify, and track what each component is doing. In this way, special behaviors can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 can be made independent of the browser zoom level. In the case of an application marked as not recognizing or intentionally not recognizing DPI, DPI can be disabled in some embodiments.
[0082] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for the execution of the AI / ML model(s) and receive outputs therefrom via the core hyper-automation system 240.
[0083] One or more robots 230 may be listeners, as described above. These listeners can provide information to the core hyper-automation system 240 regarding what users are doing when they use their computing systems. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0084] An assistant / chatbot 250 can be provided on the user computing system to enable the user to launch an RPA local robot. The assistant can be placed, for example, in the system tray. The chatbot can have a user interface so that the user can view the text of the chatbot. Alternatively, the chatbot can run in the background without a user interface and can listen for the user's utterances using the microphone of the computing system.
[0085] In some embodiments, data labeling may be performed by a user of the computing system that the robot is executing on, or on another computing system that the robot provides information to. For example, if the robot calls an AI / ML model that performs CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user may draw a rectangle around the mis-identified or non-identified component and potentially provide text with the correct identification. This information may be provided to the core hyper-automation system 240 and may then be used later for training a new version of the AI / ML model.
[0086] Figure 3 is an architecture diagram showing an expanded RPA system 300 according to an embodiment of the present invention. In some embodiments, the RPA system 300 may be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The expanded RPA system 300 may be a cloud-based system, an on-premises system, a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes, etc.
[0087] It should be noted that the client side, the server side, or both can include any desired number of computing systems without departing from the scope of the present invention. On the client side, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is executing a process. As shown in FIG. 3, a plurality of business projects can be executed simultaneously. The agent 314 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via a database server 350, an AI / ML server 360, an indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 can be a robot component.
[0088] In some embodiments, the robot represents an association between a machine name and a user name. The robot can manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots can be executed simultaneously, each running in a separate Windows® session using a unique user name. This is referred to as the HD robot described above.
[0089] Agent 314 is also responsible for sending the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. In some embodiments, the communication between Agent 314 and Conducter 340 is always initiated by Agent 314. In a notification scenario, Agent 314 may open a WebSocket channel that will later be used by Conducter 340 to send commands (e.g., start, stop, etc.) to the robot.
[0090] Listener 330 monitors and records data related to user interactions with the operation of the attended computing system and / or unattended computing system in which listener 330 resides. Listener 330 can be an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware without departing from the scope of the present invention. In fact, in some embodiments, the listener logic is implemented partially or fully via physical hardware.
[0091] On the server side, there are a presentation layer (web application 342, open data protocol (oData) representational state transfer (REST) application programming interface (API) endpoint 344, notification and monitoring 346), a service layer (API implementation / business logic 348), and a persistence layer (database server 350, AI / ML server 360, indexer server 370). The conductor 340 includes the web application 342, the oData REST API endpoint 344, notification and monitoring 346, and the API implementation / business logic 348. In some embodiments, most actions performed by a user at the interface of the conductor 340 (e.g., via the browser 320) are executed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of the present invention. The web application 342 is the visual layer of the server platform. In this embodiment, the web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, script language, or any other format may be used without departing from the scope of the present invention. The user interacts with the web page from the web application 342 via the browser 320 in this embodiment to perform various operations for controlling the conductor 340. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or process, start and stop robots, etc.
[0092] In addition to the web application 342, the conductor 340 also includes a service layer that exposes an oData REST API endpoint 344. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 342 and the agent 314. The agent 314 is, in this embodiment, a supervisor for one or more robots on a client computer.
[0093] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functions. Configuration endpoints may be used, in some embodiments, to define and configure the users, permissions, robots, assets, releases, and environments of the application. The logging REST endpoint may be used, for example, to log various information such as errors, explicit messages sent by the robots, and other environment-specific information. The deployment REST endpoint may be used by the robots to query the version of the package to be executed when a job start command is used in the conductor 340. The queuing REST endpoint may be responsible for the management of queues and queue items, such as adding data to the queue, retrieving transactions from the queue, and setting the status of transactions.
[0094] Monitoring of the REST endpoints may monitor the web application 342 and the agent 314. The notification and monitoring API 346 may be a REST endpoint used for the registration of the agent 314, the delivery of configuration settings to the agent 314, and the sending and receiving of notifications from the server and the agent 314. The notification and monitoring API 346 may use WebSocket communication in some embodiments.
[0095] In some embodiments, the API of the service layer can be accessed through the configuration of an appropriate API access path, for example, based on whether the conductor 340 and the overall hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The API for the conductor 340 can provide custom methods for querying statistics regarding various entities registered with the conductor 340. In some embodiments, each logical resource may be an oData entity. In such an entity, components such as robots, processes, queues, etc. may have properties, relationships, and operations. The API of the conductor 340 can be consumed by the web application 342 and / or the agent 314 in two ways in some embodiments: by obtaining API access information from the conductor 340 or by registering an external application to use the oAuth flow.
[0096] In this embodiment, the persistent layer includes three servers: a database server 350 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides AI / ML model providing services such as AI center functions), and an indexer server 370. The database server 350 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 342. The database server 350 may manage queues and queue items. In some embodiments, the database server 350 may store (in addition to or instead of the indexer server 370) messages recorded by robots. The database server 350 may also store, for example, process mining, task mining, and / or task capture-related data received from a listener 330 installed on the client side. Although no arrow is shown between the listener 330 and the database 350, it should be understood that in some embodiments, the listener 330 can communicate with the database 350 and vice versa. This data can be stored in the form of PDD, images, XAML files, etc. The listener 330 may be configured to eavesdrop on user actions, processes, tasks, and performance metrics on each computing system where the listener 330 resides. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, time, etc.) on its respective computing system and then convert them into a form suitable for being provided to and stored in the database server 350.
[0097] The AI / ML server 360 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can enable non-data scientists to access such capabilities. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of the AI / ML models can be monitored and trained and improved using human-verified data. The AI / ML server 360 can schedule and execute training jobs to train new versions of the AI / ML models.
[0098] The AI / ML server 360 can store data related to AI / ML models and ML packages for configuring various ML skills for users during development. The ML skills used herein are, for example, pre-built and trained ML models for processes that can be used by automation. The AI / ML server 360 can also store data related to document understanding technologies and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, NLP, speech analysis, different types of AI / ML models, etc.
[0099] Optionally in some embodiments, the indexer server 370 stores information recorded by robots and creates an index. In certain embodiments, the indexer server 370 may be disabled via configuration settings. In some embodiments, the indexer server 370 uses ElasticSearch®, an open-source project full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 370 via logging REST endpoint(s), where they are indexed for future use.
[0100] Figure 4 is an architecture diagram illustrating the relationship 400 between designer 410, activities 420, 430, 440, 450, driver 460, API 470, and AI / ML model 480 according to an embodiment of the present invention. As described above, a developer uses designer 410 to develop a workflow to be performed by a robot. Various types of activities may be presented to the developer in some embodiments. Designer 410 may be local or remote to the user's computing system (e.g., accessed via a local web browser interacting with a VM or remote web server). The workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. User-defined activities 420 and API-driven activities 440 interact with applications via their APIs. User-defined activities 420 and / or AI / ML activities 440 may call one or more AI / ML models 480 that may be located locally and / or remotely to the computing system on which the robot operates in some embodiments.
[0101] In some embodiments, non-text visual components in an image can be identified, which is referred to as CV herein. However, it should be noted that in some embodiments, CV incorporates OCR. CV can be at least partially executed by AI / ML model(s) 480. Some CV activities related to such components can include, but are not limited to, extraction of text from segmented label data using OCR, fuzzy text matching, cropping of segmented label data using ML, comparison of the extracted text in the label data with ground truth data, etc. In some embodiments, the number of activities that can be implemented in user-defined activity 420 can be hundreds or thousands. However, any number and / or type of activities can be used without departing from the scope of the present invention.
[0102] UI automation activity 450 is a subset of special low-level activities described in low-level code that facilitate interaction with the screen. UI automation activity 450 facilitates these interactions via a driver 460 that enables the robot to interact with the desired software. For example, driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, etc. In some embodiments, for performing interactions with a computing system, one or more AI / ML models 480 can be used by UI automation activity 450. In certain embodiments, the AI / ML models 480 can enhance or completely replace the drivers 460. In fact, in certain embodiments, the drivers 460 are not included.
[0103] Driver 460 can interact with the OS at a low level, such as searching for hooks and monitoring keys, via the OS driver 462. Driver 460 may facilitate integration with Chrome (registered trademark), IE (registered trademark), Citrix (registered trademark), SAP (registered trademark), etc. For example, a "click" activity serves the same role in these different applications via driver 460.
[0104] FIG. 5 is an architectural diagram showing a computing system 500 configured to perform automatic annotation and technical specification generation of an RPA workflow using AI according to an embodiment of the present invention. In some embodiments, computing system 500 may be one or more computing systems depicted and / or described herein. In certain embodiments, computing system 500 may be part of a hyperautomation system as shown in FIGS. 1 and 2. Computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to bus 505 for processing information. The processor(s) 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The processor(s) 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one processor(s) 510 can be a neuromorphic circuit that includes processing elements mimicking biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0105] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by processor(s) 510. Memory 515 can be composed of random access memory (RAM), read-only memory (ROM), flash memory, cache, a static storage device such as a magnetic disk or optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media can be any available media accessible by processor(s) 510, and can include volatile media, non-volatile media, or both. Also, the media can be removable, non-removable, or both. Computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, communication device 520 can include one or more antennas that are a single antenna, an array of antennas, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration without departing from the scope of the present invention.
[0106] The processor(s) 510 is further coupled to the display 525 via the bus 505. Without departing from the scope of the present invention, any suitable display device and tactile I / O may be used. The keyboard 530 and the cursor control device 535, such as a computer mouse, a touch pad, etc., are further coupled to the bus 505 to enable the user to interface with the computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user may interact with the device only via the display 525 and / or a touch pad (not shown). Any type and combination of input devices may be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, the user may interact with the computing system 500 remotely via another computing system communicating with the computing system 500, or the computing system 500 may operate autonomously.
[0107] The memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include an auto-generated module 545 configured to execute all or a portion of the AI / ML processes described herein or derivatives thereof. The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0108] One skilled in the art will understand that the "system" can be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a smart watch, a quantum computing system, or any other suitable computing device, or a combination of devices, without departing from the scope of the present invention. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. In fact, the methods, systems, and apparatuses disclosed herein may be implemented in a localized form and a distributed form that is consistent with computing techniques including cloud computing systems. The computing system may be part of, or accessible by, a LAN, a mobile communication network, a satellite communication network, the Internet, a public cloud or a private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of the present invention.
[0109] It should be noted that some of the system features described herein are presented as modules to emphasize implementation independence more. For example, a module can be implemented as a hardware circuit including custom very large scale integration (VLSI) circuits or gate arrays, logic chips, transistors, or other off-the-shelf semiconductors such as individual components. Also, a module can be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units.
[0110] The module can also be at least partially implemented in software for execution by various types of processors. For example, a specified unit of executable code can include one or more physical or logical blocks of computer instructions that may be organized, for example, as objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and can include modules when logically combined and can include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module can be stored on a non-transitory computer-readable medium such as, for example, a hard disk drive, a flash device, RAM, a tape, and / or any other non-transitory computer-readable medium used for storing data without departing from the scope of the present invention.
[0111] In fact, a module of executable code can be a single instruction, a number of instructions, or even be distributed among multiple different code segments, between different programs, and among multiple memory devices. Similarly, the operational data can be specified within the module, shown here, and can be embodied in any suitable form and organized within any suitable type of data structure. The operational data can be collected as a single data set or can be distributed in different locations across different storage devices and can exist, at least in part, simply as electronic signals on a system or network.
[0112] Without departing from the scope of the present invention, various types of AI / ML models can be trained and deployed. For example, FIG. 6A shows an example of a neural network 600 trained to complement the automatic annotation and technical specification generation of an RPA workflow using AI, according to an embodiment of the present invention. The neural network 600 includes a number of hidden layers. Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) typically have multiple layers, but an SLNN may sometimes have only one or two layers and usually has fewer layers than a DLNN. Typically, the architecture of a neural network includes an input layer, a plurality of intermediate layers, and an output layer, as in the case of the neural network 600.
[0113] Often, a DLNN has many layers (such as 10, 50, 200, etc.), and subsequent layers typically reuse the functions from the previous layers to compute more complex and general functions. On the other hand, an SLNN has only a few layers and tends to be trained relatively quickly because expert functions are pre-created from raw data samples. However, feature extraction is cumbersome. On the other hand, a DLNN usually does not require expert functions but takes more time to train and tends to have more layers.
[0114] In either approach, the layers are trained simultaneously on the training set and usually checked for overfitting on a separate cross-validation set. Excellent results are obtained with both techniques, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem being addressed by each neural network.
[0115] Returning to FIG. 6A, an RPA workflow, PDD and / or other documents, UI screenshots, RPA workflows and annotations, etc. are provided as the input layer and supplied as input to the J neurons of the hidden layer 1. Various other inputs are possible, including but not limited to the state information of the computing system, published automation, business rules, information related to things related to RPA workflows and / or tasks, initial definitions of automation, process automation documents, etc. In this example, all of these inputs are supplied to each neuron, but not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, noise removal autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks that do not depart from the scope of the present invention. Various architectures can be used individually or in combination.
[0116] The hidden layer 2 receives input from the hidden layer 1, the hidden layer 3 receives input from the hidden layer 2, and the same is done for all hidden layers until the last hidden layer provides its output as the input to the output layer. Although multiple proposals are shown as outputs in this specification, in some embodiments, only a single output proposal is provided. In certain embodiments, the proposals are ranked based on a confidence score.
[0117] Note that the numbers of neurons I, J, K, and L are not necessarily equal. Thus, any desired number of layers can be used in a given layer of neural network 600 without departing from the scope of the present invention. In fact, in certain embodiments, the types of neurons in a given layer may not all be the same.
[0118] Neural network 600 is trained to assign confidence score(s) to the appropriate output. To reduce inaccurate predictions, in some embodiments, only those results with a confidence score above a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs with a confidence score exceeding this amount may be used and the rest may be ignored.
[0119] A neural network is typically a probabilistic construct that has confidence score(s). This can be a score learned by the AI / ML model based on the frequency with which similar inputs were correctly identified during training. Some common types of confidence scores include a decimal number between 0 and 1 (which can also be interpreted as a confidence percentage), a numerical value between negative infinity and positive infinity, a series of expressions (e.g., "low", "medium", and "high"), etc. Various post-processing calibration techniques such as temperature scaling, batch normalization, weight decay, negative log likelihood (NLL), etc. can also be used to obtain a more accurate confidence score.
[0120] The "neurons" of a neural network are typically implemented algorithmically as mathematical functions based on the functions of biological neurons. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass the output to the next layer. This activation function can be a non-linear thresholded activity function that does nothing if the value is below the threshold, and responds linearly when the function exceeds the threshold (i.e., rectified linear unit (ReLU) non-linearity). Since actual neurons can have a nearly identical activity function, the sum function and ReLU function are used in deep learning. Through linear transformation, information can be subtracted, added, etc. Essentially, neurons function as gating functions that pass the output to the next layer governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some of the neurons.
[0121] JPEG2025097253000002.jpg84144
[0122] JPEG2025097253000003.jpg46144
[0123] JPEG2025097253000004.jpg29132
[0124] In this case, neuron 610 is a single-layer perceptron. However, without departing from the scope of the present invention, any suitable neuron type or combination of neuron types can be used. It should also be noted that the weights of the activation function and / or the range of values of the output value(s) can be different in some embodiments without departing from the scope of the present invention.
[0125] A target, i.e., a "reward function", is often adopted. The reward function guides the search of the state space and tries to achieve the target (e.g., finding the most accurate answer to the user's query based on relevant indicators) by using both short-term and long-term rewards to explore intermediate transitions and steps. During training, various labeled data are supplied through the neural network 600. When a particular one succeeds, the weights of the inputs to the neurons are strengthened, while when a particular one fails, those weights are weakened. A cost function such as the mean squared error (MSE) or gradient descent can be used to make slightly incorrect predictions cost much less than greatly incorrect predictions. If the performance of the AI / ML model is not improved after a certain number of training iterations, the data scientist can change the reward function or correct the incorrect predictions, etc.
[0126] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "pop the hood" of the hidden layer of a neural network to check how much loss each node is bearing, and then give low weights to the nodes with a high error rate and vice versa, to update the weights to minimize the loss. That is, backpropagation enables the data scientist to repeatedly adjust the weights so as to minimize the difference between the actual output and the desired output.
[0127] The algorithm of backpropagation is mathematically based on the optimization theory. In supervised learning, training data with known outputs are passed through the neural network, and the error is calculated using the cost function from the known target outputs, which gives the error for backpropagation. The error is calculated at the output, and this error is converted into the correction of the network weights that minimizes the error.
[0128] JPEG2025097253000005.jpg71144
[0129] JPEG2025097253000006.jpg22144
[0130] JPEG2025097253000007.jpg123144
[0131] JPEG2025097253000008.jpg106116
[0132] JPEG2025097253000009.jpg83144
[0133] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., above 97% using an F2 or F4 threshold for detection, about 2000 epochs). This level of accuracy can be determined in some embodiments using an F1 score, an F2 score, an F4 score, or any other suitable technique that does not depart from the scope of the present invention. Once trained on the training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps prevent the AI / ML model from becoming "overfitted" such that it performs well on the training data but not on other data.
[0134] In some embodiments, it may not be known what accuracy levels an AI / ML model can achieve. Thus, when the accuracy of the AI / ML model begins to decline when analyzing evaluation data (i.e., the model performs well on training data but its performance is starting to degrade on evaluation data), the AI / ML model can undergo additional training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when the accuracy reaches a certain level or when the accuracy of the trained AI / ML model is better than that of an existing deployed AI / ML model. In certain embodiments, a set of trained AI / ML models can be used to accomplish a task. For example, one model can be trained to recognize images, another model can be trained to recognize text, and yet another model can be trained to recognize semantic and / or ontology-related aspects, etc.
[0135] In some embodiments, a transformer network such as SentenceTransformers™, a Python™ framework for state-of-the-art sentence, text, and image embedding, can be used. Such a transformer network learns associations of words and phrases with both high and low scores. This trains the AI / ML model to determine what is close to the input and what is not, respectively. Instead of using only word / phrase pairs, the transformer network may also use field length and field type.
[0136] In some embodiments, NLP technologies such as word2vec, BERT, GPT-3, ChatGPT, and other LLMs can be used to facilitate semantic understanding and provide more accurate and human-like answers as described above. Other technologies such as clustering algorithms can be used to find similarities between groups of elements. Clustering algorithms can include, but are not limited to, density-based algorithms, distribution-based algorithms, centroid-based algorithms, and hierarchical-based algorithms. Such as the K-means clustering algorithm, DBSCAN clustering algorithm, Gaussian mixture model (GMM) algorithm, and Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) algorithm. Such techniques can also be useful for classification.
[0137] Figure 7 is a flowchart showing a process 700 for training an AI / ML model(s) according to an embodiment of the present invention. In some embodiments, the AI / ML model(s) may be a generative AI model as described above. The neural network architecture of an AI / ML model typically includes multiple layers of neurons including an input layer, an output layer, and hidden layers. For example, refer to FIGS. 6A and 6B. The intermediate hidden layers process the input data and generate an intermediate representation of the input used for the generation of the output. These hidden layers can include various types of neurons such as convolutional neurons, recurrent neurons, and / or transformer neurons.
[0138] The training process begins at 710 by providing, with or without labels, an RPA workflow, PDDs, and / or other documents, screenshots, RPA workflow activities with annotations, etc. The AI / ML model is then trained at 720 over multiple epochs, and the results are reviewed at 730. Various types of AI / ML models can be used, but LLM and other generative AI models are typically trained using a process called "supervised learning" as described above. Supervised learning involves providing the model with a large dataset, which the model uses to learn the relationship between the input and the output. During the training process, the model adjusts the weights and biases of the neurons within the neural network to minimize the difference between the predicted output and the actual output within the training dataset.
[0139] One aspect of the model in some embodiments is the use of transfer learning. For example, transfer learning can utilize a pre-trained model such as ChatGPT that is fine-tuned for a specific task or domain at step 720. This allows the model to leverage the knowledge already learned from the pre-training phase and adapt it to a specific application through the training phase at step 720.
[0140] The pre-training phase involves training the model based on an initial set of training data that may be more general. In this phase, the model learns the relationships within the data. In the fine-tuning phase (e.g., in some embodiments, when a pre-trained model is used as the initial basis for the final model, in addition to or instead of the initial training phase, and performed during step 720), the pre-trained model is adapted to a specific task or domain by training the model with a smaller dataset specific to the task. For example, in some embodiments, the model may focus on a particular type(s) of data source. This can help the model more accurately identify data elements within it than a generative AI model pre-trained alone. Through fine-tuning, the model can learn nuances of the source such as specific vocabulary and syntax, specific graphical characteristics, specific data formats, etc., without requiring as much data as would be needed to train the model from scratch. By leveraging the knowledge learned in the pre-training phase, the fine-tuned model can achieve state-of-the-art performance on a specific task with relatively little additional training data.
[0141] If the AI / ML model does not meet the desired confidence threshold at 740, the training data is supplemented and / or the reward function is modified at 750 to help the AI / ML model better achieve its objective, and the process returns to step 720. If the AI / ML model meets the confidence threshold at 740, the AI / ML model is tested against the evaluation data at 760 to confirm that the AI / ML model generalizes well and does not overfit to the training data. The evaluation data includes information that the AI / ML model has not processed before. If the confidence threshold is met for the evaluation data at 770, the AI / ML model is deployed at 780. Otherwise, the process returns to step 750 and the AI / ML model is further trained.
[0142] FIG. 8 shows an AI / ML model 800 of the cognitive AI layer according to an embodiment of the present invention. In some embodiments, the model of the cognitive AI layer 800 can be trained using the process 700 of FIG. 7. The generative AI model 810 provides results to other cognitive AI layer models in a serial configuration 840, a parallel configuration 842, or a combination 844 of serial and parallel configurations. For example, the generative AI model can generate an annotated RPA workflow, generate PDD and / or other text, provide a semantic association between texts on a screen, logically group classes of RPA workflow activities, infer subsequent activities to add based on the context of the RPA workflow under development, convert an RPA workflow from one RPA vendor's format to another, generate automation from a PDD or business process description provided by a user, and so on. The CV model 820 and the OCR model also provide the detected graphical elements and recognized text, respectively, to the generative AI model 810 and other cognitive AI layer models within the configurations 840, 842, or 844.
[0143] Other cognitive AI models of the configurations 840, 842, or 844 use the outputs from the generative AI model 810, the CV model 820, and / or the OCR model 830 to provide the automatic annotation and technical specification generation described in detail herein. The cognitive AI layer 800 can facilitate the understanding of what code to generate for an RPA workflow, what text to generate for a PDD, and so on. The output (i.e., the result 850) from other cognitive AI layer models in the configurations 840, 842, or 844 can include an RPA workflow, a PDD, other documents, and the like.
[0144] Figure 9 is a screenshot 900 showing an annotated RPA workflow according to an embodiment of the present invention. The project name is listed in field 910, and the text box 920 contains an explanation of the project generated by the generative AI model of the cognitive AI layer. For example, the generative AI model can be an LLM trained in a fine-tuning phase to understand the nature of RPA workflow activities and the interrelationships between them. In the example of FIG. 9, the generative AI model can understand that there are activities to create a new row in the vendor table, onboard a new row to Salesforce®, notify an individual's team via Slack®, and send an email confirmation to those individuals via Outlook®. Much more complex workflows are possible and can be annotated by the cognitive AI layer.
[0145] Figure 10 is a flowchart showing a process 1000 that provides automatic annotation and technical specification generation for an RPA workflow using an existing AI for RPA automation according to an embodiment of the present invention. In some embodiments, process 1000 of FIG. 10 may be part of process 1100 of FIG. 11. For example, in some embodiments, steps 1010 to 1040 of FIG. 10 may correspond to steps 1130, 1140 of FIG. 11.
[0146] In some embodiments, the process begins by generating a PDD that describes the process to be automated using a generative AI model of the cognitive AI layer or another generative AI model (e.g., one trained separately from and executed separately from the cognitive AI layer). The PDD can be generated by the generative AI model using various sources. For example, the PDD can be generated using process logs (e.g., from Salesforce® or another application), manual traces, diagrams, etc.
[0147] The RPA workflow code or PDD is provided as an input to the cognitive AI layer at 1020. For example, the XAML file of the RPA workflow can be provided to the initial AI / ML model of the cognitive AI layer such as an LLM. In some embodiments, screenshots (s) may be provided so that a CV and / or OCR model can identify images and / or text therein. The cognitive AI layer then processes the RPA workflow code or PDD (and other input information if provided) at 1030 and provides at 1040 the annotation of the RPA workflow, or the code of the RPA workflow including the annotation. For example, the generative AI model of the cognitive AI layer can be trained to annotate the RPA workflow and / or generate PDD from the content of the RPA workflow. In some embodiments, the RPA workflow annotation(s) can provide an explanation of the entire process of the RPA workflow, explain what each activity within the RPA workflow is doing, or both. In certain embodiments, the annotation can include an explanation of the changes between the current version of the RPA workflow and the previous version(s) of the RPA workflow. For example, the explanation of an activity can be divided by what that activity was doing in the previous version (if it exists) and what has been changed for the current version. In embodiments where an annotated RPA workflow is generated, the annotated RPA workflow can be displayed in the RPA designer application at 1050.
[0148] Figure 11 is a flowchart showing a process 1100 that provides automatic annotation and technical specification generation to an RPA workflow using AI when a developer creates an RPA workflow according to an embodiment of the present invention. The process begins at 1110 by monitoring RPA workflow development in an RPA designer application. This monitoring can be performed by the RPA designer application itself or by another process (e.g., a task mining process, a listener process, etc.). Next, the RPA designer application or other process determines at 1120 that an activity of the RPA workflow has been added or changed. This determination can be made when an activity is completed and the user proceeds to the next activity, when some change is made, periodically (e.g., every minute, every 10 minutes, etc.), when the user saves the RPA workflow, when the user clicks a button to annotate the RPA workflow in the RPA designer application, and so on.
[0149] The RPA workflow code is provided to and processed by a cognitive AI layer at 1130 to provide an annotated RPA workflow. For example, the XAML file of the RPA workflow can be provided to an initial AI / ML model of the cognitive AI layer such as an LLM. In some embodiments, screenshots (s) may be provided so that CV and / or OCR models can identify images and / or text therein. The RPA designer application uses this output at 1140 to update the RPA workflow and display the annotated version to the user. In certain embodiments, the annotation includes an explanation of the changes between the current version of the RPA workflow and the previous version(s) of the RPA workflow. If the user is still working on the RPA workflow, the process returns to step 1110. Otherwise, the process ends.
[0150] Figure 12 is a flowchart showing a process of creating an RPA workflow using a PDD or another natural language description of a business process as input. The process begins at 1210 by providing a description of the PDD or other business process to the cognitive AI layer. For example, a user may input a natural language description of a desired business process, or use speech-to-text conversion to indicate what he or she wishes to automate. Next, the cognitive AI layer processes the description of the PDD or other business process at 1220 and generates an annotated RPA workflow at 1230. For example, the LLM of the cognitive AI layer can be trained in a fine-tuning phase to understand the nature of the RPA workflow activities and the interrelationships between them. In some embodiments, the cognitive AI layer or the RPA designer application can generate runtime automation at 1240 using an RPA workflow that can be deployed to a production environment. In certain embodiments, documents such as PDDs, compliance documents, audit documents, etc. can be created at 1250 if they were not provided as input to the cognitive AI layer.
[0151] The process steps executed in FIGS. 7 and 10 - 12 may be executed by a computer program that encodes instructions to a processor(s) to execute at least a portion of the process(es) described in FIGS. 7 and 10 - 12, according to embodiments of the present invention. The computer program may be stored on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, a tape, and / or any other such medium or combination of media used to store data. The computer program may include encoded instructions for controlling a processor(s) of a computing system (e.g., the processor(s) 510 of the computing system 500 of FIG. 5) to implement all or a portion of the process steps described in FIGS. 7 and 10 - 12, which may also be stored on a computer-readable medium.
[0152] A computer program can be implemented in hardware, software, or a hybrid implementation. The computer program can be composed of modules that perform operable communication with each other and is designed to send information or instructions to a display. The computer program can be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0153] It will be readily understood that the components of the various embodiments of the present invention may be arranged and designed in a variety of different configurations as generally described and illustrated herein. Accordingly, the detailed description of the embodiments of the present invention as represented in the accompanying figures is not intended to limit the scope of the invention as claimed, but rather represents only selected embodiments of the present invention.
[0154] The features, structures, or characteristics of the present invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments", "some embodiments", or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. Accordingly, the appearances of "in certain embodiments", "in some embodiments", "in other embodiments", or similar language throughout this specification are not necessarily all referring to the same group of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0155] References throughout this specification to features, advantages, or similar language do not imply that all of the features and advantages that can be realized in the present invention should be in any single embodiment of the invention or in any embodiment of the invention. Rather, language referring to features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, discussions of features and advantages throughout this specification, as well as similar language, can refer to the same embodiment, but not necessarily.
[0156] Furthermore, the described features, advantages, and characteristics of the present invention can be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the present invention can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments but not in all embodiments of the present invention.
[0157] Those of ordinary skill in the art will readily understand that the present invention, as described above, can be practiced using steps in a different order and / or using hardware elements configured differently than those disclosed. Thus, while the present invention has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations will become apparent while remaining within the spirit and scope of the invention. Accordingly, the appended claims should be referred to in order to determine the scope of the present invention.
Claims
1. 1. A non-transitory computer readable medium having stored thereon a computer program for providing automated annotation in a robotic process automation (RPA) workflow, the computer program comprising: Providing the RPA workflow or process definition document (PDD) code to a cognitive artificial intelligence (AI) layer; processing the RPA workflow or the code of the PDD by the cognitive AI layer; A non-transitory computer readable medium configured to provide, by a generative AI layer, as output, annotations for the RPA workflow or RPA workflow code including the annotations.
2. The computer program product further comprises:
2. The non-transitory computer-readable medium of claim 1 configured to generate the PDD describing a process to be automated using a generative AI model of the cognitive AI layer or another generative AI model.
3. 2. The non-transitory computer-readable medium of claim 1 , wherein the annotations include a description of changes between a current version of the RPA workflow and one or more previous versions of the RPA workflow.
4. The computer program further comprises:
10. The non-transitory computer-readable medium of claim 1 , configured to display the annotated RPA workflow in a user interface of an RPA designer application.
5. 2. The non-transitory computer-readable medium of claim 1, wherein the generative AI layer includes a large-scale language model (LLM) that is fine-tuned during training to understand the nature and interrelationships between RPA workflow activities.
6. The computer program further comprises:
2. The non-transitory computer readable medium of claim 1 , configured to provide one or more screenshots comprising one or more visual representations of at least a portion of the RPA workflow to a computer vision (CV) model and / or an optical character recognition (OCR) model of the cognitive AI layer to identify text and / or images therein.
7. 2. The non-transitory computer-readable medium of claim 1, wherein the annotations provide a description of an overall process of the RPA workflow, a description of what each activity in the RPA workflow is doing, or both.
8. The computer program further comprises: monitoring the development of said RPA workflow in an RPA designer application; determining that one or more activities have been added to and / or modified in the RPA workflow; 2. The non-transitory computer-readable medium of claim 1 , configured to provide the code of the RPA workflow to the cognitive AI layer in response to a determination that the one or more activities have been added to and / or modified in the RPA workflow.
9. 9. The non-transitory computer readable medium of claim 8, wherein the determining that the one or more activities have been added to and / or changed in the RPA workflow includes determining that an activity in the RPA workflow has been completed and a user has moved on to a next activity, periodically checking for changes to the RPA workflow, determining that the user has saved the RPA workflow, and determining that the user has clicked a button in an RPA designer application to annotate the RPA workflow.
10. The computer program further comprises:
9. The non-transitory computer-readable medium of claim 8, configured to repeat the steps of claim 8 after the generative AI layer provides as output the annotations of the RPA workflow or the RPA workflow code including the annotations.
11. The cognitive AI layer is 2. The non-transitory computer readable medium of claim 1, comprising a generative AI model configured to generate annotated RPA workflows, provide semantic associations between on-screen text, logically group classes of RPA workflow activities, infer subsequent activities to add based on a context of the RPA workflow, convert an RPA workflow from one RPA vendor's format to another RPA vendor's format, or any combination thereof.
12. a memory storing computer program instructions for providing automated annotation to a robotic process automation (RPA) workflow; and at least one processor configured to execute the computer program instructions, the computer program instructions comprising: providing the RPA workflow or process definition document (PDD) code to a cognitive AI layer comprising a generative artificial intelligence (AI) model configured to generate annotated RPA workflows, provide semantic associations between on-screen text, logically group classes of RPA workflow activities, infer subsequent activities to add based on the context of the RPA workflow, translate the RPA workflow from one RPA vendor's format to another RPA vendor's format, or any combination thereof; processing the RPA workflow or the code of the PDD by the cognitive AI layer; configured to provide, by a generation AI layer, as an output, annotations of the RPA workflow or RPA workflow code including the annotations; The annotations provide a description of the overall process of the RPA workflow, a description of what each activity in the RPA workflow is doing, or both.
13. 13. The one or more computing systems of claim 12, wherein the annotations include a description of changes between a current version of the RPA workflow and one or more previous versions of the RPA workflow.
14. The computer program instructions further include causing the at least one processor to:
13. The one or more computing systems of claim 12, configured to display the annotated RPA workflow in a user interface of an RPA designer application.
15. 13. The one or more computing systems of claim 12, wherein the generative AI layer includes large language models (LLMs) fine-tuned during training to understand the nature and interrelationships between RPA workflow activities.
16. The computer program instructions further include causing the at least one processor to: configured to provide one or more screenshots comprising one or more visual representations of at least a portion of the RPA workflow to a computer vision (CV) model and / or an optical character recognition (OCR) model of the cognitive AI layer to identify text and / or images therein; 13. The one or more computing systems of claim 12, wherein the CV model and / or the OCR model are configured to provide output to a generative AI model of the cognitive AI layer.
17. The computer program instructions further include causing the at least one processor to: monitoring the development of said RPA workflow in an RPA designer application; determining that one or more activities have been added to and / or modified in the RPA workflow; 13. The one or more computing systems of claim 12, configured to provide the code of the RPA workflow to the cognitive AI layer in response to a determination that the one or more activities have been added to and / or changed in the RPA workflow.
18. 18. The one or more computing systems of claim 17, wherein the determining that the one or more activities have been added to and / or changed in the RPA workflow includes determining that an activity in the RPA workflow has been completed and a user has moved on to a next activity, periodically checking for changes to the RPA workflow, determining that the user has saved the RPA workflow, and determining that the user has clicked a button in an RPA designer application to annotate the RPA workflow.
19. 1. A computer-implemented method for providing automated annotation for a robotic process automation (RPA) workflow, comprising: providing code of the RPA workflow or process definition document (PDD) to a cognitive artificial intelligence (AI) layer by an RPA designer application executed on a computing system, the cognitive AI layer being configured to process the code of the RPA workflow and / or the PDD; receiving, by the RPA designer application, annotations for the RPA workflow or RPA workflow code including the annotations from a generation AI layer; displaying, by the RPA designer application, the RPA workflow with annotations; The computer-implemented method, wherein the annotations provide a description of the overall process of the RPA workflow, a description of what each activity of the RPA workflow does, a description of changes between a current version of the RPA workflow and one or more previous versions of the RPA workflow, or any combination thereof.
20. 20. The computer-implemented method of claim 19, wherein the generative AI layer includes a large-scale language model (LLM) that is fine-tuned during training to understand the nature and interrelationships between RPA workflow activities.
21. moreover, providing, by the RPA designer application, one or more screenshots comprising one or more visual representations of at least a portion of the RPA workflow to a computer vision (CV) model and / or an optical character recognition (OCR) model of the cognitive AI layer to identify text and / or images therein; 20. The computer-implemented method of claim 19, wherein the CV model and / or the OCR model are configured to provide output to a generative AI model of the cognitive AI layer.
22. moreover, monitoring development of the RPA workflow with the RPA designer application; determining that one or more activities have been added to and / or modified in the RPA workflow by the RPA designer application; providing the code of the RPA workflow to the cognitive AI layer in response to a determination by the RPA designer application that the one or more activities have been added to and / or modified in the RPA workflow; 20. The computer-implemented method of claim 19, wherein determining that the one or more activities have been added to and / or modified in the RPA workflow includes determining that an activity in the RPA workflow has been completed and a user has moved on to a next activity, periodically checking for changes to the RPA workflow, determining that the user has saved the RPA workflow, and determining that the user has clicked a button in an RPA designer application to annotate the RPA workflow.
23. moreover, 23. The computer-implemented method of claim 22, comprising repeating the steps of claim 22 after the generative AI layer provides as output the annotations of the RPA workflow or the RPA workflow code including the annotations, by the RPA designer application.
24. The cognitive AI layer is 20. The computer-implemented method of claim 19, comprising a generative AI model configured to generate annotated RPA workflows, provide semantic associations between on-screen text, logically group classes of RPA workflow activities, infer subsequent activities to add based on a context of the RPA workflow, convert an RPA workflow from one RPA vendor's format to another RPA vendor's format, or any combination thereof.
Citation Information
Cited By
Browser AI integration system
JP7891198B1