Generative artificial intelligence that integrates detection and automation of user interface elements including context awareness
The interface engine integrates RPA and generative AI to address language-specific limitations in UI detection, ensuring consistent automation across diverse UIs.
Patent Information
- Application Number
- JP2025003163
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2025-01-09
- Publication Date
- 2025-10-24
AI Technical Summary
Specialized computer vision algorithms for user interface detection are language-specific, failing to recognize elements on screens localized in non-English languages.
An interface engine leveraging robotic process automation (RPA) and generative artificial intelligence (AI) models to detect and automate user interface elements, incorporating context awareness, enabling language-independent UI automation and content awareness.
Enables consistent detection and automation of user interface elements across different languages and attributes, improving UI ecosystem consistency and automation efficiency.
Smart Images

Figure 2025161728000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to automation, and more specifically to leveraging automation and robotic process automation (RPA) integrated with generative artificial intelligence (AI) models to enable the detection and automation of user interface elements, including context awareness. [Background technology]
[0002] Traditionally, technologies such as specialized computer vision algorithms are trained on large volumes of computer screen images to consistently identify elements / objects on the screen. The problem with these specialized computer vision algorithm approaches is that they are specific to the images in their training set (e.g., they are highly language-specific). For example, if a user interface (UI) is localized into Japanese, Arabic, or other non-English languages, training a specialized computer vision algorithm on English screens will not return the same results.
[0003] Therefore, there is a need to provide detection and automation of user interface elements that includes context awareness. Summary of the Invention
[0004] According to one or more embodiments, a method is provided. The method is performed by an interface engine implemented as a computer program within a computing environment. The interface engine performs computer activity detection and automation. The method includes recording computer activity across one or more user interfaces (UIs) and automatically processing the computer activity using at least one generative AI model to extract patterns. The patterns include portions of the computer activity that are similar to or tolerant to change. The method includes determining a plurality of existing automations according to the patterns.
[0005] According to any of the one or more embodiments or method embodiments herein, the interface engine may be implemented as an apparatus, a system, and a computer program product. [Brief explanation of the drawings]
[0006] So that the advantages of particular embodiments of this invention may be readily understood, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments which are illustrated in the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope, but the invention will be described and explained with additional specificity and detail through the use of the following accompanying drawings, in which:
[0007] [Figure 1] FIG. 1 illustrates an architecture diagram illustrating an automation system according to one or more embodiments.
[0008] [Figure 2] FIG. 1 illustrates an architecture diagram illustrating a robotic process automation (RPA) system according to one or more embodiments.
[0009] [Figure 3]FIG. 1 illustrates an architecture diagram illustrating a deployed RPA system according to one or more embodiments.
[0010] [Figure 4] 1 illustrates an architecture diagram showing the relationships between designers, activities, and drivers, according to one or more embodiments.
[0011] [Figure 5] 1 shows an architectural diagram illustrating a computing system according to one or more embodiments.
[0012] [Figure 6] 1 illustrates an example of a neural network trained to recognize graphical elements in an image according to one or more embodiments.
[0013] [Figure 7] 1 illustrates an example of a neuron according to one or more embodiments.
[0014] [Figure 8] 1 shows a flowchart illustrating a process for training artificial intelligence and / or machine learning (AI / ML) model(s) according to one or more embodiments.
[0015] [Figure 9] 1 illustrates a process according to one or more embodiments.
[0016] [Figure 10] 1 illustrates a process according to one or more embodiments.
[0017] Unless otherwise noted, like reference characters denote corresponding features consistently throughout the accompanying drawings. DETAILED DESCRIPTION OF THE INVENTION
[0018] Detailed Description of the Embodiments According to one or more embodiments, an interface engine leverages automation and robotic process automation (RPA) and integrates with generative artificial intelligence (AI) models to enable the detection and automation of user interface elements, including context awareness. The interface engine includes software implemented as processor-executable code stored in memory. The interface engine may be implemented by a combination of the processor-executable code and hardware described herein. The interface engine is necessarily rooted in the operation of at least one processor to improve upon prior art to enable user interface (UI) and ecosystem consistency, regardless of language and other attributes within the UI.
[0019] In operation, according to one or more embodiments, the interface engine records text, events, and screenshots during computer activity (e.g., by a user, automation, or RPA). The interface engine then extracts pattern-generating AI models to find controls on individual screens and group screens that are similar and / or resistant to changes in color, theme, language, resolution, screen size, browser, and other attributes. Generative AI models may include any generative pre-trained transformer (GPT), AI agent (e.g., AutoGPT and UiPath Autopilot™), large-scale language model (LLM) gateway, and other models. Additionally, the interface engine and its generative AI models process the patterns to determine whether existing or new automation and / or RPA (e.g., multiple automations) should implement the computer activity.
[0020] Thus, the interface engine provides a mechanism for interacting with automation and RPA that enables language-independent UI automation and discovery, and content awareness as technical effects, benefits, and advantages.
[0021] FIG. 1 is an architectural diagram illustrating a hyperautomation system 100 according to one or more embodiments. As used herein, “hyperautomation” refers to an automation system that brings together process automation components, integrated tools, and technologies that amplify the ability to automate work. For example, in some embodiments, RPA is used at the core of the hyperautomation system, and in certain embodiments, automation capabilities may be extended with artificial intelligence and / or machine learning (AI / ML), process mining, analytics, and / or other advanced tools. As the hyperautomation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work may be automated, and computing systems within an organization, e.g., both those used by individuals and those operating autonomously, may all be engaged as participants in the hyperautomation process. The hyperautomation system of some embodiments enables users and organizations to efficiently and effectively discover, understand, and scale automation.
[0022] The hyperautomation system 100 includes user computing systems such as a desktop computer 102, a tablet 104, and a smartphone 106. However, any desired computing system, including, but not limited to, a smartwatch, a laptop computer, a server, an Internet of Things (IoT) device, etc., may be used without departing from the scope of one or more embodiments described herein. Also, while three user computing systems are shown in FIG. 1 , any suitable number of computing systems may be used without departing from the scope of one or more embodiments described herein. For example, in some embodiments, tens, hundreds, thousands, or millions of computing systems may be used. The user computing systems may be actively used by a user or may run automatically without much or any user input.
[0023] Each computing system 102, 104, 106 has a respective automation process(es) 110, 112, 114 executing thereon. The automation process(es) 102, 104, 106 may include, without limitation, an RPA robot, part of an operating system, downloadable application(s) for the respective computing system, any other suitable software and / or hardware, or any combination thereof, without departing from the scope of one or more embodiments described herein. In some embodiments, one or more process(es) 110, 112, 114 may be a listener. The listener may be an RPA robot, part of an operating system, a downloadable application for the respective computing system, or any other software and / or hardware without departing from the scope of one or more embodiments described herein. Indeed, in some embodiments, the logic of the listener(s) is implemented partially or fully via physical hardware.
[0024] The listeners monitor and record data related to user interactions with their respective computing systems and / or the operation of unattended computing systems and transmit the data over a network (e.g., a local area network (LAN), a mobile communications network, a satellite communications network, the Internet, any combination thereof, etc.) to the core hyperautomation system 120. The data may include, but is not limited to, which buttons were clicked, where the mouse was moved, text entered into a field, when one window was minimized and another was opened, the application associated with the window, etc. In certain embodiments, the data from the listeners may be transmitted periodically as part of a heartbeat message. In some embodiments, the data may be transmitted to the core hyperautomation system 120 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listeners and store it in a database, such as database 140.
[0025] An automation process may execute logic developed in a workflow during design time. In the case of RPA, a workflow may include a set of steps, defined herein as "activities," that are executed in sequence or some other logical flow. Each activity may include an action such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.
[0026] Long-running workflows for RPA, in some embodiments, are a master project that supports service orchestration, human intervention, and long-running transactions in unattended environments. See U.S. Patent No. 10,860,905, the entire contents of which are incorporated by reference. Human intervention occurs when a particular process requires human input for exception handling, approval, or validation before proceeding to the next step in the activity. In this case, process execution is paused, freeing up the RPA robot until the human task is completed.
[0027] Long-running workflows may support workflow fragmentation through persistence activities and may be combined with call process and non-user interaction activities to orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or numerous computing systems may participate in the execution of a long-running workflow's logic. Long-running workflows may execute in sessions to facilitate rapid execution. In some embodiments, long-running workflows may orchestrate background processes that execute application programming interface (API) calls and may include activities that execute in the long-running workflow session. These activities may, in some embodiments, be invoked by a call process activity. Processes with user interaction activities that execute in a user session may be invoked by initiating a job from a conductor activity (conductors are described in more detail later in this specification). In some embodiments, users may interact through tasks that require the completion of a form in the conductor. An activity may be included that causes the RPA robot to wait for a form task to complete and then resume the long-running workflow.
[0028] One or more automation process(es) 110, 112, 114 are in communication with the core hyperautomation system 120. In some embodiments, the core hyperautomation system 120 may execute a conductor application on one or more servers, such as server 130. While one server 130 is shown for illustrative purposes, multiple or numerous servers in close proximity to one another or in a distributed architecture may be employed without departing from the scope of one or more embodiments described herein. For example, one or more servers may be provided for conductor functionality, AI / ML model serving, certification, governance, and / or any other suitable functionality without departing from the scope of one or more embodiments described herein. In some embodiments, the core hyperautomation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, or the like. In certain embodiments, the core hyperautomation system 120 may host multiple software-based servers on one or more computing systems, such as server 130. In some embodiments, one or more servers of the core hyperautomation system 120, such as server 130, may be implemented via one or more virtual machines (VMs).
[0029] In some embodiments, one or more automation process(es) 110, 112, 114 may invoke one or more AI / ML models 132 deployed on or accessible by the core hyperautomation system 120. The AI / ML models 132 may be trained for any suitable purpose, as discussed in more detail herein, without departing from the scope of one or more embodiments described herein. Two or more AI / ML models 132 may be chained (e.g., serially, in parallel, or a combination thereof) in some embodiments so that they collectively provide collaborative output(s). The AI / ML models 132 may perform or assist in computer vision (CV), optical character recognition (OCR), document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, any combination thereof, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of one or more embodiments described herein. Using multiple AI / ML models, for example, the system may develop a complete picture of what is happening on a given computing system. For example, one AI / ML model may perform OCR, another may detect buttons, another may compare sequences, etc. Patterns may be determined by the AI / ML models individually or collectively by multiple AI / ML models. In particular embodiments, one or more AI / ML models are deployed locally on at least one computing system 102, 104, 106.
[0030] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that runs on data, and the AI / ML model itself may be, for example, a deep learning neural network (DLNN) of artificial “neurons” trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions, such as statistical modeling (e.g., hidden Markov models (HMMs)), and may utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of prior hidden states, etc.) to perform desired functions.
[0031] The hyperautomation system 100, in some embodiments, may provide four main groups of functions: (1) discovery, (2) automation build, (3) management, and (4) engagement. Automation (e.g., running on a user computing system, server, etc.) may, in some embodiments, be performed by software robots such as RPA robots. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with users to assist them with tasks (e.g., via UiPath Assistant™). Unattended robots operate independently of users and may potentially run in the background without the user's knowledge. Test robots are unattended robots that run test cases against an application or RPA workflow. Test robots, in some embodiments, may run in parallel on multiple computing systems.
[0032] A discovery function may discover various business process automation opportunities and provide automated recommendations. Such functionality may be implemented by one or more servers, such as server 130. In some embodiments, the discovery function may include providing an automation hub, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath Automation Hub™) may provide a mechanism for managing automation rollouts with visibility and control. Automation ideas may be crowdsourced from employees, for example, via a submission form. Calculations of the feasibility and return on investment (ROI) for automating these ideas may be provided, documentation for future automations may be collected, and collaboration may be provided to expedite automation discovery and construction.
[0033] Process mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Center™) refers to the process of collecting and analyzing data from applications (e.g., enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how they can be effectively automated, and the impact of automation. This data may be obtained, for example, by listeners from user computing systems 102, 104, 106 and processed by a server, such as server 130. In some embodiments, one or more AI / ML models 132 may be employed for this purpose. This information may be exported to an automation hub to speed implementation and avoid manual information transfer. The goal of process mining may be to increase business value by automating processes within an organization. Some example goals of process mining include, but are not limited to, increased profits, improved customer satisfaction, regulatory and / or contractual compliance, improved employee efficiency, etc.
[0034] Task mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Center™) identifies and aggregates workflows (e.g., employee workflows) and then applies AI to uncover patterns and variations in routine tasks and score such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to uncover repetitive task patterns within the data. Repetitive tasks ripe for automation may then be identified. This information may initially be provided by a listener and, in some embodiments, may be analyzed on a server of the core hyper-automation system 120, such as server 130. Findings from task mining (e.g., extensive application markup language (XAML) process data) may be exported to process documentation or a designer application, such as UiPath Studio™, to more quickly create and deploy automations.
[0035] Task mining in some embodiments may include taking screenshots with user actions (e.g., mouse click locations, keyboard input, application windows and graphical elements with which the user was interacting, timestamps for the interactions, etc.), collecting statistical data (e.g., performance time, number of actions, text input, etc.), editing and annotating screenshots, specifying the types of actions to be recorded, etc.
[0036] Task capture (via UiPath Automation Cloud™ and / or UiPath AI Center™) automatically documents attended processes as users work on them, or provides a framework for unattended processes. Such documentation may include process definition documents (PDDs), skeleton workflows, capturing actions for each part of the process, recording user actions and automatically generating comprehensive workflow diagrams with details about each step, tasks desired to be automated, in formats such as Microsoft Word documents, XAML files, etc. Configurable workflows, in some embodiments, can be exported directly to designer applications such as UiPath Studio™. Task capture may simplify the requirements gathering process for both subject matter experts describing the process and Center of Excellence (CoE) members delivering production-grade automation.
[0037] Building automations may be accomplished through a designer application (such as UiPath Studio™, UiPath StudioX™, or UiPath Web™). For example, an RPA developer at the PA development facility 150 may use the RPA designer application 154 on the computing system 152 to build and test automations for various applications and environments, such as web, mobile, SAP®, and virtual desktops. API integration may be provided for various applications, technologies, and platforms. Predefined activities, drag-and-drop modeling, and a workflow recorder may facilitate automation with minimal coding. Document understanding capabilities may be provided through drag-and-drop AI skills for data extraction and interpretation that invoke one or more AI / ML models 132. Such automations can handle virtually any document type and format, including tables, checkboxes, signatures, and handwriting. When data is validated or exceptions are handled, this information may be used to retrain the respective AI / ML models, improving their accuracy over time.
[0038] Integration services allow developers to seamlessly combine user interface (UI) automation and API automation, for example. Automations can be built that require APIs or span both API and non-API applications and systems. A repository (e.g., UiPath Object Repository™) or marketplace (e.g., UiPath Marketplace™) for pre-built RPA and AI templates and solutions can be provided to enable developers to more quickly automate a wide variety of processes. Thus, when building automations, hyperautomation system 100 can provide a user interface, development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. Hyperautomation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for hyperautomation system 100.
[0039] In some embodiments, components of hyper-automation system 100, such as designer application(s) and / or external rules engines, provide support for managing and enforcing governance policies to control various functions provided by hyper-automation system 100. Governance is the ability for an organization to put policies in place to prevent users from developing automation (e.g., RPA robots) that can perform actions that could harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the US Health Insurance Portability and Accountability Act (HIPAA), third-party application terms of use, etc. Because developers may otherwise create automations that violate privacy laws, terms of use, etc. during the execution of their automations, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. This may provide an additional level of security and compliance to the automation process development pipeline in some embodiments by preventing developers from taking dependencies on unauthorized software libraries that may pose security risks or operate in a manner that violates policies, regulations, privacy laws, and / or privacy policies. See U.S. Non-Provisional Patent Application No. 16 / 924,499, the entire contents of which are incorporated by reference.
[0040] Management functions may provide management, deployment, and optimization of automation across an organization. Management functions may, in some embodiments, include orchestration, test management, AI capabilities, and / or insights. Management functions of hyperautomation system 100 may also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. Management functions of hyperautomation system 100 may include, but are not limited to, facilitating provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots, among others.
[0041] Conductor applications, such as UiPath Orchestrator™ (which may be offered in some embodiments as part of UiPath Automation Cloud™, or may be offered on-premise, on a VM, in a private or public cloud, on a Linux™ VM, or as a cloud-native single-container suite via the UiPath Automation Suite™), provide orchestration capabilities to deploy, monitor, optimize, scale, and ensure the security of RPA robot deployments. Test suites (e.g., the UiPath Test Suite™) may provide test management to monitor the quality of deployed automation. Test suites may facilitate test planning and execution, requirements fulfillment, and defect traceability. Test suites may include comprehensive test reports.
[0042] Analytics software (e.g., UiPath Insights™) can track, measure, and manage the performance of deployed automation. Analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes for the organization. Analytics software can present results in a dashboard format that is more easily understood by human users.
[0043] Data services (e.g., UiPath Data Service™), stored in database 140, for example, can bring data to a single, scalable, and secure location with a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage to automation while ensuring seamless access, enterprise-grade security, and scalability of data. AI capabilities may be provided by an AI center (e.g., UiPath AI Center™), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may make such capabilities accessible to non-data scientists. Deployed automation (e.g., RPA robots) may invoke AI / ML models from the AI center, such as AI / ML model 132. The performance of AI / ML models may be monitored and trained and improved using human-verified data, such as provided by data review center 160. Human reviewers may provide labeled data to core hyperautomation system 120 via review application 152 on computing system 154. For example, a human reviewer may verify that the predictions made by the AI / ML model 132 are accurate or may otherwise provide corrections. This dynamic input may then be saved as training data for retraining the AI / ML model 132 and stored in a database, such as database 140. The AI center may then schedule and execute training jobs to train a new version of the AI / ML model using the training data. Both positive and negative examples may be stored and used to retrain the AI / ML model 132.
[0044] Engagement capabilities engage humans and automation as one team for seamless collaboration around desired processes. Low-code applications can be built (e.g., via UiPath Apps™) to connect browser tabs with legacy software, even those lacking APIs in some embodiments. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. Applications can be connected to one automation or multiple automations.
[0045] Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for handing off processes from automation to humans and vice versa. Humans can provide approvals or escalations, handle exceptions, etc. Automation can then perform the automated functions of a given workflow.
[0046] A local assistant may be provided as a launch pad for users to launch automations (e.g., UiPath Assistant™). This functionality may be provided, for example, in a tray provided by the operating system and may allow users to interact with RPA robots and RPA robot-powered applications on their computing system. The interface may list automations approved for a given user and allow the user to run them. These may include ready-to-use automations from an automation marketplace, an automation hub's internal automation store, etc. When automations are running, they may run as local instances in parallel with other processes on the computing system so that the user can use the computing system while the automation performs its actions. In certain embodiments, the assistant is integrated with task capture functionality so that users can document their soon-to-be-automated processes from the assistant's launch pad.
[0047] Chatbots (e.g., UiPath Chatbots™), social messaging applications, and / or voice commands may enable users to execute automations, simplifying access to information, tools, and resources needed to interact with customers or perform other activities. Human-to-human conversations can be automated as easily as other processes. Triggered RPA robots activated in this way could potentially perform actions such as checking order status or posting data to a CRM using plain language commands.
[0048] End-to-end measurement and governance of automation programs at any scale may be provided by the hyper-automation system 100 in some embodiments. Accordingly, analytics (e.g., via UiPath Insights™) may be employed to understand automation performance. Data modeling and analytics using any combination of available business metrics and operational insights may be used for various automation processes. Custom-designed and pre-built dashboards may visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for automations, perform telemetry monitoring on user computing systems, detect errors and anomalies, and debug automations. An automation management console (e.g., UiPath Automation Ops™) may be provided to manage automations throughout their lifecycle. Organizations may govern how automations are built, what users can do with them, and which automations users have access to.
[0049] The hyper-automation system 100, in some embodiments, provides an iterative platform where processes can be discovered, automations can be built, tested, and deployed, performance can be measured, automation usage can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the process itself can be repeated, thereby facilitating a more robust and effective set of automations.
[0050] FIG. 2 is an architecture diagram illustrating an RPA system 200 according to one or more embodiments. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 can facilitate the development of automation projects, which are graphical representations of business processes. Simply put, the designer 210 facilitates the development and deployment (indicated by arrow 211) of workflows and robots. In some embodiments, the designer 210 can be an application running on a user's desktop, an application running remotely in a VM, a web application, or the like.
[0051] An automation project enables rule-based process automation by giving developers control over the order of execution and relationships between a custom set of steps developed in a workflow, defined herein as "activities" as described above. One commercial example of an embodiment of the designer 210 is UiPath Studio™. Each activity may include an action such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.
[0052] Some types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. Sequences may be particularly well-suited for linear processes, allowing the flow of one activity from another without cluttering the workflow. Flowcharts may be particularly well-suited for more complex business logic, allowing for the integration of decisions and the connection of activities in more diverse ways through multiple branching logic operators. FSMs may be particularly well-suited for large workflows. FSMs may use a finite number of states during their execution that are triggered by conditions (i.e., transitions) or activities. Global exception handlers may be particularly well-suited for determining workflow behavior when an execution error is encountered or for debugging the process.
[0053] Once a workflow is developed in designer 210, the execution of the business process is orchestrated by conductor 220, which coordinates one or more robots 230 that execute the workflow developed in designer 210. One commercial example of an embodiment of conductor 220 is UiPath Orchestrator™. Conductor 220 facilitates the management of the creation, monitoring, and deployment of resources in an environment. Conductor 220 may act as an integration point with third-party solutions and applications. Accordingly, in some embodiments, conductor 220 may be part of the core hyper-automation system 120 of FIG. 1 .
[0054] The conductor 220 may manage all robots 230, connecting and executing them from a centralized point (indicated by arrow 231). Types of robots 230 that may be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non-production robots (similar to attended robots 232 but used for development and testing purposes). Attended robots 232 are triggered by user events and operate in parallel with humans on the same computing system. Attended robots 232 may be used with the conductor 220 for centralized process deployment and logging media. Attended robots 232 may assist human users in accomplishing various tasks and may be triggered by user events. In some embodiments, processes cannot be initiated from the conductor 220 on this type of robot, and / or they cannot be run under a locked screen. In certain embodiments, the attended robot 232 can only be launched from the robot tray or from a command prompt. The attended robot 232 preferably operates under human supervision in some embodiments.
[0055] Unattended robots 234 operate unattended in virtual environments and can automate many processes. Unattended robots 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can be performed in designer 210 in some embodiments. Both attended and unattended robots can automate a variety of systems and applications (depicted by dashed box 290), including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those produced by SAP®, Salesforce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0056] Conductor 220 may have various capabilities (indicated by arrow 232), including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining connections between robots 230 and conductor 220 (e.g., web applications). Deployment may include ensuring the correct delivery of package versions to robots 230 assigned to perform. Configuration may include maintaining and delivering robot environment and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include tracking robot-specific data and maintaining user permissions. Logging may include storing and indexing logs in a database (e.g., a Structured Query Language (SQL) database or a NoSQL database) and / or another storage mechanism (e.g., ElasticSearch®, which provides the ability to store large data sets and quickly execute queries). Conductor 220 may provide interconnectivity by operating as a centralized point of communication for third-party solutions and / or applications.
[0057] Robots 230 are execution agents that implement workflows built in designer 210. One commercial example of some embodiments of robot(s) 230 is UiPath Robots™. In some embodiments, robots 230 install the Microsoft Windows Service Control Manager (SCM) management service by default. As a result, such robots 230 can open interactive Windows sessions under the local system account and may have Windows service rights.
[0058] In some embodiments, a robot 230 can be installed in user mode, meaning that such a robot 230 has the same rights as the user to whom the given robot 230 is installed. This feature can also be available for high-density (HD) robots, ensuring maximum utilization of each machine. In some embodiments, either type of robot 230 can be configured in an HD environment.
[0059] In some embodiments, the robot 230 is divided into multiple components, each specialized for a specific automation task. In some embodiments, the robot components include, but are not limited to, an SCM-managed robot service, a user-mode robot service, an executor, an agent, and a command line. The SCM-managed robot service manages and monitors Windows sessions and acts as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 executes). These services are responsible for managing the credentials of the robot 230. A console application is launched by the SCM under Local System.
[0060] The user-mode robot service in some embodiments manages and monitors Windows sessions and acts as a proxy between the conductor 220 and the execution host. The user-mode robot service may be delegated and manage the credentials of the robot 230. If the SCM management robot service is not installed, a Windows application may be launched automatically.
[0061] An Executor may execute a given job under a Windows session (i.e., execute a workflow). An Executor may be aware of per-monitor dots-per-inch (DPI) settings. An Agent may be a Windows Presentation Foundation (WPF) application that displays available jobs in a system tray window. An Agent may be a client of a service. An Agent may request to start or stop a job or change settings. A Command Line is a client of a service. A Command Line is a console application that can request the start of a job and wait for its output.
[0062] As described above, the separation of the robot 230 components helps developers, support users, and computing systems more easily implement, identify, and track what each component is doing. In this way, special behaviors can be configured for each component, such as setting different firewall rules for executors and services. The executor may always be aware of per-monitor DPI settings in some embodiments. As a result, workflows may execute at any DPI regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 may be made independent of the browser zoom level. For applications that are not DPI-aware or are intentionally marked as not-aware, some embodiments may disable DPI.
[0063] The RPA system 200 in this embodiment is part of a hyperautomation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed to the core hyperautomation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for execution of the AI / ML model(s) and receive outputs therefrom via the core hyperautomation system 240.
[0064] One or more robots 230 may be listeners, as described above. These listeners may provide information to the core hyperautomation system 240 about what users are doing when they use their computing systems. This information may then be used by the core hyperautomation system for process mining, task mining, task capture, etc.
[0065] An assistant / chatbot 250 may be provided on a user computing system to allow the user to launch an RPA local robot. The assistant may be located, for example, in the system tray. The chatbot may have a user interface so that the user can view the chatbot's text. Alternatively, the chatbot may have no user interface, run in the background, and listen to the user's speech using the computing system's microphone.
[0066] In some embodiments, data labeling may be performed by a user of the computing system on which the robot is running, or on another computing system to which the robot provides information. For example, if the robot invokes an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user may draw a rectangle around the misidentified or unidentified component and potentially provide text with the correct identification. This information may be provided to the core hyperautomation system 240 and then later used to train a new version of the AI / ML model.
[0067] 3 is an architecture diagram illustrating a deployed RPA system 300 according to one or more embodiments. In some embodiments, the RPA system 300 may be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The deployed RPA system 300 may be a cloud-based system, an on-premise system, a desktop-based system, or the like, providing enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes.
[0068] It should be noted that the client side 301, the server side 302, or both may include any desired number of computing systems without departing from the scope of one or more embodiments described herein. On the client side 301, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 executes processes. As shown in FIG. 3, multiple business projects may run simultaneously. In this embodiment, the agent 314 (e.g., a Windows service) is the single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via the database server 355, the AI / ML server 360, the indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 may be a robotic component.
[0069] In some embodiments, a Robot represents an association between a machine name and a username. A Robot may manage multiple executors simultaneously. In computing systems that support multiple interactive sessions running simultaneously (such as Windows Server 2012), multiple Robots may run simultaneously, each running in a separate Windows session using a unique username. This is referred to as an HD Robot above.
[0070] Agent 314 is also responsible for transmitting the robot's status (e.g., periodically sending "heartbeat" messages to indicate that the robot is still functioning) and downloading required versions of packages to be fulfilled. Communication between agent 314 and conductor 340 is, in some embodiments, always initiated by agent 314. In notification scenarios, agent 314 may open a WebSocket channel that is later used by conductor 330 to send commands (e.g., start, stop, etc.) to the robot.
[0071] Listener 330 monitors and records data related to user interactions with the operation of the attended and / or unattended computing systems on which listener 330 resides. Listener 330 may be an RPA robot, part of an operating system, a downloadable application for the respective computing system, or any other software and / or hardware without departing from the scope of one or more embodiments described herein. Indeed, in some embodiments, the listener's logic is implemented partially or fully via physical hardware.
[0072] Server side 302 includes a conductor 340, as well as a presentation layer 333, a service layer 334, and a persistence layer 336. Presentation layer 333 may include a web application 342, an Open Data Protocol (OData) Representative State Transfer (REST) application programming interface (API) endpoint 344, and notifications and monitoring 346. Service layer 334 may include API implementation / business logic 348. Persistence layer 336 may include a database server 355, an AI / ML server 360, and an indexer server 370. For example, conductor 340 includes web application 342, an OData REST API endpoint 344, notifications and monitoring 346, and API implementation / business logic 348. In some embodiments, most actions a user performs in the conductor 340 interface (e.g., via browser 320) are performed by calling various APIs. Such operations may include, but are not limited to, launching jobs on robots, adding / removing data from a queue, scheduling jobs to run unattended, etc., without departing from the scope of one or more embodiments described herein. Web application 342 may be the visual layer of the server platform. In this embodiment, web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of one or more embodiments described herein. A user interacts with web pages from web application 342 via browser 320 in this embodiment to perform various operations to control conductor 340. For example, a user may create robot groups, assign packages to robots, analyze per-robot and / or per-process logs, start and stop robots, etc.
[0073] In addition to the web application 342, the conductor 340 also includes a services layer 334 that exposes an OData REST API endpoint 344. However, other endpoints may be included without departing from the scope of one or more embodiments described herein. The REST API is consumed by both the web application 342 and an agent 314, which in this embodiment is a supervisor of one or more robots on a client computer.
[0074] The REST API of this embodiment includes configuration, logging, monitoring, and queuing functionality (indicated by at least arrow 349). The configuration endpoint, in some embodiments, may be used to define and configure users, permissions, robots, assets, releases, and environments for an application. The logging REST endpoint may be used to log various information, such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by robots to query the version of a package that should be executed when a start job command is used in conductor 340. The queuing REST endpoint may be responsible for managing queues and queue items, such as adding data to a queue, retrieving transactions from a queue, and setting the status of transactions.
[0075] Monitoring REST endpoints may monitor web application 342 and agents 314. Notification and monitoring API 346 may be a REST endpoint used to register agents 314, deliver configuration settings to agents 314, and send and receive notifications from the server and agents 314. Notification and monitoring API 346 may, in some embodiments, use WebSocket communications. As shown in FIG. 3 , one or more activities / actions described herein are represented by arrows 350 and 351.
[0076] The APIs of the service layer 334, in some embodiments, may be accessed through configuration of appropriate API access paths, for example, based on whether the conductor 340 and the overall hyperautomation system have an on-premise or cloud-based deployment type. The APIs for the conductor 340 may provide custom methods for querying statistics about various entities registered with the conductor 340. Each logical resource, in some embodiments, may be an OData entity. In such entities, components such as robots, processes, queues, etc. may have properties, relationships, and behaviours. The APIs of the conductor 340, in some embodiments, may be consumed by the web application 342 and / or the agent 314 in two ways: by obtaining API access information from the conductor 340 or by registering an external application to use the OAuth flow.
[0077] In this embodiment, the persistence layer 336 includes three servers: a database server 355 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides AI / ML model provisioning services such as an AI center function), and an indexer server 370. The database server 355 in this embodiment stores configurations of robots, robot groups, associated processes, users, roles, schedules, etc. This information is managed via a web application 342 in some embodiments. The database server 355 may also manage queues and queue items. In some embodiments, the database server 355 may store messages logged by robots (in addition to or instead of the indexer server 370). The database server 355 may also store process mining, task mining, and / or task capture related data received, for example, from a listener 330 installed on the client side 301. Although no arrow is shown between the listener 330 and the database 355, it should be understood that the listener 330 can communicate with the database 355 in some embodiments, and vice versa. This data may be stored in the form of PDD, images, XAML files, etc. Listener 330 may be configured to intercept user actions, processes, tasks, and performance metrics on each computing system on which listener 330 resides. For example, listener 330 may record user actions (e.g., clicks, typed characters, location, application, active element, time, etc.) on its respective computing system and then convert these into a format suitable for being provided to and stored in database server 355.
[0078] AI / ML server 360 facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may make such capabilities accessible to non-data scientists. Deployed automation (e.g., RPA robots) may invoke AI / ML models from AI / ML server 360. The performance of AI / ML models may be trained and improved using monitored and human-validated data. AI / ML server 360 may schedule and execute training jobs to train new versions of AI / ML models.
[0079] AI / ML server 360 may store data related to AI / ML models and ML packages for configuring various ML skills for users during development. As used herein, an ML skill is a pre-built and trained ML model for a process that can be used, for example, by automation. AI / ML server 460 may also store data related to document understanding techniques and frameworks, algorithms, and software packages for various AI / ML capabilities, including, but not limited to, intent analysis, natural language processing (NLP), speech analysis, different types of AI / ML models, etc.
[0080] Optionally in some embodiments, indexer server 370 stores and indexes information logged by the robots. In particular embodiments, indexer server 370 may be disabled via a configuration setting. In some embodiments, indexer server 370 uses ElasticSearch®, a full-text search engine from an open source project. Messages logged by robots (e.g., using activities such as log messages or line writes) may be sent via logging REST endpoint(s) to indexer server 370, where they are indexed for future use.
[0081] FIG. 4 is an architecture diagram illustrating the relationships between a designer 410, activities 420, 430, 440, 450, a driver 460, an API 470, and an AI / ML model 480, according to one or more embodiments. As shown herein, a developer uses the designer 410 to develop a workflow to be performed by a robot. Various types of activities may be displayed to the developer in some embodiments. The designer 410 may be local to a user's computing system or remote thereto (e.g., accessed via a VM or a local web browser interacting with a remote web server). A workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. By way of example (as indicated by dotted lines), the user-defined activities 420 and API-driven activities 440 interact with applications via their APIs. The user-defined activities 420 and / or AI / ML activities 440 may then invoke one or more AI / ML models 480, which in some embodiments may be located locally and / or remotely relative to the computing system on which the robot is operating.
[0082] In some embodiments, non-text visual components in an image can be identified, referred to herein as CV. CV may be performed at least in part by AI / ML model(s) 480. Some CV activities related to such components may include, but are not limited to, extracting text from segmented label data using OCR, fuzzy text matching, cropping segmented label data using ML, comparing extracted text in label data to ground truth data, etc. In some embodiments, the number of activities that may be implemented in user-defined activities 420 may be in the hundreds or thousands. However, any number and / or type of activities may be used without departing from the scope of one or more embodiments described herein.
[0083] UI automation activities 450 are a subset of specialized low-level activities written in low-level code that facilitate interactions with the screen. UI automation activities 450 facilitate these interactions through drivers 460 that enable the robot to interact with desired software. For example, drivers 460 may include operating system (OS) drivers 462, browser drivers 464, VM drivers 466, enterprise application drivers 468, etc. In some embodiments, one or more AI / ML models 480 may be used by UI automation activities 450 to perform interactions with the computing system. In particular embodiments, AI / ML models 480 may augment or completely replace drivers 460. Indeed, in certain embodiments, drivers 460 are not included.
[0084] Driver 460 may interact with the OS at a low level, such as by looking for hooks or monitoring keys, via OS driver 462. Driver 460 may also facilitate integration with Chrome®, IE®, Citrix®, SAP®, etc. For example, a "click" activity plays the same role in these different applications via driver 460.
[0085] FIG. 5 is an architectural diagram illustrating a computing system 500 configured to provide automated form filling and ticket extraction from email, according to one or more embodiments. In some embodiments, computing system 500 may be one or more computing systems depicted and / or described herein. In particular embodiments, computing system 500 may be part of a hyper-automation system such as those shown in FIGS. 1 and 2. Computing system 500 includes a bus 505 or other communication mechanism for communicating information and processor(s) 510 coupled to bus 505 for processing information. Processor(s) 510 may be any type of general or application-specific processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. Processor(s) 510 may also have multiple processing cores, at least some of which may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In particular embodiments, at least one processor(s) 510 may be a neuromorphic circuit that includes processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require the typical components of a von Neumann computing architecture.
[0086] The computing system 500 further includes memory 515 for storing information and instructions executed by the processor(s) 510. The memory 515 may be comprised of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage such as a magnetic or optical disk, or any other type of non-transitory computer-readable medium, or any combination thereof. The non-transitory computer-readable medium may be any available medium that can be accessed by the processor(s) 510 and may include volatile media, non-volatile media, or both. Also, the medium may be removable, non-removable, or both.
[0087] Additionally, the computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, the communications device 520 may support any of the following radio technologies: Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), Global System for Mobile (GSM) communications, General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced (LTE-A), LTE-Advanced (LTE-B), LTE-Advanced (LTE-C), LTE-Advanced (LTE-B), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-C), LTE-Advanced (LTE-A), LTE-Advanced (LTE-C ... Advanced), 802.11x, Wi-Fi, Zigbee, Ultra-Wideband (UWB), 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near-Field Communications (NFC), Fifth Generation (5G), New Radio (NR), any combination thereof, and / or any other currently existing or future-implemented communication standard and / or protocol without departing from the scope of one or more embodiments described herein.In some embodiments, the communications device 520 may include one or more antennas that are a single antenna, an array of antennas, a panel of antennas, a phased antenna, a switched antenna, a beamforming antenna, a beamsteering antenna, a combination thereof, and / or any other antenna configuration without departing from the scope of one or more embodiments described herein.
[0088] The processor(s) 510 are further coupled via bus 505 to a display 525, such as a plasma display, a liquid crystal display (LCD), a light-emitting diode (LED) display, a field emission display (FED), an organic light-emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high-definition display, a Retina® display, an in-plane switching (IPS) display, or any other suitable display for displaying information to a user. The display 525 may be configured as a touch (haptic) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-touch display, or the like, using resistive, capacitive, surface acoustic wave (SAW) capacitive, infrared, optical imaging, dispersive signaling, acoustic pulse recognition, frustrated total internal reflection, or the like. Any suitable display device and haptic I / O may be used without departing from the scope of one or more embodiments described herein.
[0089] A keyboard 530 and cursor control device 535, such as a computer mouse, touchpad, etc., are further coupled to bus 505 to allow a user to interface with computing system 500. However, in certain embodiments, a physical keyboard and mouse may not be present, and the user may interact with the device solely through display 525 and / or a touchpad (not shown). Any type and combination of input devices may be used as a matter of design choice. In certain embodiments, no physical input devices and / or displays are present. For example, a user may interact with computing system 500 remotely through another computing system in communication with it, or computing system 500 may operate autonomously.
[0090] The memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include a module 545 (implementing an interface engine) configured to perform all or part of the processes described herein or derivatives thereof.
[0091] According to one or more embodiments, module 545 performs computer activity detection and automation. For example, module 545 records computer activity across one or more user interfaces (UIs), utilizes at least one generative AI model to automatically process the computer activity to extract patterns including portions of computer activity that resemble or are resistant to change, and determines multiple existing automations according to the patterns. Automatically may refer to autonomously performing computer operations without human intervention or initiation, where the speed and data scale of the autonomous computer operations exceed human capabilities.
[0092] The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0093] Those skilled in the art will understand that a “system” may be embodied as a server, embedded computing system, personal computer, console, personal digital assistant (PDA), mobile phone, tablet computing device, quantum computing system, or any other suitable computing device or combination of devices without departing from the scope of one or more embodiments described herein. Presenting the above-described functions as being performed by a “system” is not intended to limit the scope of the embodiments described herein in any way, but rather to provide an example of many possible embodiments. Indeed, the methods, systems, and apparatuses disclosed herein may be implemented in localized and distributed forms consistent with computing techniques, including cloud computing systems. The computing system may be part of or otherwise accessible by a local area network (LAN), a mobile communications network, a satellite communications network, the Internet, a public or private cloud, a hybrid cloud, a server farm, any combination thereof, or the like. Any local or distributed architecture may be used without departing from the scope of one or more embodiments described herein.
[0094] It should be noted that some of the system features described herein are presented as modules to further emphasize implementation independence. For example, a module may be implemented as a hardware circuit comprising custom very large scale integrated (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, etc.
[0095] Modules may also be implemented at least partially in software for execution by various types of processors. For example, an identified unit of executable code may include one or more physical or logical blocks of computer instructions, which may be organized, for example, as an object, a procedure, or a function. Nevertheless, the identified executable modules need not be physically located together; they may include separate instructions stored in different locations that, when logically combined, comprise a module to achieve the purpose stated for the module. Furthermore, modules may be stored on non-transitory computer-readable media, such as, for example, a hard disk drive, a flash device, RAM, tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of one or more embodiments described herein.
[0096] Indeed, a module of executable code may be a single instruction, many instructions, or even distributed across several different code segments, different programs, and multiple memory devices. Similarly, operational data may be identified and depicted herein within a module and may be embodied in any suitable form and organized within any suitable type of data structure. Operational data may be collected as a single data set or may be distributed in different locations across different storage devices, or may exist, at least in part, simply as electronic signals on a system or network.
[0097] Various types of AI / ML models may be trained and deployed without departing from the scope of one or more embodiments described herein. For example, FIG. 6 illustrates an example neural network 600 trained to recognize graphical elements in an image according to one or more embodiments. Here, neural network 600 receives pixels (shown in column 610) of a screenshot image of a 1920x1080 screen as input for input “neurons” 1 through I of its input layer (shown in column 620). In this case, I is 2,073,600, the total number of pixels in the screenshot image.
[0098] Neural network 600 also includes a number of hidden layers (represented by columns 630 and 640). While both DLNNs and shallow learning neural networks (SLNNs) typically have multiple layers, SLNNs may sometimes have only one or two layers, and typically have fewer than DLNNs. Typically, a neural network architecture includes an input layer, multiple intermediate layers (e.g., hidden layers), and an output layer (represented by column 650), as in neural network 600.
[0099] DLNNs often have many layers (10, 50, 200, etc.), with subsequent layers typically reusing features from previous layers to compute more complex and general functions. SLNNs, on the other hand, tend to train relatively quickly because they have only a few layers and expert features are pre-created from raw data samples, although feature extraction is tedious. DLNNs, on the other hand, typically do not require expert features, but they take longer to train and tend to have more layers.
[0100] In both approaches, layers are trained simultaneously on a training set and are typically checked for overfitting on a separate cross-validation set. Both techniques produce excellent results, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem being addressed by each neural network.
[0101] 6, the pixels provided as the input layer are fed as inputs to J neurons in hidden layer 1. In this example, every pixel is fed to each neuron, however, a variety of architectures are possible that may be used individually or in combination, including, but not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks without departing from the scope of one or more embodiments described herein.
[0102] Hidden layer 2 (630) receives input from hidden layer 1 (620), hidden layer 3 receives input from hidden layer 2 (630), and so on for all hidden layers until the last hidden layer (represented by oval 655) provides its output as input for the output layer. Note that the numbers of neurons I, J, K, and L are not necessarily equal, and thus any desired number of layers can be used for a given layer of neural network 600 without departing from the scope of one or more embodiments described herein. Indeed, in certain embodiments, the types of neurons in a given layer need not all be the same.
[0103] Neural network 600 is trained to assign confidence scores to graphical elements believed to have been found in the image. To reduce matches with unacceptably low likelihoods, in some embodiments, only those results with confidence scores equal to or greater than a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs with confidence scores above this amount may be used, while the remainder may be ignored. In this case, the output layer indicates that two text fields (represented by outputs 661 and 662), a text label (represented by output 663), and a submit button (represented by output 665) were found. Neural network 600 may provide the location, dimensions, image, and / or confidence score of these elements, which may then be used by an RPA robot or another process that uses this output for a predetermined purpose, without departing from the scope of one or more embodiments described herein.
[0104] Note that neural networks are probabilistic constructs that typically have a confidence score. This can be a score that the AI / ML model learned based on how often similar inputs were correctly identified during training. For example, text fields often have a rectangular shape and a white background. A neural network can learn to identify graphical elements with these characteristics with high confidence. Common types of confidence scores include a decimal number between 0 and 1 (which can be interpreted as a percentage of confidence), a number between negative infinity and positive infinity, or a set of expressions (e.g., "low," "medium," and "high"). Various post-processing calibration techniques, such as temperature scaling, batch normalization, weight decay, and negative log-likelihood (NLL), can also be employed in an attempt to obtain more accurate confidence scores.
[0105] A "neuron" in a neural network is typically a mathematical function based on the function of a biological neuron. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass an output to the next layer. This activation function may be a nonlinear, thresholded activity function that does nothing if the value is below a threshold, and responds linearly when the function exceeds the threshold (i.e., rectified linear unit (ReLU) nonlinearity). Sum and ReLU functions are used in deep learning because real neurons may have roughly similar activity functions. Information may be subtracted, added, etc. through linear transformations. Essentially, neurons act as gating functions that pass their output to the next layer, governed by their underlying mathematical function. In some embodiments, different functions may be used for at least some neurons.
[0106] JPEG2025161728000002.jpg64170
[0107] JPEG2025161728000003.jpg44170
[0108] JPEG2025161728000004.jpg29136
[0109] In this case, neuron 700 is a single-layer perceptron. However, any suitable neuron type or combination of neuron types may be used without departing from the scope of one or more embodiments described herein. It should also be noted that the range of values of the activation function weights and / or output value(s) may vary in some embodiments without departing from the scope of one or more embodiments described herein.
[0110] For example, in this case, successfully identifying a graphical element in an image, a goal, or "reward function," is often used. The reward function guides the search of the state space, exploring intermediate transitions and steps with both short-term and long-term rewards in an attempt to achieve the goal (e.g., successful identification of a graphical element, successful identification of the next sequence of activities in an RPA workflow, etc.).
[0111] During training, various labeled data (in this case, images) are fed through the neural network 600. Successful identification strengthens the input weights to neurons, while failed identification weakens those weights. A cost function such as mean squared error (MSE) or gradient descent may be used to ensure that slightly incorrect predictions are punished much less than significantly incorrect predictions. If the performance of an AI / ML model does not improve after a certain number of training iterations, a data scientist may modify the reward function to indicate where misidentified graphical elements are located, provide corrections for misidentified graphical elements, and so on.
[0112] Backpropagation is a technique for optimizing synaptic weights in feedforward neural networks. Backpropagation can be used to "pop the hood" of a neural network's hidden layers to see how much loss each node is carrying, and then update the weights to minimize loss, giving lower weights to nodes with higher error rates and vice versa. In other words, backpropagation allows data scientists to iteratively adjust the weights to minimize the difference between the actual output and the desired output.
[0113] The backpropagation algorithm is mathematically based on optimization theory. In supervised learning, training data with known outputs is passed through a neural network, and the error is calculated from the known target output using a cost function, which gives the backpropagation error. The error is calculated at the output, and this error is converted into a modification of the network weights that minimizes the error.
[0114] JPEG2025161728000005.jpg51170
[0115] JPEG2025161728000006.jpg19170
[0116] JPEG2025161728000007.jpg90170
[0117] JPEG2025161728000008.jpg117136
[0118] JPEG2025161728000009.jpg67170
[0119] The AI / ML model may be trained over multiple epochs until it reaches a good level of accuracy (e.g., 97% or greater using an F2 or F4 threshold for detection, approximately 2,000 epochs). This accuracy level may, in some embodiments, be determined using an F1 score, an F2 score, an F4 score, or any other suitable technique without departing from the scope of one or more embodiments described herein. Once trained on training data, the AI / ML model may be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps ensure that the AI / ML model does not "overfit," identifying graphical elements in the training data well but not generalizing well to other images.
[0120] In some embodiments, it may not be known what accuracy level an AI / ML model is capable of achieving. Thus, if the accuracy of the AI / ML model begins to degrade upon analyzing the evaluation data (i.e., the model performs well on the training data but begins to degrade on the evaluation data), the AI / ML model can undergo further epochs of training on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when a certain level of accuracy is reached or if the accuracy of the trained AI / ML model is superior to the existing deployed AI / ML model.
[0121] In particular embodiments, a collection of trained AI / ML models may be used to accomplish tasks such as employing an AI / ML model for each type of graphical element of interest, employing an AI / ML model to perform OCR, deploying yet another AI / ML model to recognize proximity relationships between graphical elements, and employing yet another AI / ML model to generate RPA workflows based on output from other AI / ML models, for example, so that the AI / ML models collectively enable semantic automation.
[0122] In some embodiments, a Transformer network such as SentenceTransformers™, a state-of-the-art Python™ framework for sentence, text, and image embedding, can be used. Such a Transformer network learns associations between words and phrases with both high and low scores. This trains an AI / ML model to determine what is close to the input and what is not, respectively. Rather than using only word / phrase pairs, the Transformer network may also use field lengths and field types.
[0123] FIG. 8 is a flowchart illustrating a process 800 for training an AI / ML model(s) according to one or more embodiments. Note that process 800 is also applicable to other UI learning operations, such as NLP and chatbots. The process begins at block 810 with training data that provides labeled data, such as that shown in FIG. 8 , such as labeled screens (e.g., with identified graphical elements and text), words and phrases, a “thesaurus” of semantic associations between words and phrases so that similar words and phrases to a given word or phrase can be identified, etc. The nature of the training data provided depends on the objectives the AI / ML model is intended to achieve. The AI / ML model is then trained for multiple epochs at block 820, and the results are reviewed at block 830.
[0124] If the AI / ML model does not meet the desired confidence threshold at decision block 840 (process 800 proceeds according to the no arrow), the training data is supplemented and / or the reward function is modified to help the AI / ML model better achieve its objective at block 850, and the process returns to block 820. If the AI / ML model meets the confidence threshold at decision block 840 (process 800 proceeds according to the yes arrow), the AI / ML model is tested against evaluation data at block 860 to ensure the AI / ML model generalizes well and does not overfit with respect to the training data. The evaluation data may include screens, source data, etc. that the AI / ML model has not previously processed. If the confidence threshold for the evaluation data is met at decision block 870 (process 800 proceeds according to the yes arrow), the AI / ML model is deployed at block 880. Otherwise (process 800 proceeds according to the no arrow), the process returns to block 880, and the AI / ML model is further trained.
[0125] 9-10 are flowcharts illustrating processes 900 and 1000 according to one or more embodiments. Generally, processes 900 and 1000 leverage RPA integrated with automated and generative artificial intelligence (AI) models for UI element detection / automation. For example, processes 900 and 1000 include performing computer activity detection and automation through an interface engine. Process 800 performed in FIG. 8 and processes 900 and 1000 performed in FIGS. 9-10 may be performed by an interface engine implemented in a computer program according to one or more embodiments. The computer program may be stored on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other such medium or combination of media used to store data. The computer program may include coded instructions for controlling a processor(s) of a computing system (e.g., module 545 of FIG. 5) to implement all or part of processes 800, 900, and 1000 described in FIGS. 8-10, which may also be stored on a computer-readable medium.
[0126] Referring to FIG. 9 , process 900 begins at block 910, where an interface engine records computer activity. Note that the recorded computer activity is in the form of non-standard data that spans languages, contexts, and the like. Examples of computer activity include, but are not limited to, one or more occurrences of a UI from a textual perspective (e.g., on the desktop of a computer display or the home screen of a mobile device display), one or more events, and one or more screenshots (e.g., UI controls such as buttons, text boxes, check boxes, etc.). As further examples, computer activity or events may include window opening, browser navigation, file opening and / or closing, folder opening and / or closing, and user interactions with the mouse and keyboard (click, type, copy, paste, etc.). Additionally, computer activity or events may be identified by one or more of a timestamp, user, category, application, application type, process name, URL screen, window title, action, action location, and action details.
[0127] At block 960, the interface engine automatically processes the computer activity. According to one or more embodiments, the interface engine automatically processes the computer activity by utilizing at least one generative AI model. The generative AI model may include any generative pre-trained transformer (GPT), AI agent (e.g., AutoGPT and UiPath Autopilot™), large language model (LLM) gateway, and other models. In this regard, the interface engine and / or generative AI model manipulates the non-standardized data of the computer activity and converts it into standardized data, i.e., patterns. Upon extraction, the patterns can be expressed in machine language, computer code, BPMN, or other processor-understandable programs (e.g., activity logs). The interface engine provides the technical effects, advantages, and benefits of documenting patterns (and their common variations), providing users with a valuable resource otherwise unavailable through conventional techniques. Thus, users can use patterns for process improvement through training, automation (as seen in block 990), and / or process change.
[0128] In block 990, the interface engine determines multiple existing automations according to the pattern. According to one or more embodiments, the interface engine and the generative AI model of the interface engine process the pattern to determine whether existing or new automations and / or RPAs (e.g., multiple automations) implement the computer activity.
[0129] 10, process 1000 begins at block 1010, where an interface engine records computer activity. The recording of computer activity by the interface engine may span one or more user interfaces (UIs). The recording of computer activity by the interface engine may be considered task mining. As discussed herein, the recorded computer activity is in the form of non-standardized data across languages, contexts, etc.
[0130] Examples of computer activity include, but are not limited to, one or more occurrences of a UI in terms of text (e.g., on the desktop of a computer display or the home screen of a mobile device display), one or more events, and one or more screenshots (e.g., UI controls such as buttons, text boxes, check boxes, etc.) For example, sub-blocks 1020, 1030, and 1040 further describe task mining by the interface engine for occurrences, events, and screenshots, respectively.
[0131] In subblock 1020, the interface engine records everything that is happening on the desktop of the home screen from a textual perspective. In this regard, computer activity includes text changes made by the desktop of the computer display or the home screen of the mobile device display.
[0132] In subblock 1030, the interface engine records all events. In this regard, computer activity includes one or more events. An event may include an action or occurrence that is generated or triggered within the computing environment and recognized by software in the computing environment. Examples of actions or occurrences include, but are not limited to, an application error, an application launch, an application termination, a change or switch between UIs, saving an instance, selecting an icon, cutting and pasting, and activities that occur asynchronously and externally from the computing environment.
[0133] At sub-block 1040, the interface engine captures screenshots. In this regard, computer activity includes one or more screenshots of a display (e.g., the desktop of a computer display or the home screen of a mobile device display). According to one or more embodiments, the interface engine can selectively capture an active window, an active interface, the entire desktop, or a combination thereof (e.g., if a user is typing in a word processing window and also has an email window open, the interface engine can selectively capture a screenshot of the word processing window).
[0134] At block 1060, the interface engine automatically processes the computer activity. According to one or more embodiments, the interface engine automatically processes the computer activity by utilizing at least one generative AI model. The interface engine and / or the generative AI model manipulates the non-standardized data of the computer activity and converts it into standard data, i.e., patterns. Upon extraction, the patterns can be expressed in machine language, computer code, BPMN, or other processor-understandable programs (e.g., activity logs). Examples of what patterns can represent include, but are not limited to, filling out a form (e.g., creating a ticket) and a multi-step / screen process (receiving and entering an invoice). Note that patterns do not specify control, as the generative AI model (LLM) provides this action. As an example, the interface engine manipulates the computer activity to create a new, computer-understandable activity log that would not otherwise exist, improving the operation of at least one processor running the interface engine. Furthermore, using a generative AI approach can avoid limitations of conventional techniques, and results can be language-independent. Additionally, as a technical effect, benefit, and advantage, it is noted that client-side software generally does not have a back-end mechanism for describing all user activity (e.g., there is no log of user actions within a spreadsheet program), and the interface engine solves this limitation with a new computer-understandable activity log.
[0135] According to one or more embodiments, when processing computer activity, the interface engine and / or generative AI model extracts patterns (e.g., extracts patterns from text, events, images (using GPT)). The patterns may be language-independent, which is a technical effect, advantage, and benefit over prior art, as the interface engine can detect similarities between interfaces exhibiting different languages. The patterns may include portions of the computer activity that are similar or resistant to change. The changes may include changes in one or more of color, theme, language, resolution, screen size, and browser across the UI. Examples of similarities may include the exact same article opened on two different devices looking different in different browsers, different screen resolutions, different window sizes, different font scaling, different images reflowing in a different order, different scroll points on the page, etc. In this regard, opening these two instances is the same activity (e.g., users A and B looking at article X). Examples of similarity may include when a form can be reflowed (e.g., a label moves from the left side of a text box to above the text box, a save button scrolls from the bottom of the screen, or different fields are visible at certain times) or when opening an issue in Jira when there are multiple views of the same issue (e.g., a full-screen view and a side panel). The pattern may identify one or more controls on individual screens that are similar or tolerant to change and group multiple screens as part of a computer activity according to the one or more controls. According to one or more technical effects, benefits, and advantages, the interface engine extracts a pattern-generating AI model to find controls on individual screens and group multiple screens that are similar and / or tolerant to changes in color, theme, language, resolution, screen size, and browser.
[0136] Optionally, in block 1070, the interface engine monitors the network layer. The interface engine monitors the network layer to track where data packets are being forwarded. These data packet traces can be utilized by the interface engine as additional computer activity to detect and automate computer activity. According to one or more embodiments, the interface engine can monitor API calls made by a browser or client software and extract key information from the API calls. For example, a user navigating to customer management software to open a customer account can perform a search for the customer using text. Note that customers are stored by identification number (e.g., rather than name), and the identification number is clearly indicated in the “GET API request” that the interface engine can record from the network. Other examples may include, but are not limited to, monitoring all uniform resource locators (URLs) hit when opening a web page, all post commands that create artifacts, all put / patch commands that make updates, etc. Because network monitoring is very noisy (e.g., hundreds of thousands of calls per day), the interface engine's generative AI models are well-suited to detecting patterns.
[0137] Optionally, at block 1080, the interface engine can integrate structured data generated from the events with unstructured data from the occurrences, events, and screenshots. For example, the interface engine may identify unstructured data from web pages navigated by the user and related documents displayed, as well as identify open accounts, and integrate the unstructured data with the structured data.
[0138] In block 1090, the interface engine determines multiple existing automations according to the pattern. According to one or more embodiments, the interface engine searches a repository of existing automations or RPAs according to descriptions or titles associated with the existing automations or RPAs to find one or more matches to the pattern. If an existing automation match is found, the interface engine adds the existing automation match to the automation plan. If an existing automation match is not found, the interface engine uses a designer application (e.g., UiPath Studio™) to create one or more new automations and / or RPAs to accomplish the computer activity.
[0139] The computer program may be implemented in hardware, software, or a hybrid implementation. The computer program may be composed of modules in operable communication with each other and designed to send information or instructions to a display. The computer program may be configured to run on a general-purpose computer, an ASIC, or any other suitable device.
[0140] It will be readily understood that the components of the various embodiments, as generally described and illustrated herein, may be arranged and designed in a wide variety of different configurations. Thus, the detailed description of the embodiments, as represented in the accompanying figures, is not intended to limit scope, as claimed, but is merely representative of selected embodiments.
[0141] The features, structures, or characteristics described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments," "some embodiments," or similar language mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of "certain embodiments," "some embodiments," "other embodiments," or similar language throughout this specification do not necessarily refer to the same group of all embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0142] It should be noted that references to features, advantages, or similar language throughout this specification do not imply that all of the features and advantages that may be realized should be present in any single embodiment, or that any embodiment is. Rather, language referring to features and advantages is understood to mean that the particular feature, advantage, or characteristic described in connection with an embodiment is included in one or more embodiments. Thus, discussions of features and advantages throughout this specification, and similar language, can, but do not necessarily, refer to the same embodiment.
[0143] Furthermore, the features, advantages, and characteristics of one or more embodiments described herein may be combined in any suitable manner. Those skilled in the relevant art will recognize that the present disclosure may be practiced without certain features or advantages of one or more particular embodiments. In other instances, additional features and advantages may be recognized in certain embodiments, even though they may not be present in all embodiments.
[0144] Those of ordinary skill in the art will readily appreciate that the present disclosure can be practiced using steps in a different order and / or with hardware elements in a different configuration than that disclosed. Thus, while the present disclosure has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain modifications, variations, and alternative configurations will become apparent while remaining within the spirit and scope of the present disclosure. Accordingly, reference should be made to the appended claims to determine the scope of the present disclosure.
Claims
1. 1. A method performed by an interface engine implemented as a computer program within a computing environment, the interface engine performing detection and automation of computer activity, the method comprising: recording said computer activity across one or more user interfaces (UIs); utilizing at least one generative AI model to automatically process the computer activity to extract patterns including portions of the computer activity that are similar to or resistant to change; determining a plurality of existing automations according to the pattern.
2. The method of claim 1 , wherein the computer activity includes changing text on a desktop or home screen of a display.
3. 10. The method of claim 1, wherein the computer activity comprises one or more events that include actions or occurrences generated or triggered within a computing environment and recognized by software of the computing environment, examples of the actions or occurrences include, but are not limited to, application errors, application launches, application terminations, changes or switches between UIs, saving instances, icon selections, cut and paste, and activities that occur asynchronously and externally from the computing environment.
4. The method of claim 1 , wherein the computer activity includes one or more screenshots of a display.
5. 10. The method of claim 1, wherein the generative AI model comprises one or more of a generative pre-trained transform (GPT), an AI agent, and a large language model (LLM) gateway.
6. The method of claim 1 , wherein the patterns extracted by the interface engine are language agnostic.
7. The method of claim 1 , wherein the changes include changing one or more of the following across the UI: color, theme, language, resolution, screen size, and browser.
8. 2. The method of claim 1, wherein the pattern identifies one or more controls on individual screens that are similar to or resistant to change, and groups multiple screens as part of the computer activity according to the one or more controls.
9. The method of claim 1 , wherein the plurality of automations includes new automation or robotic process automation (RPA).
10. 10. The method of claim 1, wherein the interface engine searches a repository of existing automations or RPAs according to descriptions or titles associated with the existing automations or RPAs to find one or more matches with the pattern.
11. a memory for storing program code of an interface engine for performing computer activity detection and automation; at least one processor that executes the program code to cause the interface engine to: recording said computer activity across one or more user interfaces (UIs); utilizing at least one generative AI model to automatically process the computer activity to extract patterns including portions of the computer activity that are similar to or resistant to change; Determining a number of existing automations according to the pattern.
12. The computing system of claim 11 , wherein the computer activity includes changing text on a desktop or home screen of a display.
13. 12. The computing system of claim 11, wherein the computer activity comprises one or more events comprising actions or occurrences generated or triggered within the computing environment and recognized by software of the computing environment, examples of the actions or occurrences including, but not limited to, an application error, an application launch, an application termination, a change or switch between UIs, saving an instance, selecting an icon, cutting and pasting, and activities occurring asynchronously and externally from the computing environment.
14. The computing system of claim 11 , wherein the computer activity includes one or more screenshots of a display.
15. 12. The computing system of claim 11, wherein the generative AI model comprises one or more of a generative pre-trained transform (GPT), an AI agent, and a large language model (LLM) gateway.
16. The computing system of claim 11 , wherein the patterns extracted by the interface engine are language agnostic.
17. The computing system of claim 11 , wherein the modifications include changing one or more of the following across the UI: color, theme, language, resolution, screen size, and browser.
18. 12. The computing system of claim 11, wherein the pattern identifies one or more controls on individual screens that are similar to or resistant to change, and groups multiple screens as part of the computer activity according to the one or more controls.
19. 12. The computing system of claim 11, wherein the plurality of automations includes new automation or robotic process automation (RPA).
20. 12. The computing system of claim 11, wherein the interface engine searches a repository of existing automations or RPAs according to descriptions or titles associated with the existing automations or RPAs to find one or more matches with the pattern.
Citation Information
Cited By
Browser AI integration system
JP7891198B1