User interface automation using robotic process automation for detecting user interface element not displayed on display and filling form
The interface engine addresses the limitations of conventional clipboard technologies by using computer vision and DOM analysis to detect and automate UI elements, including those not visible on the screen, thereby enhancing UI automation efficiency.
Patent Information
- Application Number
- JP2024070399
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-06
- Filing Date
- 2024-04-24
- Publication Date
- 2025-05-19
AI Technical Summary
Conventional clipboard technologies struggle to detect and fill out input fields that are not visible on the screen, limiting their ability to automate the completion of long forms.
An interface engine that uses computer vision and Document Object Model (DOM) analysis to detect and automatically fill UI elements, including those not visible on the screen, by extracting DOM text and estimating target UI elements in invisible regions.
Enables the identification and automation of UI elements not displayed on the screen without scrolling, improving UI automation efficiency and reducing user intervention.
Smart Images

Figure 2025077955000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to automation, and more particularly to an interface engine that provides automation of a user interface (UI) using one or more robotic process automations (RPAs) to detect and fill out forms.
Background Art
[0002] Generally, long forms can be presented in a UI where parts of the long form may or may not be visible on the screen. Conventional clipboard technologies use computer vision (CV)-based detection to perform destination analysis of the screen and determine the input fields there. Thus, conventional clipboard technologies can focus on visible input fields and cannot detect input fields that are not visible on the screen.
[0003] Another problem with conventional clipboard technologies is that when a user wants to paste source field data into a long form using a conventional clipboard technology, the user can only paste the source field data into the visible input fields of the long form. And input fields that are not displayed on the screen (for example, are only displayed when manually scrolled) cannot be mapped to source field data with conventional clipboard technologies. This problem becomes even more severe when the user repeatedly performs partial scrolls and re-uses the conventional clipboard technology to map new iterations of visible input fields onto the screen.
[0004] There is a need for a solution to automatically fill out and / or complete long forms.
Summary of the Invention
[0005] According to one or more embodiments, a method is provided. The method is for detecting and automatically filling one or more user interface (UI) elements of a page that are not displayed on the screen. The method is performed by an interface engine implemented as a computer program within a computing environment. The method includes analyzing, by the interface engine, a document object model (DOM) of the page to extract DOM text of one or more relevant UI elements among the one or more UI elements. The method includes performing computer vision (CV) analysis to determine one or more types of the one or more relevant UI elements and estimating one or more target UI elements of an invisible area of the page from the one or more relevant UI elements.
[0006] According to one or more embodiments, the above method may be implemented in a system, a computer product, and an apparatus.
Brief Description of the Drawings
[0007] To facilitate an understanding of the advantages of the specific embodiments of the present specification, a more specific description is depicted with reference to the specific embodiments illustrated in the accompanying drawings. It should be understood that these drawings depict only typical embodiments and are not considered to limit the scope. However, the one or more embodiments described in the present specification will be described and explained in more detail and specificity by using the following accompanying drawings.
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
[0012]
Figure 5
[0013]
Figure 6
[0014]
Figure 7
[0015]
Figure 8
[0016]
Figure 9
[0017]
Figure 10
[0018]
Figure 11
[0019]
Figure 12
[0020]
Figure 13
[0021] Unless otherwise noted, like reference characters consistently refer to corresponding features throughout the accompanying drawings.
DETAILED DESCRIPTION OF THE INVENTION
[0022] (Detailed Description of Embodiments) The present disclosure generally relates to automation, and more particularly to an interface engine that provides automation of a user interface (UI) using one or more robotic process automations (RPA) to fill out forms. As an example, the interface engine detects UI elements that are not displayed on the display (e.g., outside the displayable area of the UI presented on the screen) and provides UI automation that uses one or more RPAs to input / complete a form therein. The UI elements can include, but are not limited to, one or more input fields of a form that are not visible within the displayable area of the UI displayed on the screen. The input / completion of the one or more input fields can be performed by the interface engine without moving or scrolling the UI displayed on the screen. The implementation of the interface engine can be performed by a computing system (as described herein).
[0023] According to one or more embodiments, the interface engine performs one or more operations to detect UI elements not displayed on the screen while automatically entering a long form / page. The interface engine analyzes the Document Object Model (DOM) of the long form / page to extract the DOM text of the relevant UI elements, performs computer vision (CV) analysis to determine the type of the relevant UI elements, and may estimate the target UI elements of the invisible regions of the long form / page from the relevant UI elements. The interface engine may also include linking the DOM text to the target UI elements based on the type determined by the CV analysis and the relevant UI elements. The estimation by the interface engine includes searching the DOM for similar Hypertext Markup Language (HTML) structures (e.g., the relevant DOM text corresponding to the target UI elements). Further, the interface engine can find the fields (i.e., UI elements) of the long form / page from both the visible and invisible regions, and can extract data from the source for pasting into the fields (e.g., on a subsequent or destination screen). Additionally, the interface engine also has AI / ML capabilities to detect UI elements and automatically enter them into the long form / page.
[0024] Thus, one or more advantages, technical effects, and / or benefits of the interface engine include being able to identify UI elements not displayed on the screen without scrolling or moving the page presented by the screen, which is not available or currently not performed by conventional clipboard techniques. In this regard, the interface engine improves UI automation for identifying input field elements and reduces the time consumed to overcome the scroll activity performed by the user to show the target UI elements.
[0025] FIG. 1 is an architectural diagram showing a hyper - automation system 100 according to one or more embodiments. As used herein, "hyper - automation" refers to an automation system that combines components of process automation, integration tools, and technologies that amplify the ability to automate work. For example, in some embodiments, RPA is used at the core of the hyper - automation system, and in certain embodiments, the automation capabilities can be extended by artificial intelligence and / or machine learning (AI / ML), process mining, analytics, and / or other advanced tools. When a hyper - automation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work can be automated, and both computing systems within an organization, such as those used by individuals and those operating autonomously, can all participate as participants in the hyper - automation process. The hyper - automation systems of some embodiments enable users and organizations to discover, understand, and expand automation efficiently and effectively.
[0026] The hyper - automation system 100 includes user computing systems such as, for example, desktop computer 102, tablet 104, and smartphone 106. However, any desired computing system, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc., may be used without departing from the scope of one or more embodiments described herein. Also, although three user computing systems are shown in FIG. 1, any suitable number of computing systems may be used without departing from the scope of one or more embodiments described herein. For example, in some embodiments, dozens, hundreds, thousands, or millions of computing systems may be used. The user computing systems may be actively used by users or may be automatically executed with little or no user input.
[0027] Each computing system 102, 104, 106 has its respective automation process(es) 110, 112, 114 running thereon. The automation process(es) 102, 104, 106 can include, without limitation, and without departing from the scope of one or more embodiments described herein, RPA robots, a part of an operating system, downloadable application(s) for each computing system, any other suitable software and / or hardware, or any combination thereof. In some embodiments, one or more process(es) 110, 112, 114 can be listeners. The listener can be, without departing from the scope of one or more embodiments described herein, an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the listener is implemented partially or fully via physical hardware.
[0028] The listener monitors and records data related to the user interaction with each computing system and / or the operation of the unattended computing system, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data can include, but is not limited to, which buttons were clicked, where the mouse moved, the text entered in a field, that one window was minimized and another window was opened, the application associated with the window, etc. In certain embodiments, the data from the listener can be transmitted periodically as part of a heartbeat message. In some embodiments, the data can be transmitted to the core hyper-automation system 120 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0029] The automation process can perform the logic developed in the workflow during design time. In the case of RPA, the workflow can include a set of steps performed in a sequence or some other logical flow, defined herein as an "activity". Each activity can include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, the workflows can be nested or embedded.
[0030] The long-running workflows for RPA in some embodiments are master projects that support service orchestration, human intervention, and long-running transactions in an unattended environment. See U.S. Patent No. 10,860,905. Also, all of this content is incorporated by reference. Human intervention occurs when a particular process requires human input for exception handling, approval, or verification before proceeding to the next step of an activity. In this case, the execution of the process is paused and the RPA robot is released until the human task is completed.
[0031] The long-running workflows may support fragmentation of the workflow via persistence activities, combined with call processes and non-user interaction activities, and orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or a large number of computing systems may participate in the execution of the logic of the long-running workflow. The long-running workflow may be executed in a session to facilitate rapid execution. In some embodiments, the long-running workflow may orchestrate a background process that executes application programming interface (API) calls and may include activities that execute in the long-running workflow session. These activities may, in some embodiments, be called by call process activities. A process having user interaction activities that execute in a user session may be called by starting a job from a conductor activity (the conductor is described in more detail later in this specification). The user may, in some embodiments, interact through tasks that require the user to complete a form in the conductor. Activities may be included that cause the RPA robot to wait for the form task to be completed and then resume the long-running workflow.
[0032] One or more automation processes 110, 112, 114 communicate with a core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may execute a conductor application on one or more servers, e.g., server 130. Although one server 130 is shown for illustration purposes, multiple or numerous servers in close proximity to each other or in a distributed architecture may be employed without departing from the scope of one or more embodiments described herein. For example, one or more servers may be provided for conductor functionality, AI / ML model provisioning, authentication, governance, and / or any other suitable functionality without departing from the scope of one or more embodiments described herein. In some embodiments, the core hyper-automation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyper-automation system 120 may host multiple software-based servers on one or more computing systems, e.g., server 130. In some embodiments, one or more servers, such as the core hyper-automation system 120, e.g., server 130, may be implemented via one or more virtual machines (VMs).
[0033] In some embodiments, one or more automation processes 110, 112, 114 may invoke one or more AI / ML models 132 deployed on or accessible by a core hyper-automation system 120. The AI / ML models 132 can be trained for any suitable purpose without departing from the scope of one or more embodiments described herein, as discussed in detail herein. In some embodiments, two or more AI / ML models 132 may be chained (e.g., in series, in parallel, or a combination thereof) such that they collectively provide a collaborative output(s). The AI / ML models 132 may perform or assist with computer vision (CV), optical character recognition (OCR), document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, combinations of any of these, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of one or more embodiments described herein. By using multiple AI / ML models, for example, a system can develop a holistic picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, etc. Patterns may be determined individually by an AI / ML model or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one computing system 102, 104, 106.
[0034] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that executes on data, and the AI / ML model itself may be, for example, a deep learning neural network (DLNN) of artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that execute various functions of, for example, statistical modeling (such as a hidden Markov model (HMM)), and may utilize deep learning techniques (such as long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to execute the desired functions.
[0035] In some embodiments, the hyper-automation system 100 may provide four main functional groups: (1) discovery, (2) automation construction, (3) management, and (4) engagement. Automation (e.g., executed on user computing systems, servers, etc.) may be executed, in some embodiments, by software robots such as RPA robots. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with the user to assist the user in tasks (e.g., via UiPath Assistant™). Unattended robots operate independently of the user and may potentially execute in the background without the user's knowledge. Test robots are unattended robots that execute test cases against an application or an RPA workflow. Test robots may be executed in parallel on multiple computing systems in some embodiments.
[0036] The discovery function can discover various opportunities for business process automation and provide automated recommendations thereof. Such a function can be implemented by one or more servers, for example, server 130. The discovery function may include, in some embodiments, providing an automation hub, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath Automation Hub (trademark)) can provide a mechanism for managing an automation rollout with visibility and controllability. Automation ideas can be crowdsourced from employees, for example, via a submission form. Feasibility and return on investment (ROI) calculations for automating these ideas are provided, documents for future automation are collected, and collaboration for quickly performing automation from discovery to construction can be provided.
[0037] Process mining (e.g., via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) refers to the process of collecting and analyzing data from applications (enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how to effectively automate them, and the potential impacts that automation may bring. This data can be obtained from user computing systems 102, 104, 106 by a listener, for example, and processed by a server, for example, server 130. In some embodiments, one or more AI / ML models 132 can be employed for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of the goals of process mining include, but are not limited to, increased profits, improved customer satisfaction, regulatory and / or compliance, and improved employee efficiency.
[0038] Task mining (e.g., via UiPath Automation Cloud™ and / or UiPath AI Center™) identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in routine tasks and scores such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to reveal repetitive task patterns within the data. Repetitive tasks ripe for automation can then be identified. This information may initially be provided by a listener and, in some embodiments, analyzed on a server of the core hyperautomation system 120, such as server 130. Discoveries from task mining (e.g., extensive application markup language (XAML) process data) are exported to a process document or a designer application such as UiPath Studio™ to enable more rapid creation and deployment of automation.
[0039] Task mining in some embodiments may include taking screenshots with user actions (e.g., mouse click location, keyboard input, application windows and graphical elements the user interacted with, timestamps for the interaction, etc.), collecting statistical data (e.g., execution time, number of actions, text input, etc.), editing and annotating the screenshots, specifying the types of actions being recorded, and the like.
[0040] (Via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) Task capture automatically documents attended processes while the user is working or provides a framework for unattended processes. Such documentation may include process definition documents (PDDs), skeleton workflows, capture of actions for each part of the process, recording of user actions and automatic generation of comprehensive workflow diagrams including details for each step, tasks that are desirable to automate in formats such as Microsoft Word (registered trademark) documents, XAML files, and other documentation. Constructible workflows can, in some embodiments, be directly exported to designer applications such as, for example, UiPath Studio (trademark). Task capture can simplify the requirements gathering process for both subject matter experts who describe the process and Center of Excellence (CoE) members who provide production grade automation.
[0041] Automation can be achieved through designer applications (such as UiPath Studio™, UiPath StudioX™, UiPath Web™, etc.). For example, RPA developers at the PA development facility 150 can use the RPA designer application 154 on the computing system 152 to build and test automation for various applications and environments such as web, mobile, SAP®, and virtual desktop. API integration can be provided for various applications, technologies, and platforms. Pre-defined activities, drag-and-drop modeling, and workflow recorders can facilitate automation with minimal coding. The document understanding feature can be provided through drag-and-drop AI skills for data extraction and interpretation that call one or more AI / ML models 132. Such automation can process virtually any document type and format, including tables, checkboxes, signatures, and handwritten. When data is validated or exceptions are handled, this information can be used to retrain the respective AI / ML models, improving their accuracy over time.
[0042] With the integrated service, developers can seamlessly combine, for example, the automation of the user interface (UI) and the automation of APIs. Automations that require APIs or that cross both API and non-API applications and systems can be built. A repository (e.g., UiPath Object Repository (trademark)) or marketplace (e.g., UiPath Marketplace (trademark)) for pre-built RPA and AI templates and solutions can be provided so that developers can automate a wide variety of processes more quickly. Thus, when building an automation, the hyper-automation system 100 can provide a user interface, a development environment, API integration, pre-built and / or custom-built AI / ML models, development templates, an integrated development environment (IDE), and advanced AI capabilities. The hyper-automation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.
[0043] In some embodiments, components of the hyper-automation system 100, such as, for example, designer applications and / or an external rule engine, provide support for managing and enforcing governance policies for controlling the various functions provided by the hyper-automation system 100. Governance is the ability of an organization to introduce policies to prevent the development of automation (such as RPA robots) that can harm the organization by users, for example, in violation of the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), the terms of use of third-party applications, etc. Otherwise, developers could create automation that violates privacy laws, terms of use, etc. during the execution of their automation. Thus, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. This can provide an additional level of security and compliance in the automation process development pipeline in some embodiments by preventing developers from introducing security risks or taking dependencies on unapproved software libraries that could operate in a way that violates policies, regulations, privacy laws, and / or privacy policies. See U.S. Non-Provisional Patent Application No. 16 / 924,499. Also, all of this content is incorporated by reference.
[0044] The management function can provide management, deployment, and optimization of automation across the entire organization. The management function may include, in some embodiments, orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management function of the hyper-automation system 100 can include, among other things, but not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots.
[0045] For example, a conductor application such as UiPath Orchestrator (trademark) (which may be provided as part of UiPath Automation Cloud (trademark) in some embodiments, or on-premises, VM, private or public cloud, on a Linux (trademark) VM, or as a cloud-native single-container suite via UiPath Automation Suite (trademark)) provides orchestration capabilities to deploy, monitor, optimize, scale, and secure RPA robot deployments. A test suite (e.g., UiPath Test Suite (trademark)) can provide test management for monitoring the quality of deployed automation. The test suite can facilitate test planning and execution, requirement fulfillment, and defect traceability. The test suite can include comprehensive test reports.
[0046] Analysis software (e.g., UiPath Insights (trademark)) can track, measure, and manage the performance of deployed automation. The analysis software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes of the organization. The analysis software can present results in dashboard form for easier understanding by human users.
[0047] A data service (e.g., UiPath Data Service (trademark)) can, for example, be stored in a database 140 and bring data into a single, scalable, and secure place using a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities may be provided by an AI center (e.g., UiPath AI Center (trademark)), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may enable non-data scientists to access such capabilities. Deployed automation (e.g., an RPA robot) can call an AI / ML model from an AI center, such as AI / ML model 132. The performance of the AI / ML model is monitored and can be trained and improved using human-verified data provided, for example, by a data review center 160. A human reviewer may provide labeled data to the core hyperautomation system 120 via a review application 152 on a computing system 154. For example, a human reviewer may verify that the predictions by the AI / ML model 132 are accurate or, if not, provide corrections. For example, this dynamic input may then be saved as training data for retraining the AI / ML model 132 and may be stored in a database, such as database 140. The AI center can then schedule and execute a training job to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining the AI / ML model 132.
[0048] The engagement function involves humans and automation as one team for seamless collaboration regarding a desired process. Low-code applications can be built (e.g., via UiPath Apps™) even if they lack an API in some embodiments to connect browser tabs and legacy software. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. An application can be connected to one automation or multiple automations.
[0049] The Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for passing a process from automation to humans or vice versa. Humans can provide approvals or escalations and perform exception handling, etc. Then, the automation can execute the automated functions of a given workflow.
[0050] The local assistant can be provided as a launch pad for the user to launch automation (e.g., UiPath Assistant (trademark)). This feature may be provided, for example, in the tray provided by the operating system, and may enable the user to interact with RPA robots and RPA robot - enabled applications on their computing system. The interface may list the automations approved for a given user and allow the user to execute them. These may include off - the - shelf automations from an automation marketplace, an internal automation store in an automation hub, etc. While the automation is running, they may execute as a local instance in parallel with other processes on the computing system so that the user can use the computing system while the automation performs its actions. In certain embodiments, the assistant is integrated with a task capture function so that the user can document the processes that will soon be automated from the assistant's launch pad.
[0051] Chatbots (e.g., UiPath Chatbots (trademark)), social messaging applications, and / or voice commands may enable the user to execute automations. This can simplify access to the information, tools, and resources necessary to conduct customer interactions or other activities. Human - to - human conversations can be easily automated just like other processes. The triggered RPA robots launched in this way may be able to perform actions such as order status checks, data posting to CRM, etc., using plain - language commands.
[0052] End-to-end measurement of automation programs at any scale and governance can be provided by the hyper-automation system 100 in some embodiments. As such, analysis (e.g., via UiPath Insights™) may be employed to understand the performance of the automation. Data modeling and analysis using any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for automation, perform remote measurement monitoring on the user's computing system, detect errors and anomalies, and debug the automation. An automation management console (e.g., UiPath Automation Ops™) may be provided to manage the automation throughout its lifecycle. An organization may govern how the automations are built, what users can do with them, and which automations users can access.
[0053] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes can be discovered, automations can be built, tested, and deployed, performance can be measured, use of the automations can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the processes themselves can be repeated. This promotes a more robust and effective set of automations.
[0054] FIG. 2 is an architectural diagram showing an RPA system 200 according to one or more embodiments. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 may facilitate the development of an automation project that is a graphical representation of a business process. Put simply, the designer 210 facilitates the development and deployment of workflows and robots (indicated by arrow 211). In some embodiments, the designer 210 may be an application running on a user's desktop, an application running remotely on a VM, a web application, or the like.
[0055] An automation project enables the automation of a rule-based process by giving the developer control over the order of execution and relationships between a custom set of steps developed in a workflow, as defined herein as an "activity" above. A commercial example of an embodiment of the designer 210 is UiPath Studio (trademark). Each activity may include actions such as clicking a button, reading a file, writing to a log panel, and the like. In some embodiments, workflows may be nested or embedded.
[0056] Some types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. A sequence may be particularly suitable for a linear process that enables the flow from one activity to another without cluttering the workflow. A flowchart may be particularly suitable for more complex business logic and enables the integration of decision-making and the connection of activities in more diverse ways through multiple branching logic operators. An FSM may be particularly suitable for large-scale workflows. An FSM may use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. A global exception handler may be particularly suitable for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0057] When a workflow is developed within the designer 210, the execution of the business process is coordinated by the conductor 220, which coordinates one or more robots 230 that execute the workflow developed within the designer 210. A commercial example of an embodiment of the conductor 220 is UiPath Orchestrator (trademark). The conductor 220 facilitates the management of the generation, monitoring, and deployment of resources in an environment. The conductor 220 may operate as an integration point with third-party solutions and applications. As such, in some embodiments, the conductor 220 may be part of the core hyper-automation system 120 of FIG. 1.
[0058] The conductor 220 can manage all the robots 230 and connect and execute the robots 230 from a centralized point (indicated by arrow 231). The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non-production robots (similar to attended robots 232 but used for development and testing purposes). The attended robot 232 is triggered by user events and operates in parallel with humans on the same computing system. The attended robot 232 can be used together with the conductor 220 for centralized process deployment and logging media. The attended robot 232 may assist a human user in achieving various tasks and may be triggered by user events. In some embodiments, the process cannot start from the conductor 220 on this type of robot and / or they cannot execute under a locked screen. In certain embodiments, the attended robot 232 can only be activated from a robot tray or from a command prompt. The attended robot 232 preferably operates under human supervision in some embodiments.
[0059] The unattended robot 234 can operate unmanned in a virtual environment and automate many processes. The unattended robot 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can be performed by the designer 210 in some embodiments. Both attended and unattended robots can automate various systems and applications (shown by the dashed box 290) including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0060] The conductor 220 may have various capabilities including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity (as indicated by arrow 232). Provisioning may include creating and maintaining a connection between the robot 230 and the conductor 220 (e.g., a web application). Deployment may include ensuring the correct delivery of a package version to the robot 230 assigned for execution. Configuration may include maintaining and delivering the robot environment and process configuration. Queuing may include providing management of queues and queue items. Monitoring may include tracking specific data of the robot and maintaining user permissions. Logging may include saving and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a NoSQL database) and / or another storage mechanism (e.g., ElasticSearch® which provides the ability to store large datasets and execute queries quickly). The conductor 220 may provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0061] The robot 230 is an execution agent that implements the workflow constructed by the designer 210. One commercial example of some embodiments of the robot(s) 230 is UiPath Robots™. In some embodiments, the robot 230, by default, installs the Microsoft Windows® Service Control Manager (SCM) management service. As a result, such a robot 230 can open an interactive Windows® session under a local system account and may have the rights of a Windows® service.
[0062] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, it means having the same rights as the user in which a given robot 230 is installed. This feature may also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0063] The robot 230 in some embodiments is divided into multiple components, each specialized for a specific automation task. Robot components in some embodiments include, but are not limited to, SCM management robot service, user mode robot service, executor, agent, and command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 is executed). These services are entrusted with managing the qualification information of the robot 230. The console application is launched by the SCM under the local system.
[0064] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host. The user mode robot service may be entrusted with managing the qualification information of the robot 230. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0065] The executor can execute a job given under a Windows (registered trademark) session (i.e., can execute a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows (registered trademark) Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0066] As described above, the fact that the components of the robot 230 are divided helps developers, support users, and computing systems to more easily execute, identify, and track what each component is doing. In this way, for example, special behaviors can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 can be made independent of the browser zoom level. In the case of applications marked as not recognizing or not intentionally recognizing DPI, DPI can be disabled in some embodiments.
[0067] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for the execution of the AI / ML model(s) and receive outputs therefrom via the core hyper-automation system 240.
[0068] One or more robots 230 may be listeners, as described above. These listeners can provide information to the core hyper-automation system 240 regarding what the user is doing when they use their computing system. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0069] An assistant / chatbot 250 can be provided on the user computing system to enable the user to launch an RPA local robot. The assistant can be placed, for example, in the system tray. The chatbot can have a user interface so that the user can view the text of the chatbot. Alternatively, the chatbot can run in the background without a user interface and can listen for the user's utterances using the microphone of the computing system.
[0070] In some embodiments, data labeling may be performed by a user of the computing system that the robot is executing on, or on another computing system to which the robot provides information. For example, if the robot calls an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user may draw a rectangle around the mis-identified or non-identified component and potentially provide text with the correct identification. This information can be provided to the core hyper-automation system 240 and can then be used later for training a new version of the AI / ML model.
[0071] Figure 3 is an architecture diagram showing a deployed RPA system 300 according to one or more embodiments. In some embodiments, the RPA system 300 can be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The deployed RPA system 300 can be a cloud-based system, an on-premises system, a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes, etc.
[0072] Note that either the client side 301, the server side 302, or both may include any desired number of computing systems without departing from the scope of one or more embodiments described herein. On the client side 301, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is executing a process. As shown in FIG. 3, multiple business projects can be executed simultaneously. The agent 314 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via a database server 355, an AI / ML server 360, an indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 can be a robot component.
[0073] In some embodiments, the robot represents an association between a machine name and a user name. The robot can manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots can be run simultaneously, each running in a separate Windows® session using a unique user name. This is referred to as the HD robot described above.
[0074] Agent 314 is also responsible for sending the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. The communication between Agent 314 and Conducter 340 is, in some embodiments, always initiated by Agent 314. In a notification scenario, Agent 314 may open a WebSocket channel that is later used by Conducter 330 to send commands (e.g., start, stop, etc.) to the robot.
[0075] Listener 330 monitors and records data related to user interactions with the operation of the attended computing system and / or unattended computing system in which listener 330 resides. Listener 330 can be an RPA robot, part of an operating system, a downloadable application for each computing system, or any other software and / or hardware without departing from the scope of one or more embodiments described herein. In fact, in some embodiments, the logic of the listener is implemented partially or fully via physical hardware.
[0076] In addition to the conductor 340, the server side 302 includes a presentation layer 333, a service layer 334, and a persistence layer 336. The presentation layer 333 may include a web application 342, an Open Data Protocol (OData) Representative State Transfer (REST) application programming interface (API) endpoint 344, and notifications and monitoring 346. The service layer 334 may include an API implementation / business logic 348. The persistence layer 336 may include a database server 355, an AI / ML server 360, and an indexer server 370. For example, the conductor 340 includes the web application 342, the OData REST API endpoint 344, notifications and monitoring 346, and the API implementation / business logic 348. In some embodiments, most of the actions that a user performs at the interface of the conductor 340 (e.g., via the browser 320) are performed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of one or more embodiments described herein. The web application 342 may be the visual layer of the server platform. In this embodiment, the web application 342 uses HTML and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of one or more embodiments described herein. The user interacts with the web page from the web application 342 via the browser 320 in this embodiment to perform various operations for controlling the conductor 340. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or process, start and stop robots, etc.
[0077] In addition to the web application 342, the conductor 340 also includes a service layer 334 that exposes an OData REST API endpoint 344. However, other endpoints may be included without departing from the scope of one or more embodiments described herein. The REST API is consumed by both the web application 342 and the agent 314. The agent 314 is, in this embodiment, a supervisor of one or more robots on a client computer.
[0078] The REST API of this embodiment includes configuration, logging, monitoring, and queuing functions (at least as indicated by arrow 349). The configuration endpoint may be used in some embodiments to define and configure the users, permissions, robots, assets, releases, and environments of the application. For example, the logging REST endpoint may be used to log various information such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by robots to query the version of the package to be executed when a job start command is used in the conductor 340. The queuing REST endpoint may be responsible for managing queues and queue items, such as adding data to a queue, retrieving transactions from a queue, and setting the status of a transaction.
[0079] Monitoring of the REST endpoint may monitor the web application 342 and the agent 314. The notification and monitoring API 346 may be a REST endpoint used for registering the agent 314, distributing configuration settings to the agent 314, and sending and receiving notifications from the server and the agent 314. The notification and monitoring API 346 may use WebSocket communication in some embodiments. As shown in FIG. 3, one or more activities / actions described herein are represented by arrows 350 and 351.
[0080] In some embodiments, the API of the service layer 334 can be accessed through the configuration of an appropriate API access path, for example, based on whether the conductor 340 and the overall hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The API for the conductor 340 can provide custom methods for querying statistics regarding various entities registered with the conductor 340. Each logical resource may be an OData entity in some embodiments. In such an entity, components such as, for example, robots, processes, queues may have properties, relationships, and actions. The API of the conductor 340 can be consumed by the web application 342 and / or the agent 314 in two ways in some embodiments: by obtaining API access information from the conductor 340 or by registering an external application to use the OAuth flow.
[0081] In this embodiment, the persistent layer 336 includes three servers: a database server 355 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides an AI / ML model providing service such as an AI center function), and an indexer server 370. The database server 355 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 342. The database server 355 may manage queues and queue items. In some embodiments, the database server 355 may store (in addition to or instead of the indexer server 370) messages recorded by robots. The database server 355 may also store, for example, process mining, task mining, and / or task capture-related data received from a listener 330 installed on the client side 301. Although no arrow is shown between the listener 330 and the database 355, it should be understood that in some embodiments, the listener 330 can communicate with the database 355 and vice versa. This data can be stored in the form of PDD, images, XAML files, etc. The listener 330 may be configured to eavesdrop on user actions, processes, tasks, and performance metrics on each computing system where the listener 330 resides. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, time, etc.) on its respective computing system and then convert them into a form suitable for being provided to and stored in the database server 355.
[0082] The AI / ML server 360 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can enable non-data scientists to access such capabilities. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of the AI / ML models can be monitored and trained and improved using human-verified data. The AI / ML server 360 can schedule and execute training jobs to train new versions of the AI / ML models.
[0083] The AI / ML server 360 can store data related to AI / ML models and ML packages for configuring various ML skills for users during development. The ML skills used herein are, for example, pre-built and trained ML models for processes that can be used by automation. The AI / ML server 460 can also store data related to document understanding technologies and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, natural language processing (NLP), voice analysis, different types of AI / ML models, etc.
[0084] Optionally in some embodiments, the indexer server 370 stores information recorded by robots and creates an index. In certain embodiments, the indexer server 370 may be disabled via configuration settings. In some embodiments, the indexer server 370 uses ElasticSearch®, an open-source project full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 370 via logging REST endpoint(s), where they are indexed for future use.
[0085] FIG. 4 is an architectural diagram illustrating the relationship between a designer 410, activities 420, 430, 440, 450, a driver 460, an API 470, and an AI / ML model 480, according to one or more embodiments. As shown herein, a developer uses the designer 410 to develop a workflow to be performed by a robot. Various types of activities may be presented to the developer in some embodiments. The designer 410 may be local to or remote from the user's computing system (e.g., accessed via a local web browser that interacts with a VM or a remote web server). The workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. By way of example (shown in dashed lines), the user-defined activities 420 and the API-driven activities 440 interact with an application via their APIs. The user-defined activities 420 and / or the AI / ML activities 440 may then call one or more AI / ML models 480 that may be located locally to and / or remotely from the computing system on which the robot operates, in some embodiments.
[0086] In some embodiments, non-text visual components in an image can be identified, which is referred to herein as CV. CV can be at least partially performed by an AI / ML model(s) 480. Some CV activities related to such components can include, but are not limited to, extraction of text from segmented label data using OCR, fuzzy text matching, cropping of segmented label data using ML, comparison of the extracted text in the label data with ground truth data, etc. In some embodiments, the number of activities that can be implemented in the user-defined activity 420 can be hundreds or thousands. However, any number and / or type of activities can be used without departing from the scope of one or more embodiments described herein.
[0087] The UI automation activity 450 is a subset of special low-level activities described in low-level code that facilitate interaction with the screen. The UI automation activity 450 facilitates these interactions via a driver 460 that enables the robot to interact with the desired software. For example, the driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, etc. In some embodiments, one or more AI / ML models 480 can be used by the UI automation activity 450 to perform interactions with the computing system. In certain embodiments, the AI / ML models 480 can enhance or completely replace the drivers 460. In fact, in certain embodiments, the drivers 460 are not included.
[0088] Driver 460 can interact with the OS at a low level, such as searching for hooks and monitoring keys, via the OS driver 462. Driver 460 may facilitate integration with Chrome (registered trademark), IE (registered trademark), Citrix (registered trademark), SAP (registered trademark), etc. For example, a "click" activity serves the same role in these different applications via driver 460.
[0089] Figure 5 is an architectural diagram showing a computing system 500 configured to provide an interface engine for RPA according to one or more embodiments. In some embodiments, the computing system 500 may be one or more computing systems depicted and / or described herein. In certain embodiments, the computing system 500 may be part of a hyperautomation system, such as shown in FIGS. 1 and 2. The computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to the bus 505 for processing information. The one or more processors 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The one or more processors 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one of the one or more processors 510 can be a neuromorphic circuit that includes processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require the typical components of a von Neumann computing architecture.
[0090] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by processor(s) 510. Memory 515 can be composed of random access memory (RAM), read-only memory (ROM), flash memory, cache, a static storage device such as a magnetic disk or optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media can be any available media accessible by processor(s) 510 and can include volatile media, non-volatile media, or both. Also, the media can be removable, non-removable, or both.
[0091] Furthermore, computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via a wireless and / or wired connection. In some embodiments, communication device 520 is Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), Global System for Mobile (GSM) communication, General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced (LTE-A), 802.11x, Wi-Fi, Zigbee, Ultra-WideBand (UWB), 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near-Field Communications (NFC), Fifth Generation (5G), New Radio (NR), any combination thereof, and / or may be configured to use any other currently existing or future-implemented communication standard and / or protocol without departing from the scope of one or more embodiments described herein.In some embodiments, the communication device 520 may include one or more antennas that, without departing from the scope of one or more of the embodiments described herein, are a single antenna, an array antenna, a panel antenna, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration.
[0092] The processor(s) 510 is further coupled via the bus 505 to a display 525 (i.e., a screen), such as, for example, a plasma display, a liquid crystal display (LCD), a light emitting diode (LED) display, a field emission display (FED), an organic light emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high definition display, a Retina (registered trademark) display, an in-plane switching (IPS) display, or any other suitable display for presenting information to a user. The display 525 may be configured as a touch (haptic) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-touch display, etc., using, for example, a resistive method, a capacitive method, a surface acoustic wave (SAW) capacitive method, an infrared method, an optical imaging method, a diffuse signal method, an acoustic pulse recognition method, a frustrated total internal reflection method, etc. Any suitable display device and haptic I / O may be used without departing from the scope of one or more of the embodiments described herein.
[0093] The keyboard 530 and cursor control devices 535 such as, for example, a computer mouse, touchpad, etc., are further coupled to the bus 505 to enable a user to interface with the computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user can interact with the device only via the display 525 and / or a touchpad (not shown). Any type and combination of input devices can be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, the user may interact remotely with the computing system 500 via another computing system that is communicating with it, or the computing system 500 may operate autonomously.
[0094] The memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include modules 545 (e.g., modules implementing an interface engine) configured to execute all or part of the processes described herein or derivatives thereof.
[0095] According to one or more embodiments, the module 545 may perform one or more operations such as, for example, DOM analysis, DOM text extraction, CV analysis, and schema extraction.
[0096] According to one or more embodiments, module 545 may also perform one or more operations such as, for example, detecting a framework, starting CV analysis, detecting a top layer, extracting data, determining nodes (e.g., UiNodes), determining positions, detecting element types (by DOM, framework, CV, and / or element definitions), filling in automatic anchors, filling in date and time, cleaning elements, enriching elements, creating a data structure (e.g., a DOM tree, etc.), collecting relationships (e.g., DOM relationships), collecting dropdown information, collecting geometric anchors, collecting radios, collecting values, collecting date and time groups, and determining a schema.
[0097] According to one or more embodiments, the DOM can be a programming API for documents, web pages, and other files, while the DOM tree can be regarded as the data structure, schema, or structural model of the DOM (e.g., of a web page). As an example, an attribute can be regarded as a node of the DOM (e.g., the API of a document), but not a node of the DOM tree (e.g., the structure of a document), while an element of the DOM can be provided as a node of the DOM tree that includes attributes, tags, and children.
[0098] According to one or more embodiments, module 545 also has AI / ML and RPA capabilities for performing one or more operations herein. Computing system 500 may include one or more additional functional modules 550 that include additional capabilities.
[0099] One skilled in the art will understand that the "system" can be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices, without departing from the scope of one or more embodiments described herein. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the embodiments described herein in any way, but rather to provide an example of many embodiments. In fact, the methods, systems, and devices disclosed herein may be implemented in a localized form and a distributed form that is consistent with computing techniques including cloud computing systems. The computing system may be part of, or accessible by, a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public cloud or a private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of one or more embodiments described herein.
[0100] It should be noted that some of the system features described herein are presented as modules to emphasize implementation independence more. For example, a module can be implemented as a custom very large scale integration (VLSI) circuit or a hardware circuit including off-the-shelf semiconductors such as, for example, a gate array, a logic chip, a transistor, or other individual components. Also, a module can be implemented in a programmable hardware device such as, for example, a field programmable gate array, a programmable array logic, a programmable logic device, a graphics processing unit, or other devices.
[0101] The module can also be at least partially implemented in software for execution by various types of processors. For example, a specified unit of executable code can include one or more physical or logical blocks of computer instructions that may be organized, for example, as objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and can include modules when logically combined and can include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module can be stored in a non-transitory computer-readable medium such as, for example, a hard disk drive, a flash device, RAM, a tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of one or more embodiments described herein.
[0102] In fact, a module of executable code can be a single instruction, or multiple instructions, and can even be distributed among multiple different code segments, between different programs, and among multiple memory devices. Similarly, the operational data can be specified within the module, can be shown here, can be embodied in any suitable form, and can be organized within any suitable type of data structure. The operational data can be collected as a single data set or can be distributed in different locations across different storage devices and can exist, at least in part, simply as electronic signals on a system or network.
[0103] Without departing from the scope of one or more embodiments described herein, various types of AI / ML models can be trained and deployed. For example, FIG. 6 shows an example of a neural network 600 trained to recognize graphical elements within an image according to one or more embodiments. Here, the neural network 600 receives pixels (shown in column 610) of a 1920×1080 screen screenshot image as inputs for the input “neurons” 1 to I of the input layer (shown in column 620). In this case, I is 2,073,600, which is the total number of pixels in the screenshot image.
[0104] The neural network 600 also includes a number of hidden layers (represented in columns 630 and 640). Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) typically have multiple layers, but an SLNN may sometimes have only one or two layers and is usually less than that of a DLNN. Typically, a neural network architecture includes an input layer, multiple intermediate layers (e.g., hidden layers), and an output layer (represented in column 650), as in the case of the neural network 600.
[0105] Often, a DLNN has many layers (such as 10, 50, 200, etc.), and subsequent layers usually reuse the functions from the previous layers to compute more complex and general functions. On the other hand, an SLNN has only a few layers and tends to be trained relatively quickly because expert functions are pre-created from raw data samples. However, feature extraction is cumbersome. On the other hand, a DLNN usually does not require expert functions but takes time to train and tends to have more layers.
[0106] In either approach, the layers are trained simultaneously on the training set and usually checked for overfitting on a separate cross-validation set. Excellent results are obtained with both techniques, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem being addressed by each neural network.
[0107] Returning to FIG. 6, the pixels provided as the input layer are supplied as inputs to the J neurons of hidden layer 1. In this example, all pixels are supplied to each neuron, but not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks not departing from the scope of one or more embodiments described herein. Various architectures can be used, either individually or in combination.
[0108] Hidden layer 2 (630) receives inputs from hidden layer 1 (620), hidden layer 3 receives inputs from hidden layer 2 (630), and so on, up to the last hidden layer (represented by ellipse 655) providing its output as the input to the output layer. The number of neurons I, J, K, and L are not necessarily equal, and thus it should be noted that any desired number of layers can be used for a given layer of neural network 600 without departing from the scope of one or more embodiments described herein. In fact, in certain embodiments, the types of neurons in a given layer may not all be the same.
[0109] The neural network 600 is trained to assign a confidence score to graphical elements that are thought to be found within an image. To reduce matches with unacceptably low likelihoods, in some embodiments, only those results having a confidence score above a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs having a confidence score exceeding this amount may be used and the rest may be ignored. In this case, the output layer indicates that two text fields (represented by outputs 661 and 662), a text label (represented by output 663), and a send button (represented by output 665) have been found. The neural network 600 may provide the positions, dimensions, images, and / or confidence scores of these elements without departing from the scope of one or more embodiments described herein, which may then be used by an RPA robot or another process that uses this output for a given purpose.
[0110] Note that a neural network is typically a probabilistic construct that has a confidence score. This can be a score learned by an AI / ML model based on the frequency with which similar inputs were correctly identified during training. For example, a text field often has a rectangular shape and a white background. A neural network can learn to identify graphical elements having these features with a high degree of confidence. Common types of confidence scores include a decimal number between 0 and 1 (interpretable as a percentage of confidence), a number between negative infinity and positive infinity, or a set of expressions (e.g., "low", "medium", and "high"). Also, as an attempt to obtain a more accurate confidence score, various post-processing calibration techniques such as temperature scaling, batch normalization, weight decay, negative log likelihood (NLL), etc. may be employed.
[0111] The "neurons" of a neural network are usually mathematical functions based on the functions of biological neurons. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass the output to the next layer. This activation function can be a non-linear thresholded activity function that does nothing if the value is below the threshold, and responds linearly when the function exceeds the threshold (i.e., rectified linear unit (ReLU) non-linearity). Since actual neurons can have a nearly identical activity function, the sum function and ReLU function are used in deep learning. Through linear transformation, information can be subtracted, added, etc. Essentially, neurons function as gating functions that pass the output to the next layer governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some of the neurons.
[0112] JPEG2025077955000002.jpg75162
[0113] JPEG2025077955000003.jpg52161
[0114] JPEG2025077955000004.jpg21153
[0115] In this case, neuron 700 is a single-layer perceptron. However, without departing from the scope of one or more embodiments described herein, any suitable neuron type or combination of neuron types can be used. It should also be noted that the range of values of the weights and / or output value(s) of the activation function can be different in some embodiments without departing from the scope of one or more embodiments described herein.
[0116] For example, in the case where the identification of graphical elements within an image is successful, a goal, or "reward function", is often used. The reward function guides the search in the state space and attempts to achieve the goal (e.g., successful identification of graphical elements, successful identification of the next sequence of activities in an RPA workflow, etc.) by using both short-term and long-term rewards to explore intermediate transitions and steps.
[0117] During training, various labeled data (in this case, images) are supplied through the neural network 600. When the identification is successful, the weights of the inputs to the neurons are strengthened, while when the identification fails, those weights are weakened. A cost function such as the mean squared error (MSE) or gradient descent can be used to make slightly incorrect predictions cost much less than greatly incorrect predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, the data scientist can change the reward function, indicate where un-identified graphical elements are, provide corrections for mis-identified graphical elements, etc.
[0118] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "pop up" the hidden layers of a neural network to see how much loss each node is bearing, and then assign low weights to nodes with a high error rate and vice versa to update the weights to minimize the loss. That is, backpropagation enables the data scientist to repeatedly adjust the weights so as to minimize the difference between the actual output and the desired output.
[0119] The algorithm of backpropagation is mathematically based on optimization theory. In supervised learning, training data with known outputs is passed through the neural network, the error is calculated using a cost function from the known target outputs, and this gives the error for backpropagation. The error is calculated at the output, and this error is converted into a correction of the network's weights that minimizes the error.
[0120] JPEG2025077955000005.jpg60162
[0121] JPEG2025077955000006.jpg23162
[0122] JPEG2025077955000007.jpg112164
[0123] JPEG2025077955000008.jpg189162
[0124] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., above 97% using an F2 or F4 threshold for detection, about 2,000 epochs). This level of accuracy can, in some embodiments, be determined using an F1 score, an F2 score, an F4 score, or any other suitable technique that does not deviate from the scope of one or more of the embodiments described herein. Once trained on training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps ensure that the AI / ML model does not "overfit" such that it performs well on the graphical elements in the training data but does not generalize well to other images.
[0125] In some embodiments, it may not be known what level of accuracy the AI / ML model can achieve. Thus, if the accuracy of the AI / ML model begins to decline when analyzing the evaluation data (i.e., the model performs well on the training data but its performance begins to degrade on the evaluation data), the AI / ML model can undergo additional training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is only deployed if it reaches a certain level of accuracy or if the accuracy of the trained AI / ML model is better than an existing deployed AI / ML model.
[0126] In certain embodiments, for example, the collection of trained AI / ML models can be used to perform tasks such as employing an AI / ML model for each type of target graphical element, performing OCR by employing an AI / ML model, deploying yet another AI / ML model to recognize proximity relationships between graphical elements, and employing yet another AI / ML model to generate an RPA workflow based on the output from other AI / ML models. For example, this can enable semantic automation collectively by the AI / ML models.
[0127] In some embodiments, transformer networks such as SentenceTransformers™, a Python™ framework, can be used for state-of-the-art sentence, text, and image embedding. Such transformer networks learn associations of words and phrases with both high and low scores. This trains the AI / ML model to determine what is close to the input and what is not, respectively. Instead of using only word / phrase pairs, the transformer network may also use field length and field type.
[0128] FIG. 8 is a flowchart showing a process 800 for training an AI / ML model(s) according to one or more embodiments. Note that process 800 can also be applied to other UI learning operations such as, for example, NLP and chatbots. The process begins, for example, by training, at block 810, on data that provides labeled data such as, for example, labeled screens (with, e.g., specified graphical elements and text), words and phrases, a "thesaurus" of semantic relatedness between words and phrases such that similar words and phrases can be identified for a given word or phrase, as shown in FIG. 8. The nature of the training data provided can depend on the purpose the AI / ML model is to achieve. The AI / ML model is then trained, at block 820, over multiple epochs, and the results are reviewed at block 830.
[0129] If the AI / ML model does not meet the desired confidence threshold at decision block 840 (process 800 proceeds according to the no arrow), the training data is supplemented and / or the reward function is modified, at block 850, to help the AI / ML model better achieve its purpose, and the process returns to block 820. If the AI / ML model meets the confidence threshold at decision block 840 (process 800 proceeds according to the yes arrow), the AI / ML model is tested against evaluation data, at block 860, to confirm that the AI / ML model generalizes well and does not overfit to the training data. The evaluation data may include screens, source data, etc. that the AI / ML model has not previously processed. If the confidence threshold of the evaluation data is met at decision block 870 (process 800 proceeds according to the yes arrow), the AI / ML model is deployed, at block 880. Otherwise (process 800 proceeds according to the no arrow), the process returns to block 880 and the AI / ML model is further trained.
[0130] FIG. 9 is a flowchart showing a process 900 for providing an interface engine according to one or more embodiments. Generally, process 900 provides exemplary operations of an interface engine that detects UI elements not displayed on a display (e.g., display 525 of FIG. 5, or a screen presenting a UI including UI elements).
[0131] The process 800 executed in FIG. 8 and the process 900 executed in FIG. 9 may be executed by a computer program encoding instructions for a processor(s) according to one or more embodiments. The computer program may be stored on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, a tape, and / or other such medium or combination of media used to store data. The computer program may include encoded instructions for controlling a processor(s) of a computing system (e.g., processor(s) 510 of computing system 500 of FIG. 5) to implement all or part of the process steps described in FIGS. 8-9, which may also be stored on a computer-readable medium.
[0132] Process 900 begins at block 910 where a screen presents a UI. The screen is the display described herein. The UI can include one or more web browser windows or interface frames, and icons, toolbars, etc.
[0133] The UI may include a page. The page may be a web page within one of a web browser window or an interface frame. The page may have one or more layers that create a stacking structure. The web browser window or interface frame can be expanded to the edge of the screen or occupy that portion. The web browser window or interface frame may include a scroll bar that enables browsing of the page. The page can be made larger than the boundaries of the web browser or interface frame. The page can be made larger than the screen. As a result, the visible area of the page is bounded by the boundaries of the web browser window or interface frame (regardless of whether the web browser window or interface frame is fully expanded to the edge of the screen), and the page can be partially displayed on the screen such that the invisible area of the page comes into view by operating the scroll bar or the screen. For ease of explanation, with respect to process 900, the boundaries of the web browser window or interface frame and the edge of the screen are contemporaneous.
[0134] The page may include a long form. The page and / or the long form may include UI elements. The page, the long form, and the UI elements may be identified, itemized, and described by the DOM. The DOM is a cross-platform and language-independent mechanism that provides UI elements as a tree structure where each node represents a part of the page and / or the long form. In an example, the HTML structure of each UI element may be included in the DOM. The UI elements may include input fields such as, for example, dropdowns, checkboxes, radio buttons, and other interactive elements. Each UI element may be identified by a type such as, for example, a DOM type, a framework type, a CV type, or an element definition type. Thus, the type corresponding to the input field can be found within the DOM by the interface engine.
[0135] In block 930, the interface engine performs DOM text extraction. DOM text extraction includes cases where the interface engine analyzes the DOM of a page and determines UI elements from the DOM text therein that may be related to auto-completion or filling.
[0136] According to one or more embodiments, the interface engine analyzes the DOM text of the DOM of a page to find all nodes for UI elements. The interface engine extracts the DOM text associated with each node corresponding to a UI element in the DOM for possible related elements. For example, the related elements can be input fields of a long form and their respective labels to the input fields. Note that the problem with the DOM of a page is that it may not be clear what an "input field" is and what a "container" is. The interface engine solves this problem by automatically filtering out and removing all unnecessary context during the analysis / extraction of the DOM and capturing only the relevant text.
[0137] According to one or more embodiments, the interface engine generates a filtered DOM. The filtered DOM includes filtered elements, for example, one or more related UI elements determined from possible related elements. For example, when the first related element is determined, the interface engine generates (i.e., creates) a filtered DOM (e.g., an element list). When each subsequent related element is determined, the interface engine constructs (i.e., adds to) the element list. Thus, the element list is an itemization of one or more related UI elements using the corresponding text from the DOM. The corresponding text can include, but is not limited to, the depth hierarchy, geometric layout, type, and text values of the filtered elements.
[0138] According to one or more embodiments, the interface engine may analyze a page using an ML model, such as an anchoring / relation model. For example, depth hierarchy, geometric layout, type, and text values may be aspects of an anchoring / relation model that determine DOM text (e.g., HTML structure) associated with various input fields. The anchoring / relation model may determine the associated DOM text, as further described herein.
[0139] In block 950, the interface engine performs CV analysis. The CV analysis is performed on the visible region. According to one or more embodiments, the interface engine performs CV analysis to determine the type of one or more associated UI elements within the visible region. For example, to determine the type of an input field, the visible region of the page is captured by the interface engine as a screenshot. The interface engine positions the screenshot through CV analysis and identifies the type of all input fields on the visible region.
[0140] Turning to FIG. 10, an interface 1000 according to one or more embodiments is shown. Interface 1000 is an example of a page shown within the boundaries of an interface frame contemporaneous with the edge of the screen. Thus, the visible region 1001 of the page of interface 1000 is within the boundaries of the interface frame and the edge of the screen, while the invisible region is not shown. Interface 1000 includes a plurality of elements, such as elements 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, and 1019. Note that element 1018 is a scroll bar that may be manipulated to move a non-visible region of the page into a part of the visible region 1001.
[0141] According to one or more embodiments, the interface engine identifies some of the plurality of elements as unnecessary context. For example, element 1019 (i.e., the text "Assign to me") has been previously filtered. According to one or more embodiments, the interface engine identifies some of the plurality of elements as related UI elements or input fields. For example, element 1012 (i.e., the "Issue type" field) requires type specification. According to one or more embodiments, the interface engine may use the list of elements together with the filtered DOM to identify the related UI elements of the visible region 1001.
[0142] The interface engine performs CV analysis on the visible region 1001 to identify the types of these related UI elements. As an example, interface 1000 shows that element 1012 can be determined by the interface engine as a dropdown type from the CV analysis. Further, the interface engine determines the structure of element 1012 from the filtered DOM.
[0143] Returning to FIG. 9, as represented by sub-block 955, the CV analysis can be stored in the cache by the interface engine. The cache can be, for example, a part of the memory within the memory 515 of FIG. 5. The interface engine can use the information in the cache (e.g., when a similar long form is shown as the destination form on the next page (e.g., a second page that can be in a different interface frame than the current page)). The technical effects, advantages, and benefits of the interface engine include that the interface engine uses the cache to infer and identify the input fields of the invisible region and other pages / destination forms without performing additional CV analysis.
[0144] In block 970, the interface engine performs an estimation. Through the estimation by the interface engine, the interface engine can determine (e.g., identify and understand) the types of existing off-screen elements. The interface engine estimates one or more target UI elements in the invisible region. The one or more target UI elements are a subset of one or more UI elements of a long form or page. The invisible region can be a part of a long form or page outside the web browser window or interface frame. The estimation of the one or more target UI elements by the interface engine includes searching for similar HTML structures from the DOM or the filtered DOM.
[0145] According to one or more embodiments, the ML model of the interface engine determines the HTML structure for a specified input field in the visible region and uses that HTML structure to find similar HTML structures within the invisible region. Next, using the ML model, the interface engine estimates the relevant UI elements in the visible region as target UI elements in the invisible region. As an example, the interface engine determines the structure of the input field from the visible region using relevant DOM-extracted text and CV analysis. The structure of the input field can include the classes and attributes of the HTML portion related to the field identified from the relevant DOM-extracted text and CV analysis. Next, the relevant DOM-extracted text is searched for similar structures to identify other similar input fields within the invisible region of the screen. For example, if the CV analysis determines that an element on the visible region of the screen is of the dropdown type of the input field, the corresponding HTML structure of the input field is determined from the filtered DOM text. The HTML structure related to the dropdown type is searched from the DOM text related to the invisible region, and thus, all dropdown-type input fields within the invisible region are identified.
[0146] In block 980, the interface engine performs schema extraction. According to one or more embodiments, schema extraction enables a user to anchor specific types of elements to interaction-capable elements. In this way, through anchoring, the interface engine links DOM text to the corresponding types of relevant UI elements determined by CV analysis that connects or creates associations between fields and labels.
[0147] In an example regarding the ML model of the interface engine, the extracted DOM and the types of all input fields are input into the ML model. The ML model of the interface engine determines the DOM text (e.g., HTML structure) associated with the various input fields. To determine the relevant DOM text, the interface engine embeds information by encoding and incorporating transformation and scale-invariant features of the ML model (note that the features can be based on information already extracted by the interface engine). Next, transformation and scale-invariant features, such as appropriate attention, convolution, fully connected, etc., are fed through the various layers of the ML model. The predictor of the ML model calculates logits that assign an appropriate relationship between the input fields and the labels.
[0148] Turning to FIG. 11, an interface 1100 according to one or more embodiments is shown. Interface 1100 is an example of a page shown within the boundaries of an interface frame contemporaneous with the edge of the screen. Thus, the visible region 1101 of the page of interface 1100 is within the boundaries of the interface frame and the edge of the screen, while the invisible region is not shown. Interface 1100 includes a plurality of elements, such as elements 1122, 1123, 1124, 1125, 1126, 1127, and 1128. Note that element 1128 is a scroll bar that can be manipulated to move the invisible region of the page into a part of the visible region 1101. Note that before manipulating element 1128, elements 1122, 1123, 1124, 1125, 1126, 1127, and 1128 were not part of the visible region 1001 of FIG. 10. Thus, the interface engine identifies an element (e.g., element 1140) from the invisible region of interface 1000 of FIG. 10 having an HTML structure similar to element 1020 of FIG. 10, and then, as an example, scrolls and displays it. In this regard, interface 1100 indicates that element 1140 is identified as a drop-down type input field.
[0149] Turning to FIG. 12, an exemplary code 1200 according to one or more embodiments is shown. Code example 1200 is an example of the HTML structure of the DOM of a drop-down field in the visible region.
[0150] Problems can arise when there is no similar input field in the visible region for UI elements in the invisible region. In conventional clipboard technology, the user had to scroll to the input field in the invisible region to take a screenshot for CV analysis to determine the input field type. According to one or more embodiments, scrolling can be avoided by the interface engine, and the estimation of the structural elements can be completed using a cache. Thus, the interface solves the problem that UI elements in the invisible region do not have a similar input field in the visible region.
[0151] Returning to FIG. 9, in block 990, the interface engine inputs data. The interface engine extracts data from a source for attachment to at least one or more target UI elements. According to one or more embodiments, the interface engine inputs data from the source to one or more target UI elements. That is, when a field is identified, various methods may be used by the interface engine to type data into the input field, for example, using RPA. Data entry can be completed by the interface engine regardless of whether there is scrolling of the page. The interface engine may receive a scrolling option so as to be able to display the automation of data entry. Data can be entered into the same long form and / or a similar long form (e.g., a destination form) on subsequent pages.
[0152] FIG. 13 shows a flowchart illustrating a process 1300 according to one or more embodiments. The process 1300 executed at 13 can be executed by a computer program encoding instructions for a processor(s) according to one or more embodiments. Generally, the process 1300 provides an exemplary operation of an interface engine that detects UI elements not displayed on a display (e.g., the display 525 of FIG. 5, or a screen presenting a UI including UI elements). To facilitate the description of the process 1300, a page (e.g., a web page) is represented by a UI presented on a display, and the page is larger than the UI. Thus, the page has an invisible region and a visible region.
[0153] Process 1300 begins with block 1302 where the interface engine detects the framework. According to one or more embodiments, the interface engine may identify the framework by placing "detect-Framework.ts" in a page (e.g., PageWorld). One or more examples of frameworks include, but are not limited to, SAP and Workday.
[0154] In block 1306, the interface engine starts the CV analysis. According to one or more embodiments, the interface engine may start the CV analysis on the visible content of the page. The CV analysis may be used for detecting element types by CV.
[0155] In block 1310, the interface engine detects one or more layers. Note that the page may have multiple layers forming a stacking structure. For example, interacting with one UI element (e.g., a non-hidden element) may reveal other UI elements (e.g., hidden elements). Further, since only the UI elements in the topmost layer are displayed by the UI, not all elements (e.g., covered elements) can be captured by the CV analysis. For example, according to one or more embodiments, the interface engine may interact with UI elements covered by other elements (in other detected layers). The remaining operations of process 1300 may be applied to each layer detected by the interface engine and may be repeated or looped as needed by the interface engine. One or more of the remaining operations of process 1300 may be executed by the interface engine and may be executed in any order.
[0156] In block 1314, the interface engine extracts data. According to one or more embodiments, the interface engine may extract all non-hidden elements from the DOM using "semantic-extractData.ts". The interface engine may create an array of WebElementData that can be used for DOM tree and schema calculations. According to one or more embodiments, the interface engine may extract data by sequentially refining the DOM tree of the page to reach the schema. The schema may include mappings between anchoring text and controls (e.g., input fields, dropdowns, etc.). Examples of schemas include: "First Name" -> input field 1; "Last Name" -> input field 2; "Country" -> dropdown 1; etc. Schema elements may include input elements where new functionality can be displayed within the UI by interaction.
[0157] In block 1318, the interface engine determines nodes (e.g., UiNode). According to one or more embodiments, the interface engine can obtain the UiNodes of all WebElementData.
[0158] In block 1322, the interface engine determines positions. According to one or more embodiments, the interface engine can obtain the positions of all WebElementData.
[0159] In block 1330, the interface engine determines element types by the DOM. According to one or more embodiments, the interface engine can detect element types using information from the DOM, such as tags, types, roles, and aria attributes.
[0160] In block 1334, the interface engine determines the element type by the framework. According to one or more embodiments, the interface engine may detect the element type using framework-specific rules. The framework-specific rules define a series of conditions for searching or identifying elements. According to one or more embodiments, and by way of example, all websites have a DOM tree. Further, each element of the DOM tree has tags and attributes, and in some cases, one or more children. The framework-specific rules define a series of conditions for searching or identifying elements based on tags, attributes, or children. For example, the first rule identifies elements with a specific tag (e.g., x = input), and the second rule may identify elements with specific attributes (e.g., data automation identification by a specific value, input box, date selection, etc.). Many of the framework-specific rules are managed by the page / document. If there are 10 different control types on the page / document, the framework-specific rules may include 10 different rules.
[0161] In block 1338, the interface engine determines the element type by the CV. According to one or more embodiments, the interface engine can detect the element type using information from the CV analysis. Further, according to one or more embodiments, the interface engine may create an element definition for the CV element. The element definition is specific to the design of the page and may be composed of all tags, attributes, children, and child tags and attributes. The interface engine may match what is within the bounding box (e.g., the boundary around the element displayed during the design process) with the elements on the page.
[0162] In block 1342, the interface engine determines the element type by the element definition. According to one or more embodiments, the interface engine may detect the element type using the element definition. Note that the CV operates at the screenshot level, but the page may have off-screen elements. The interface engine operates off-screen to identify other elements such as elements within the view based on the tags and attributes of the elements.
[0163] In block 1346, the interface engine fills in the auto anchor. According to one or more embodiments, the interface engine can detect the anchor of an element using DOM element attributes and then provide values corresponding to one or more regions of the page. DOM element attributes can include, but are not limited to, label[for], aria-labelled by, aria-label, and placeholder.
[0164] In block 1350, the interface engine enters the date and time. According to one or more embodiments, the interface engine can detect the date and time type, date and time component type, and date format based on attributes, classes, and values, and provide values corresponding to one or more regions of the page.
[0165] In block 1358, the interface engine cleans the elements. According to one or more embodiments, the interface engine can remove duplicate elements that refer to the same logical element (e.g., an input element can be detected in both the CV and the DOM).
[0166] In block 1362, the interface engine enriches elements. According to one or more embodiments, the interface engine can facilitate information from any internal input element to a parent (container) element (for example, CV can detect a DIV element representing a text box). As an example, during FillElementTypeByCV, the DIV element can be promoted to an InputBox. The DIV element can include a hidden <input> with an aria-labelledby attribute. During the enrichment of UI elements, an anchor can be promoted to a container element. Note that blocks 1302 - 1362 can be executed for each layer and / or each iframe within any layer. The resulting DOMRawUiNodeElements can be aggregated later.
[0167] In block 1366, the interface engine creates a data structure such as, for example, a DOM tree. According to one or more embodiments, the interface engine can generate a simplified DOM tree structure from an aggregated array of DOMRawUiNodeElements.
[0168] In block 1370, the interface engine collects DOM relationships. According to one or more embodiments, the interface engine can detect the anchors of elements using a DOM relationship model (for example, the relationship anchor generated an ML model). The DOM relationship model is a custom ML model of the interface engine that finds the anchors of input elements from a simplified DOM tree structure. Thus, the interface engine can infer a "schema" from the simplified DOM using the custom ML model.
[0169] In block 1374, the interface engine collects dropdown information. According to one or more embodiments, the interface engine can detect the relative position of the dropdown arrow and transfer it to the DropdownInfo property of the DOMRawUiNodeElement.
[0170] In block 1378, the interface engine collects geometric anchors. According to one or more embodiments, an anchor is a label for an input. When the interface engine detects something in the region (e.g., an input box), the interface engine labels that something (e.g., the input box). An anchor, which is an automatic anchor or auto-anchor, can be inferred from the DOM since the DOM has text that helps with labeling. A machine learning anchor includes the case where the interface engine provides the DOM of the page to a model and the model determines which input the anchor is for. A shape anchor is a label determined by using position alignment with the DOM and elements (e.g., having fixed rules). For an input UI element without an anchor, if possible, a shape anchor is assigned by the interface engine. Further, the shape anchor can be searched using position-based, alignment-based, and hierarchical-based algorithms of the interface engine (e.g., the DOM tree can include a container, text, and an input box within the container). For example, the position-based algorithm of the interface engine performs a process of searching for the position of the input UI element in the page so that a shape anchor can be assigned.
[0171] In block 1382, the interface engine collects radio groups. According to one or more embodiments, the interface engine can group all radio button elements into a group and search for the label of the group.
[0172] In block 1386, the interface engine collects values. According to one or more embodiments, for each input element (e.g., dropdown, chat box, radio, and anything that can interact to set a value), the value is collected by the interface engine. For example, the interface engine obtains the content entered in the name input box. According to one or more embodiments, the interface engine may input the value property of the DOMRawUiNodeElement.
[0173] In block 1390, the interface engine collects a date and time group. For example, the date and time group can be a set of alphanumeric characters in a predetermined format representing the year, month, day of the month, hour of the day, minute of the hour, and one or more of the time zones. According to one or more embodiments, the interface engine can group all input elements of the date and time group.
[0174] In block 1394, the interface engine determines a schema. According to one or more embodiments, the interface engine can create a schema from DOMRawUiNodeElements. Note that the schema can include a mapping between the anchor text and the control as described herein.
[0175] According to one or more embodiments, the processes 900 and 1300 described herein can be implemented as activities in a designer application, such as UiPath Studio (trademark). The activity can be used for semantic copy and paste from a set of source fields (e.g., in the visible / invisible area of the screen) to a set of destination fields. Further, all caches captured during CV analysis can also be supplied to an activity for dynamically detecting UI elements in the invisible area. Further, the processes 900 and 1300 described herein can be implemented with respect to task mining.
[0176] A computer program can be implemented in hardware, software, or a hybrid implementation. The computer program can be composed of modules that communicate operably with each other and is designed to send information or instructions to a display. The computer program can be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0177] It is readily understood that the components of the various embodiments may be arranged and designed in a variety of different configurations as generally described and illustrated herein. Accordingly, the detailed description of the embodiments as represented in the accompanying figures is not intended to limit the scope as claimed, but is merely representative of the selected embodiments.
[0178] The features, structures, or characteristics described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments", "some embodiments", or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment. Thus, the appearances of "in certain embodiments", "in some embodiments", "in other embodiments", or similar language throughout this specification are not necessarily referring to the same group of all embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0179] References throughout this specification to features, advantages, or similar language do not imply that all of the features and advantages that can be realized should be, or are in, any single embodiment. Rather, the language referring to the features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in one or more embodiments. Thus, discussions of the features and advantages throughout this specification, as well as similar language, may refer to the same embodiment, but do not necessarily refer to the same embodiment.
[0180] Furthermore, the features, advantages, and characteristics of one or more embodiments described herein can be combined in any suitable manner. One of ordinary skill in the relevant art will recognize that the disclosure can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments but not in all embodiments.
[0181] One of ordinary skill in the art will readily understand that the disclosure can be practiced using steps in a different order and / or using hardware elements in a configuration different from that disclosed. Accordingly, although the disclosure has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations will become apparent while remaining within the spirit and scope of the disclosure. Therefore, reference should be made to the appended claims to determine the scope of the disclosure.
Claims
1. 1. A method for detecting and automatically filling in one or more user interface (UI) elements of a page that is not being displayed on a screen, the method being performed by an interface engine implemented as a computer program within a computing environment, the method comprising: analyzing, by the interface engine, a Document Object Model (DOM) of the page to extract DOM text for one or more associated UI elements of the one or more UI elements; performing a computer vision (CV) analysis by the interface engine to determine one or more types of the one or more associated UI elements; inferring, by the interface engine, one or more target UI elements in the invisible area of the page from the one or more associated UI elements.
2. 2. The method of claim 1, wherein the method includes linking the DOM text to the one or more target UI elements based on the one or more types and the one or more associated UI elements determined by the CV analysis.
3. The method of claim 1 , wherein the inference of the one or more target UI elements by the interface engine includes searching the DOM for similar HyperText Markup Language (HTML) constructs.
4. The method of claim 1 , wherein the one or more UI elements include one or more input fields on the page.
5. The method of claim 1 , wherein the interface engine extracts data from a source for application to at least the one or more target UI elements.
6. The method of claim 1 , wherein the one or more types of the one or more associated UI elements include a DOM type, a framework type, a CV type, or an element definition type.
7. The method of claim 1 , wherein the execution of the CV by the interface engine includes capturing a viewable area of the page and identifying the one or more types of the one or more associated UI elements within the viewable area.
8. The method of claim 1 , wherein the interface engine inputs data into the one or more target UI elements without scrolling the page.
9. The method of claim 1 , wherein the page includes a long form.
10. The method of claim 1 , wherein the non-visible area of the page includes a subsequent destination form.
11. 1. A system comprising: a memory storing code for an interface engine for detecting and automatically filling in one or more user interface (UI) elements of a page that is not displayed on a screen; In the system, analyzing, by the interface engine, a Document Object Model (DOM) of the page to extract DOM text for one or more associated UI elements of the one or more UI elements; performing a computer vision (CV) analysis by the interface engine to determine one or more types of the one or more associated UI elements; at least one processor configured to execute the code that causes the interface engine to infer one or more target UI elements in an invisible area of the page from the one or more associated UI elements.
12. 12. The system of claim 11, wherein the interface engine links the DOM text to the one or more target UI elements based on the one or more types and the one or more associated UI elements determined by the CV analysis.
13. The system of claim 11 , wherein the inference of the one or more target UI elements by the interface engine includes searching the DOM for similar HyperText Markup Language (HTML) constructs.
14. The system of claim 11 , wherein the one or more UI elements include one or more input fields on the page.
15. The system of claim 11 , wherein the interface engine extracts data from a source for application to at least the one or more target UI elements.
16. The system of claim 11 , wherein the one or more types of the one or more associated UI elements include a DOM type, a framework type, a CV type, or an element definition type.
17. The system of claim 11 , wherein the execution of the CV by the interface engine includes capturing a viewable area of the page and identifying the one or more types of the one or more associated UI elements within the viewable area.
18. The system of claim 11 , wherein the interface engine inputs data into the one or more target UI elements without scrolling the page.
19. The system of claim 11 , wherein the page includes a long form.
20. The system of claim 11 , wherein the non-visible area of the page includes a subsequent destination form.