User interface automation using robotic process automation to detect invisible user interface elements on display and populate forms

By analyzing the DOM and using CV analysis technology, automatically detecting and filling the invisible form input fields on the screen, the inefficiency problem in the existing technology is solved and efficient UI automation is achieved.

CN119940314APending Publication Date: 2025-05-06UIPATH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410640309.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-06
Filing Date
2024-05-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to automatically detect and fill in form input fields that are not visible on the screen, resulting in users needing to repeatedly scroll and manually paste data, which is inefficient.

Method used

The DOM text of the relevant UI elements is extracted by analyzing the page's document object model (DOM), and the type of UI elements is determined in combination with computer vision (CV) analysis, thereby extrapolation of the target UI elements in the invisible area of ​​the page and automatically populating it.

Benefits of technology

It enables the recognition and fill in invisible UI elements on the screen without scrolling or moving the page, improving the efficiency and accuracy of UI automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940314A_ABST
    Figure CN119940314A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to user interface automation using robotic process automation to detect invisible user interface elements on a display and populate a form. A method is provided. The method is used for detecting and automatically populating user interface (UI) elements of a page that is not visible on a screen. The method is performed by an interface engine implemented as a computer program within a computing environment. The method includes analyzing a Document Object Model (DOM) of the page to extract DOM text for related ones of the UI elements. The method includes performing computer vision (CV) analysis to determine a type of relevant UI element, and extrapolating a target UI element of an invisible region of the page from the relevant UI element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to automation, and more particularly to an interface engine that uses one or more Robotic Process Automations (RPAs) to detect and fill out forms, thereby providing user interface (UI) automation. Background Art

[0002] Typically, a long form may be presented in a UI such that portions of the long form are both visible and invisible on the screen. Conventional clipboard techniques use computer vision (CV)-based detection to perform destination analysis of the screen to determine input fields on the screen. In this manner, conventional clipboard techniques focus on visible input fields and fail to detect input fields that are not visible on the screen.

[0003] Another problem with conventional clipboard techniques is that when a user wants to paste source field data into a long form using conventional clipboard techniques, the user will only be able to paste the source field data into the visible input fields of the long form. Conversely, input fields that are not visible on the screen (e.g., visible only when manually scrolled to) cannot be mapped with source field data using conventional clipboard techniques. This problem becomes more complex when the user repeatedly partially scrolls and repeatedly uses conventional clipboard techniques to map new iterations of visible input fields on the screen.

[0004] Need a solution to fill out and / or complete long forms in automation. Summary of the invention

[0005] According to one or more embodiments, a method is provided. The method is used to detect and automatically fill one or more user interface (UI) elements of a page that is not visible on the screen. The method is performed by an interface engine implemented as a computer program within a computing environment. The method includes analyzing a document object model (DOM) of a page to extract DOM text of one or more related UI elements in one or more UI elements. The method includes performing a computer vision (CV) analysis to determine one or more types of one or more related UI elements, and extrapolating one or more target UI elements of an invisible area of ​​the page from the one or more related UI elements.

[0006] According to one or more embodiments, the above method may be implemented in a system, a computer product, and a device. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to easily understand the advantages of certain embodiments herein, a more specific description will be given with reference to specific embodiments shown in the accompanying drawings. Although it should be understood that these drawings depict only typical embodiments and are not therefore to be considered limiting of the scope thereof, one or more embodiments herein will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:

[0008] Figure 1 depicts an architectural diagram illustrating an automation system according to one or more embodiments;

[0009] Figure 2 depicts an architectural diagram illustrating an RPA system according to one or more embodiments;

[0010] Figure 3 depicts an architectural diagram of an RPA system illustrating a deployment according to one or more embodiments;

[0011] Figure 4 depicts an architectural diagram illustrating relationships between designers, activities, and drivers according to one or more embodiments;

[0012] Figure 5 depicts an architectural diagram illustrating a computing system in accordance with one or more embodiments;

[0013] Figure 6 shows an example of a neural network that has been trained to recognize graphical elements in an image in accordance with one or more embodiments;

[0014] Figure 7 illustrates an example of a neuron according to one or more embodiments;

[0015] Figure 8 depicts a flowchart illustrating a process for training AI / ML model(s) according to one or more embodiments;

[0016] Fig. 9 depicts a flowchart illustrating a process according to one or more embodiments;

[0017] Fig.10 depicts an interface according to one or more embodiments;

[0018] Fig.11 depicts an interface according to one or more embodiments;

[0019] Fig.12 depicts example code according to one or more embodiments; and

[0020] Fig.13 A flow chart illustrating a process in accordance with one or more embodiments is depicted.

[0021] Unless otherwise stated, similar reference numerals denote corresponding features consistently throughout the drawings. DETAILED DESCRIPTION

[0022] The present disclosure relates generally to automation, and more specifically to an interface engine that uses one or more robotic process automation (RPA) to fill in forms, thereby providing user interface (UI) automation. For example, the interface engine uses one or more RPAs to provide UI automation, which detects UI elements that are not visible on the display (e.g., outside the visible area of ​​the UI presented on the screen) and fills / completes the form therein. The UI elements may include, but are not limited to, one or more input fields of a form that are not visible in the visible area of ​​the UI presented on the screen. The filling / completion of one or more input fields can be performed by the interface engine without moving or scrolling the UI presented on the screen. Implementing the interface engine can be performed by a computing system (as described herein).

[0023] According to one or more embodiments, the interface engine performs one or more operations while automatically filling in a long form / page, which is used to detect UI elements that are not visible on the screen. The interface engine can analyze the document object model (DOM) of the long form / page to extract the DOM text of the relevant UI element, perform computer vision (CV) analysis to determine the type of the relevant UI element, and extrapolate the target UI element of the invisible area of ​​the long form / page from the relevant UI element. The interface engine can also include linking the DOM text to the target UI element based on the type and relevant UI element determined by the CV analysis. Extrapolation by the interface engine includes searching for similar hypertext markup language (HTML) structures (e.g., relevant DOM text corresponding to the target UI element) from the DOM. In addition, the interface engine can find the domain (i.e., UI element) of the long form / page from both the visible and invisible areas, and can extract data from the source to paste into the domain (e.g., on a subsequent or destination screen). In addition, the interface engine also has AI / ML functions to detect UI elements and automatically fill in long forms / pages.

[0024] Thus, one or more advantages, technical effects, and / or benefits of the interface engine include the ability to identify UI elements that are not visible on the screen without scrolling or moving the page presented by the screen (which is currently unavailable or not currently performed by conventional clipboard technologies). In this regard, the interface engine improves UI automation for identifying input field elements and reduces the time spent overcoming the scrolling activity performed by the user to indicate the target UI element.

[0025] Figure 1is an architectural diagram illustrating a hyper-automation system 100 according to one or more embodiments. "Hyper-automation" as used herein refers to an automation system that combines components of process automation, integration tools, and technologies that enhance work automation capabilities. For example, in some embodiments, RPA can be used at the core of a hyper-automation system, and in certain embodiments, automation capabilities can be extended using artificial intelligence and / or machines (AI / ML), process mining, analytics, and / or other advanced tools. As the hyper-automation system learns processes, trains AI / ML models, and employs analytics, for example, more and more knowledge work can be automated, and computing systems in an organization (e.g., both computing systems used by individuals and computing systems that operate autonomously) can participate in the hyper-automation process. The hyper-automation system of some embodiments allows users and organizations to efficiently and effectively discover, understand, and extend automation.

[0026] The hyper-automated system 100 includes user computing systems, such as a desktop computer 102, a tablet computer 104, and a smartphone 106. However, any desired computing system may be used without departing from the scope of one or more embodiments herein, including but not limited to smart watches, laptop computers, servers, Internet of Things (IoT) devices, etc. Figure 1 , but any suitable number of computing systems may be used without departing from the scope of one or more embodiments herein. For example, in some embodiments, dozens, hundreds, thousands, or millions of user computing systems may be used. The user computing systems may be actively used by the user, or may run automatically without much or any user input.

[0027] Each computing system 102, 104, 106 has (multiple) corresponding automated processes 110, 112, 114 running thereon. Without departing from the scope of one or more embodiments herein, (multiple) automated processes 102, 104, 106 may include, but are not limited to, an RPA robot, a portion of an operating system, (multiple) downloadable applications of the corresponding computing system, any other suitable software and / or hardware, or any combination of these. In some embodiments, one or more of the (multiple) processes 110, 112, 114 may be a listener. Without departing from the scope of one or more embodiments herein, the listener may be an RPA robot, a portion of an operating system, a downloadable application of the corresponding computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the (multiple) listeners is partially or completely implemented via physical hardware.

[0028] The listener monitors and records data related to the user's interaction with the corresponding computing system and / or the operation of the unattended computing system, and sends the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data may include, but is not limited to, which buttons are clicked, where the mouse is moved, the text entered in the field, one window is minimized while another window is open, the application associated with the window, etc. In some embodiments, the data from the listener may be sent periodically as part of a heartbeat message. In some embodiments, the data may be sent to the core hyper-automation system 120 once a predetermined amount of data has been collected, after a predetermined time period has passed, or both. For example, one or more servers (e.g., server 130) receive the data from the listener and store it in a database (e.g., database 140).

[0029] The automated process can execute the logic developed in the workflow at design time. In the case of RPA, a workflow can include a set of steps (defined as "activities" in this article) that are executed in a sequence or some other logical flow. Each activity can include an action, such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows can be nested or embedded.

[0030] In some embodiments, long-running workflows for RPA are the main project that supports service orchestration, human intervention, and long-running transactions in an unattended environment. See U.S. Patent No. 10,860,905, which is incorporated herein by reference in its entirety. Human intervention comes into play when certain processes require human input to handle exceptions, approvals, or verifications before moving to the next step of the activity. In this case, process execution is paused, freeing up the RPA robot until the human task is completed.

[0031] Long-running workflows can support workflow fragmentation via persistent activities, and can be combined with calling processes and non-user interactive activities to orchestrate manual tasks with RPA robot tasks. In some embodiments, multiple or many computing systems may participate in the logic of executing long-running workflows. Long-running workflows can run in sessions for fast execution. In some embodiments, long-running workflows can orchestrate background processes that may include activities that execute application programming interface (API) calls and run in long-running workflow sessions. In some embodiments, these activities can be called by calling process activities. A process with user interactive activities running in a user session can be called by starting a job from a commander activity (the commander will be described in more detail later in this article). In some embodiments, a user can interact with tasks that require completing a form in the commander. It may include activities that cause the RPA robot to wait for the form task to complete and then resume the long-running workflow.

[0032] One or more of the (multiple) automated processes 110, 112, 114 communicate with the core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may run a commander application on one or more servers (e.g., server 130). Although one server 130 is shown for illustrative purposes, multiple or many servers in close proximity or distributed architectures may be employed without departing from the scope of one or more embodiments herein. For example, without departing from the scope of one or more embodiments herein, one or more servers may be provided for commander functions, AI / ML model services, authentication, governance, and / or any other suitable functions. In some embodiments, the core hyper-automation system 120 may be incorporated into or as part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In some embodiments, the core hyper-automation system 120 may host multiple software-based servers (e.g., server 130) on one or more computing systems. In some embodiments, one or more servers (e.g., server 130) of the core hyper-automation system 120 may be implemented via one or more virtual machines (VMs).

[0033] In some embodiments, one or more of the (multiple) automated processes 110, 112, 114 may call one or more AI / ML models 132 deployed on or accessible by the core hyper-automation system 120. Without departing from the scope of one or more embodiments discussed in detail herein, the AI / ML model 132 may be trained for any suitable purpose. In some embodiments, two or more AI / ML models 132 may be linked (e.g., in series, in parallel, or a combination thereof) so that they jointly provide (multiple) collaborative outputs. The AI / ML model 132 may perform or assist computer vision (CV), optical character recognition (OCR), document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automatic RPA workflow generation, sequence extraction, cluster detection, audio to text conversion, any combination thereof, etc. However, without departing from the scope of one or more embodiments herein, any desired number and / or (multiple) types of AI / ML models may be used. For example, using multiple AI / ML models may allow the system to develop a global picture of what is happening on a given computing system. For example, one AI / ML model may perform OCR, another AI / ML model may detect buttons, another AI / ML model may compare sequences, and so on. The pattern may be determined by one AI / ML model alone or by multiple AI / ML models together. In some embodiments, one or more AI / ML models are locally deployed on at least one of the computing systems 102, 104, 106.

[0034] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that operates on data, and for example, the AI / ML model itself may be a deep learning neural network (DLNN) of trained artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions, such as statistical modeling (e.g., hidden Markov models (HMMs)), and utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform the desired function.

[0035] In some embodiments, the hyperautomation system 100 can provide four main groups of functions: (1) discovery; (2) building automation; (3) management; and (4) engagement. In some embodiments, the automation (e.g., running on a user computing system, server, etc.) can be run by a software robot (e.g., an RPA robot). For example, attended robots, unattended robots, and / or testing robots can be used. Attended robots work with users to assist them in completing tasks (e.g., via UiPath Assistant).TM ). Unattended robots work independently of users and can run in the background without the user's knowledge. Test robots are unattended robots that are used to run test cases against applications or RPA workflows. In some embodiments, test robots can run in parallel on multiple computing systems.

[0036] The discovery functionality may discover and provide automatic recommendations for different opportunities for business process automation. Such functionality may be implemented by one or more servers (e.g., server 130). In some embodiments, the discovery functionality may include providing an automation hub, process mining, task mining, and / or task capture. The automation hub (e.g., UiPath Automation Hub) TM ) can provide a mechanism for managing automation rollouts with visibility and control. For example, automation ideas can be crowdsourced from employees via a submission form. Feasibility and return on investment (ROI) calculations for automating these ideas can be provided, documentation for future automation can be collected, and collaboration can be provided to get faster builds from automation discoveries.

[0037] Process mining (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TM ) refers to the process of collecting and analyzing data from applications (e.g., enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify which end-to-end processes exist in the organization and how to effectively automate these processes and indicate the possible impact of automation. This data can be collected by a listener from user computing systems 102, 104, 106 and processed by a server (e.g., server 130). In some embodiments, one or more AI / ML models 132 can be used for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of process mining goals include, but are not limited to, increasing profits, improving customer satisfaction, complying with regulations and / or contracts, improving employee efficiency, and the like.

[0038] Task mining (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TM) identifies and aggregates workflows (e.g., employee workflows) and then applies AI to reveal patterns and variations in daily tasks, scoring these tasks for automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 can be employed to reveal repetitive task patterns in the data. Repetitive tasks that are ripe for automation can then be identified. In some embodiments, this information can initially be provided by a listener and analyzed on a server (e.g., server 130) of the core hyperautomation system 120. The results of task mining (e.g., Extensive Application Markup Language (XAML) process data) can be exported to a process document or designer application (e.g., UiPath Studio TM ) to create and deploy automation faster.

[0039] In some embodiments, task mining can include taking screenshots of user actions (e.g., mouse click locations, keyboard input, application windows and graphical elements with which the user is interacting, timestamps of interactions, etc.), collecting statistics (e.g., execution time, number of actions, text entries, etc.), editing and annotating screenshots, specifying the types of actions to be recorded, and so on.

[0040] Task capture (e.g., via UiPath Automation Cloud TM and / or UiPath AI Center TM ) automatically records the process of participation as the user works, or provides a framework for unattended processes. Such documents can include the desired tasks to be automated in the form of a process definition document (PDD), a framework workflow, capturing actions for each part of the process, recording user actions, and automatically generating a comprehensive workflow diagram including details of each step, Microsoft Documents, XAML files, and other documents. In some embodiments, build-ready workflows can be exported directly to a designer application, such as UiPath Studio TM Task capture can streamline the requirements gathering process for subject matter experts who interpret the process and Center of Excellence (CoE) members who provide production-grade automation.

[0041] Building automation can be applied via a designer (e.g., UiPath Studio TM 、UiPath StudioX TM or UiPath Web TM ) implementation. For example, RPA developers at PA development facility 150 can use RPA designer application 154 of computing system 152 to build and test automation for various applications and environments, such as web, mobile, and virtual desktops. API integration can be provided for a variety of applications, technologies, and platforms. Predefined activities, drag-and-drop modeling, and a workflow recorder make automation easier with minimal coding. Document understanding capabilities can be provided through drag-and-drop AI skills for data extraction and interpretation that invoke one or more AI / ML models 132. This automation can handle nearly any document type and format, including forms, checkboxes, signatures, and handwriting. When data is validated or exceptions are handled, this information can be used to retrain the corresponding AI / ML model to improve its accuracy over time.

[0042] For example, integration services can allow developers to seamlessly combine user interface (UI) automation with API automation. Automations can be built that require APIs or traverse both API and non-API applications and systems. Repositories for pre-built RPA and AI templates and solutions (e.g., UiPath Object Repository) can be provided. TM ) or a marketplace (e.g., UiPathMarketplace TM ) to allow developers to automate various processes faster. Therefore, when building automation, the hyper-automation system 100 can provide a user interface, development environment, API integration, pre-built and / or customized AI / ML models, development templates, integrated development environment (IDE), and advanced AI capabilities. In some embodiments, the hyper-automation system 100 enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.

[0043] In some embodiments, components of the hyperautomation system 100 (e.g., designer application(s) and / or external rule engines) provide support for managing and enforcing governance policies that control the various functions provided by the hyperautomation system 100. Governance is the ability of an organization to establish policies to prevent users from developing automations (e.g., RPA robots) that are capable of taking actions that could harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), third-party application terms of service, etc. Because developers may create automations that violate privacy laws, terms of service, etc. when executing their automations, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. In some embodiments, this can provide an additional level of security and compliance to the automation process development pipeline by preventing developers from relying on unapproved software libraries that may introduce security risks or work in a manner that violates policies, regulations, privacy laws, and / or privacy policies. See U.S. Non-Provisional Patent Application No. 16 / 924,499, the entire contents of which are incorporated herein by reference.

[0044] The management functions may provide management, deployment, and optimization of automation throughout the organization. In some embodiments, the management functions may include orchestration, test management, AI capabilities, and / or insights. The management functions of the hyper-automation system 100 may also serve as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management capabilities of the hyper-automation system 100 may include, but are not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnection of RPA robots.

[0045] For example, UiPath Orchestrator TM (In some embodiments, this can be implemented as UiPath AutomationCloud TM as part of a , either locally, in a virtual machine, in a private or public cloud, on Linux TM In a VM or via UiPath Automation Suite TM Commander applications such as the UiPath TestSuite provide orchestration capabilities for deploying, monitoring, optimizing, scaling, and securing RPA robot deployments. Test suites (e.g., UiPath TestSuite TM ) can provide test management to monitor the quality of deployed automation. Test suites can facilitate test planning and execution, requirements fulfillment, and defect traceability. Test suites can include comprehensive test reporting.

[0046] Analytics software (e.g., UiPath Insights TM ) can track, measure, and manage the performance of deployed automation. Analytics software can align automated actions with specific key performance indicators (KPIs) and strategic outcomes for the organization. Analytics software can present results in a dashboard format for better understanding by human users.

[0047] For example, data services (e.g., UiPath Data Service TM ) can be stored in database 140 and brought to a single, scalable, secure place through a drag-and-drop storage interface. Some embodiments can provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities can be powered by an AI center (e.g., UiPath AI Center TM) is provided, which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may make these functions accessible even to non-data scientists. Deployed automation (e.g., RPA robots) can call AI / ML models, such as AI / ML model 132, from the AI ​​center. The performance of the AI / ML model can be monitored, trained, and improved using, for example, human verification data provided by the data review center 160. A human reviewer can provide labeled data to the core hyper-automation system 120 via a review application 152 on a computing system 154. For example, a human reviewer can verify that the predictions made by the AI / ML model 132 are accurate or otherwise provide corrections. This dynamic input can then be saved as training data for retraining the AI / ML model 132 and can be stored in a database (e.g., database 140). The AI ​​center can then schedule and execute training jobs to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining of the AI / ML model 132.

[0048] Engagement capabilities bring humans and automation together as a team to seamlessly collaborate on desired processes. You can build low-code applications (e.g., via UiPath Apps TM ) to connect browser tabs and traditional software, even in the absence of an API in some embodiments. For example, applications can be quickly created using a web browser through a rich library of drag-and-drop controls. Applications can be connected to a single automation or multiple automations.

[0049] Action Center (for example, UiPath Action Center TM ) provides a simple and efficient mechanism to hand over automated processes to humans and vice versa. Humans can provide approvals or escalations, make exceptions, etc. Automation can then perform the automated functions of a given workflow.

[0050] A local assistant can be provided as a launchpad for users to start automation (e.g., UiPath Assistant TM). This functionality may be provided in a tray provided, for example, by an operating system, and may allow users to interact with RPA robots and RPA robot-driven applications on their computing systems. The interface may list automations approved for a given user and allow the user to run them. These may include off-the-shelf automations from an automation marketplace, an internal automation shop in an automation center, and the like. When automations run, they may run as local instances in parallel with other processes on the computing system, so that the user can use the computing system while the automation performs its actions. In some embodiments, the assistant is integrated with a task capture function so that users can record their processes that are about to be automated from the assistant launchpad.

[0051] Chatbots (e.g., UiPath Chatbots TM ), social messaging apps, and / or voice commands can enable users to run automations. This can simplify access to the information, tools, and resources that users need to interact with customers or perform other activities. Conversations between people can be easily automated, as can other processes. Triggered RPA robots launched in this way can perform actions, such as checking order status, posting data in a CRM, etc., possibly using plain language commands.

[0052] In some embodiments, the hyperautomation system 100 can provide end-to-end measurement and governance of automation programs of any size. In accordance with the above, analytics can be used to understand the performance of automation (e.g., via UiPathInsights TM ). Data modeling and analysis using any combination of available business metrics and operational insights can be used for a variety of automation processes. Custom designed and pre-built dashboards allow visualization of data across desired metrics, discovery of new analytical insights, tracking of performance metrics, discovery of ROI for automation, performance monitoring on user computing systems, detection of errors and anomalies, and debugging of automations. Automation management consoles (e.g., UiPath Automation Ops) can be provided TM ) to manage automations throughout the automation lifecycle. Organizations can govern how automations are built, what users can do with them, and which automations users can access.

[0053] In some embodiments, the hyperautomation system 100 provides an iterative platform. Processes can be discovered, automation can be built, tested, and deployed, performance can be measured, use of automation can be easily provided to users, feedback can be captured, AI / ML models can be trained and retrained, and processes can repeat themselves. This helps achieve a more robust and effective automation suite.

[0054] Figure 22 is an architectural diagram showing an RPA system 200 according to one or more embodiments. Figure 1 The RPA system 200 includes a designer 210 that allows developers to design and implement workflows. The designer 210 can provide solutions for application integration and automation of third-party applications, management of information technology (IT) tasks, and business IT processes. The designer 210 can facilitate the development of automation projects, which are graphical representations of business processes. Simply put, the designer 210 facilitates the development and deployment of workflows and robots (as shown by arrow 211). In some embodiments, the designer 210 can be an application running on a user's desktop, an application running remotely in a VM, a network application, etc.

[0055] Automation projects enable automation of rule-based processes by giving developers control over the order of execution and the relationships between sets of custom steps (defined herein as "activities" above) developed in a workflow. A commercial example of an embodiment of the designer 210 is UiPath Studio TM Each activity may include an action, such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.

[0056] Some types of workflows may include, but are not limited to, sequences, flow charts, finite state machines (FSMs), and / or global exception handlers. Sequences may be particularly useful for linear processes, thereby enabling flow from one activity to another without confusing the workflow. Flow charts may be particularly useful for more complex business logic, enabling decision integration and activity connections in a more diverse manner through multiple branching logic operators. FSMs may be particularly useful for large workflows. FSMs may use a limited number of states in their execution, which are triggered by conditions (i.e., transitions) or activities. Global exception handlers may be particularly useful for determining workflow behavior and debugging processes when encountering execution errors.

[0057] Once a workflow is developed in the designer 210, the execution of the business process is orchestrated by the director 220, which orchestrates one or more robots 230 that execute the workflow developed in the designer 210. One commercial example of an embodiment of the director 220 is the UiPath Orchestrator TM The director 220 facilitates the creation, monitoring, and deployment of resources in the management environment. The director 220 may serve as an integration point with third-party solutions and applications. As described above, in some embodiments, the director 220 may be Figure 1 Part of the core hyper-automation system 120.

[0058] The commander 220 can manage a fleet of robots 230, connecting and executing (as shown by arrow 231) robots 230 from a centralized point. The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234, but for development and testing purposes), and non-production robots (similar to attended robots 232, but for development and testing purposes). Attended robots 232 are triggered by user events and operate with people on the same computing system. Attended robots 232 can be used with commander 220 for centralized process deployment and logging media. Attended robots 232 can help manual users complete various tasks and can be triggered by user events. In some embodiments, processes cannot be started from the commander 220 on this type of robot, and / or they cannot be run under a locked screen. In some embodiments, attended robots 232 can only be started from the robot tray or from a command prompt. In some embodiments, attended robots 232 should be run under human supervision.

[0059] Unattended robots 234 run unattended in a virtual environment and can automate many processes. Unattended robots 234 can be responsible for remote execution, monitoring, scheduling, and providing support for work queues. In some embodiments, debugging of all robot types can be run in the designer 210. Both attended and unattended robots can automate (as shown in dashed box 290) various systems and applications, including but not limited to mainframes, network applications, VMs, enterprise applications (e.g., applications produced by the computer industry) and computing system applications (e.g., desktop and laptop computer applications, mobile device applications, wearable computer applications, etc.).

[0060] The commander 220 may have various capabilities (as indicated by arrow 232), including, but not limited to, provisioning, deploying, configuring, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining connections (e.g., web applications) between the robots 230 and the commander 220. Deployment may include ensuring that branch versions are properly delivered to assigned robots 230 for execution. Configuration may include maintenance and delivery of robot environment and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include tracking robot identification data and maintaining user permissions. Logging may include storing logs to a database (e.g., a Structured Query Language (SQL) or NoSQL database) and / or other storage mechanisms (e.g., a SQL database that provides the ability to store and quickly query large data sets). ) and indexed. Director 220 can provide interconnectivity by acting as a centralized communication point for third-party solutions and / or applications.

[0061] Robot 230 is an execution agent that implements the workflow constructed in designer 210. One commercial example of some embodiments of robot(s) 230 is UiPath Robots TM In some embodiments, the robot 230 installs Microsoft Services managed by the Service Control Manager (SCM). Therefore, such a robot 230 can open interactive session, and has Permissions for services.

[0062] In some embodiments, robots 230 can be installed in user mode. For such robots 230, this means that they have the same rights as the user who installed the given robot 230. This feature can also be used for high density (HD) robots to ensure that the maximum potential of each machine is fully utilized. In some embodiments, any type of robot 230 can be configured in an HD environment.

[0063] In some embodiments, the robot 230 is split into several components, each dedicated to a specific automation task. In some embodiments, the robot components include but are not limited to SCM-managed robot services, user-mode robot services, executors, agents, and command lines. SCM-managed robot service management and monitoring The SCM starts a console application under the local system.

[0064] In some embodiments, user-mode robot services manage and monitor session and acts as a proxy between the commander 220 and the execution host. The user mode robot service can be trusted and manage the credentials for the robot 230. If the SCM managed robot service is not installed, then Applications can be started automatically.

[0065] The actuator can be The executor can be aware of the dots per inch (DPI) setting of each monitor. The agent can be a window that displays available jobs in the system tray. A Windows Presentation Foundation (WPF) application. An agent can be a client of a service. An agent can request to start or stop a job and change settings. A command line is a client of a service. A command line is a console application that can request to start a job and wait for its output.

[0066] As described above, splitting the components of the robot 230 helps developers, support users, and computing systems to more easily run, identify, and track what each component is executing. Special behaviors can be configured for each component in this way, such as setting different firewall rules for executors and services. In some embodiments, the executor can always know the DPI setting of each monitor. Therefore, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow is created. In some embodiments, projects from the designer 210 can also be independent of the browser zoom level. In some embodiments, DPI can be disabled for applications that are DPI unaware or deliberately marked as unaware.

[0067] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for executing (multiple) AI / ML models and receive outputs therefrom via the core hyper-automation system 240.

[0068] As described above, one or more robots 230 may be listeners. These listeners may provide information to the core hyper-automation system 240 about what users are doing when using their computing systems. This information may then be used by the core hyper-automation system for process mining, task mining, task capture, and the like.

[0069] An assistant / chatbot 250 may be provided on the user computing system to allow the user to launch an RPA local robot. For example, the assistant may be located in the system tray. The chatbot may have a user interface so that the user can see the text in the chatbot. Alternatively, the chatbot may not have a user interface but run in the background to listen to the user's voice using the computing system's microphone.

[0070] In some embodiments, data labeling can be performed by a user of the computing system on which the robot is executing, or on another computing system to which the robot provides information. For example, if the robot calls an AI / ML model that performs CV on an image of a VM user, but the AI / ML model does not correctly identify a button on the screen, the user can draw a rectangle around the incorrectly identified or unidentified component and potentially provide text with the correct identification. This information can be provided to the core hyperautomation system 240 and then later used to train a new version of the AI / ML model.

[0071] Figure 3 is an architectural diagram illustrating an RPA system 300 deployed according to one or more embodiments. In some embodiments, the RPA system 300 may be Figure 2 RPA system 200 and / or Figure 1 The deployed RPA system 300 may be a cloud-based system, a local system, a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for different computing process automation, etc.

[0072] It should be noted that the client side 301, the server side 302, or both may include any desired number of computing systems without departing from the scope of one or more embodiments herein. On the client side 301, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is running a process. Multiple business projects may be running simultaneously, such as Figure 3 Agent 314 (e.g., The Executor 312 service is the single point of contact for all executors 312 in this embodiment. All messages in this embodiment are logged to the Director 340, which further processes the messages via the database server 355, AI / ML server 360, indexer server 370, or any combination thereof. Figure 2 As discussed, the actuator 312 may be a robotic component.

[0073] In some embodiments, a robot represents an association between a machine name and a user name. A robot can manage multiple executors simultaneously. On a computing system that supports running multiple interactive sessions simultaneously (e.g., Server 2012), multiple robots can run simultaneously, each with a unique username in a separate This is the HD robot mentioned above.

[0074] The agent 314 is also responsible for sending the status of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still running) and downloading the desired version of the group to be executed. In some embodiments, communication between the agent 314 and the commander 340 is always initiated by the agent 314. In the notification scenario, the agent 314 can open a WebSocket channel that is later used by the commander 340 to send commands to the robot (e.g., start, stop, etc.).

[0075] The listener 330 monitors and records data related to the user's interaction with the attended computing system and / or the operation of the unattended computing system in which the listener 330 resides. Without departing from the scope of one or more embodiments herein, the listener 330 may be an RPA robot, part of an operating system, a downloadable application for a corresponding computing system, or any other software and / or hardware. In fact, in some embodiments, the logic of the listener is partially or completely implemented via physical hardware.

[0076] On the server side 302, a presentation layer 333, a service layer 334, a persistence layer 336, and a commander 340 are included. The presentation layer 333 may include a network application 342, an open data protocol (OData) representational state transfer (REST) ​​application programming interface (API) endpoint 344, and notification and monitoring 346. The service layer 334 may include an API implementation / business logic 348. The persistence layer 336 may include a database server 355, an AI / ML server 360, and an indexer server 370. For example, the commander 340 includes a network application 342, an OData REST API endpoint 344, notification and monitoring 346, and an API implementation / business logic 348. In some embodiments, most actions performed by a user in the interface of the commander 340 (e.g., via a browser 320) are performed by calling various APIs. Without departing from the scope of one or more embodiments herein, such actions may include, but are not limited to, starting a job on a robot, adding / deleting data in a queue, scheduling a job to run unattended, etc. The network application 342 may be the visual layer of the server platform. In this embodiment, web application 342 uses HTML and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of one or more embodiments herein. In this embodiment, a user interacts with a web page from web application 342 via browser 320 to perform various actions to control commander 340. For example, a user can create robot groups, assign groups to robots, analyze logs per robot and / or per process, start and stop robots, and the like.

[0077] In addition to the web application 342, the director 340 also includes a service layer 334 that exposes an OData REST API endpoint 344. However, other endpoints may be included without departing from the scope of one or more embodiments herein. The REST API is consumed by both the web application 342 and the agent 314. In this embodiment, the agent 314 is a manager of one or more robots on a client computer.

[0078] The REST API in this embodiment includes configuration, logging, monitoring, and queuing functionality (as shown by at least arrow 349). In some embodiments, the configuration endpoint can be used to define and configure application users, permissions, robots, assets, releases, and environments. The logging REST endpoint can be used to log different information, such as errors, explicit messages sent by the robot, and other environment-specific information. If the start job command is used in the commander 340, the robot can use the deployment REST endpoint to query the grouped version that should be executed. The queuing REST endpoint can be responsible for queue and queue item management, such as adding data to the queue, getting transactions from the queue, setting the status of transactions, and so on.

[0079] The monitoring REST endpoint can monitor the network application 342 and the agent 314. The notification and monitoring API 346 can be a REST endpoint for registering the agent 314, delivering configuration settings to the agent 314, and for sending / receiving notifications from the server and the agent 314. In some embodiments, the notification and monitoring API 346 can also use WebSocket communication. Figure 3 As shown, one or more activities / actions described herein are represented by arrows 350 and 351 .

[0080] In some embodiments, the APIs in the service layer 334 can be accessed by configuring an appropriate API access path, for example, based on whether the commander 340 and the entire hyper-automation system have a local deployment type or a cloud-based deployment type. The API for the commander 340 may provide customized methods for querying statistical information about various entities registered in the commander 340. In some embodiments, each logical resource may be an OData entity. In such an entity, components such as robots, processes, queues, etc. may have attributes, relationships, and operations. In some embodiments, the network application 342 and / or the agent 314 may consume the commander 340 API in two ways: by obtaining API access information from the commander 340, or by registering an external application to use the OAuth flow.

[0081] In this embodiment, the persistence layer 336 includes three servers - a database server 355 (e.g., a SQL server), an AI / ML server 360 (e.g., a server providing AI / ML model services, such as an AI center function), and an indexer server 370. The database server 355 in this embodiment stores configurations of robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed by a network application 342. The database server 355 can manage queues and queue items. In some embodiments, the database server 355 can store messages recorded by the robot log (in addition to or instead of the indexer server 370). The database server 355 can also store process mining, task mining, and / or task capture related data, such as received from the listener 330 installed on the client side 301. Although no arrow is shown between the listener 330 and the database 355, it should be understood that in some embodiments, the listener 330 can communicate with the database 355, and vice versa. The data can be stored in the form of a PDD, an image, a XAML file, etc. The listener 330 may be configured to intercept user actions, processes, tasks, and performance metrics on the corresponding computing system where the listener 330 is located. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, times, etc.) on its corresponding computing system and then convert these actions into a suitable format to be provided to and stored in the database server 355.

[0082] The AI / ML server 360 facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may make these capabilities accessible even to non-data scientists. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of the AI / ML models can be monitored and trained and improved using manually validated data. The AI / ML server 360 can schedule and execute training jobs to train new versions of the AI / ML models.

[0083] The AI / ML server 360 may store data related to AI / ML models and ML groups for configuring various ML skills for users at the time of development. ML skills as used herein are pre-built and trained ML models for processes, for example, which may be used by automation. The AI / ML server 460 may also store data related to document understanding techniques and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, natural language processing (NLP), speech analysis, different types of AI / ML models, and the like.

[0084] Indexer server 370 (which is optional in some embodiments) stores and indexes information logged by the robot logs. In some embodiments, indexer server 370 can be disabled through configuration settings. In some embodiments, indexer server 370 uses (which is an open source project full-text search engine.) Messages logged by the robot (e.g., using activities such as logging messages or writing lines) can be sent to the indexer server 370 through the logging REST endpoint(s), where they are indexed for future use.

[0085] Figure 4 4 is an architectural diagram showing the relationship 400 between the designer 410, activities 420, 430, 440, 450, driver 460, API 470, and AI / ML model 480 according to one or more embodiments. As described herein, developers use the designer 410 to develop workflows executed by the robot. In some embodiments, various types of activities can be displayed to the developer. The designer 410 can be local or remote to the user's computing system (e.g., accessed via a VM or a local web browser interacting with a remote web server). The workflow can include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. For example (as shown by the dotted lines), user-defined activities 420 and API-driven activities 440 interact with the application via its API. In turn, in some embodiments, user-defined activities 420 and / or AI / ML activities 440 can call one or more AI / ML models 480, which can be located locally and / or remotely from the computing system on which the robot operates.

[0086] Some embodiments are capable of identifying non-textual visual components in images, referred to herein as CV. CV may be performed at least in part by (multiple) AI / ML models 480. Some CV activities associated with such components may include, but are not limited to, extracting text from segmented label data using OCR, fuzzy text matching, cropping segmented label data using ML, comparing text extracted from label data with basic fact data, and the like. In some embodiments, hundreds or even thousands of activities may be implemented in user-defined activities 420. However, any number and / or type of activities may be used without departing from the scope of one or more embodiments herein.

[0087] UI automation activities 450 are a subset of special lower-level activities written in lower-level code and facilitating interaction with the screen. UI automation activities 450 facilitate these interactions via drivers 460, which allow the robot to interact with the desired software. For example, drivers 460 may include operating system (OS) drivers 462, browser drivers 464, VM ​​drivers 466, enterprise application drivers 468, etc. In some embodiments, one or more of the AI / ML models 480 may be used by UI automation activities 450 to perform interactions with the computing system. In some embodiments, AI / ML models 480 may enhance drivers 460 or replace them completely. In fact, in some embodiments, drivers 460 are not included.

[0088] Driver 460 can interact with the OS at a low level via OS driver 462 to find hooks, monitor keys, etc. Driver 460 can facilitate communication with For example, the "click" activity performs the same role in these different applications via the driver 460.

[0089] Figure 5 is an architectural diagram illustrating a computing system 500 configured to provide an interface engine for RPA according to one or more embodiments. In some embodiments, computing system 500 may be one or more of the computing systems depicted and / or described herein. In some embodiments, computing system 500 may be, for example, Figure 1 and Figure 2 Part of the hyper-automation system shown. Computing system 500 includes a bus 505 or other communication mechanism for transmitting information, and (multiple) processors 510 coupled to bus 505 for processing information. (Multiple) processors 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. (Multiple) processors 510 can also have multiple processing cores, and at least some of these cores can be configured to perform specific functions. In some embodiments, multi-parallel processing can be used. In some embodiments, at least one of (multiple) processors 510 can be a neuromorphic circuit including a processing element that simulates a biological neuron. In some embodiments, the neuromorphic circuit may not require typical components of the von Neumann computing architecture.

[0090] The computing system 500 also includes a memory 515 for storing information and instructions to be executed by the processor(s) 510. The memory 515 may include any combination of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage (e.g., a magnetic disk or optical disk), or any other type of non-transitory computer-readable medium or a combination thereof. Non-transitory computer-readable media may be any available media that can be accessed by the processor(s) 510, and may include volatile media, non-volatile media, or both. The media may also be removable, non-removable, or both.

[0091] In addition, the computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, without departing from the scope of one or more embodiments herein, the communication device 520 can be configured to use frequency division multiple access (FDMA), single carrier FDMA (SC-FDMA), time division multiple access (TDMA), code division multiple access (CDMA), orthogonal frequency division multiplexing (OFDM), orthogonal frequency division multiple access (OFDMA), global system for mobile (GSM) communication, general packet radio service (GPRS), universal mobile telecommunications system (UMTS), cdma2000, wideband CDMA (W-CDMA), high speed downlink packet access (HSDPA), high speed uplink packet access (HSUPA), high speed packet access (HSPA), long term evolution (LTE), advanced LTE A), 802.11x, Wi-Fi, Zigbee, Ultra Wideband (UWB), 802.16x, 802.15, Home Node B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near Field Communication (NFC), Fifth Generation (5G) New Radio (NR), any combination thereof, and / or any other currently existing or future implemented communication standards and / or protocols. In some embodiments, without departing from the scope of one or more embodiments herein, the communication device 520 may include one or more antennas that are singular, array, panel, phased, switched, beamformed, beamsteered, combinations thereof, and / or any other antenna configuration.

[0092] The processor(s) 510 are also coupled via bus 505 to a display 525 (i.e., screen), such as a plasma display, a liquid crystal display (LCD), a light emitting diode (LED) display, a field emission display (FED), an organic light emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high definition display, Display, in-plane switching (IPS) display, or any other suitable display for displaying information to a user. Display 525 can be configured as a touch (tactile) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-point touch display, etc., which use resistance, capacitance, surface acoustic wave (SAW) capacitance, infrared, optical imaging, dispersion signal technology, acoustic pulse recognition, frustrated total internal reflection, etc. Any suitable display device and tactile I / O can be used without departing from the scope of one or more embodiments herein.

[0093] A keyboard 530 and a cursor control device 535 (e.g., a computer mouse, touchpad, etc.) are further coupled to bus 505 to enable a user to interface with computing system 500. However, in some embodiments, a physical keyboard and mouse may not be present, and the user may interact with the device only through display 525 and / or touchpad (not shown). Any type and combination of input devices may be used as a matter of design choice. In some embodiments, there are no physical input devices and / or displays. For example, a user may interact with computing system 500 remotely via another computing system in communication with it, or computing system 500 may operate autonomously.

[0094] Memory 515 stores software modules that provide functionality when executed by processor(s) 510. These modules include an operating system 540 for computing system 500. The modules also include a module 545 (e.g., a module implementing an interface engine) that is configured to perform all or part of the processes described herein or derivatives thereof.

[0095] According to one or more embodiments, module 545 may perform one or more operations such as DOM analysis, DOM text extraction, CV analysis, and pattern extraction.

[0096] According to one or more embodiments, module 545 may also perform one or more operations, such as detecting a frame, starting CV analysis, detecting top levels, extracting data, determining nodes (e.g., UiNodes), determining locations, detecting element types (defined by DOM, frame, CV, and / or element), populating automatic anchors, populating date and time, cleaning elements, enriching elements, creating data structures (e.g., DOM trees, etc.), collecting relationships (e.g., DOM relationships), collecting drop-down information, collecting geometric anchors, collecting options, collecting values, collecting date and time groups, and determining patterns.

[0097] According to one or more embodiments, the DOM may be a programming API for documents, web pages, and other files, and a DOM tree may be considered as a data structure, schema, or structural model of a DOM (e.g., of a web page). For example, attributes may be considered as nodes in a DOM (e.g., an API of a document), but not nodes in a DOM tree (e.g., a structure of a document), while elements of a DOM may be provided as nodes of a DOM tree including attributes, tags, and children.

[0098] According to one or more embodiments, module 545 also has AI / ML and RPA capabilities to perform one or more operations herein.Computing system 500 may include one or more additional function modules 550 including additional functions.

[0099] Those skilled in the art will appreciate that, without departing from the scope of one or more embodiments herein, a "system" may be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices. Presenting the above functions as being performed by a "system" is not intended to limit the scope of one or more embodiments herein in any way, but is intended to provide an example of many embodiments. In fact, the methods, systems, and devices disclosed herein can be implemented in a localized and distributed form consistent with computing technology, including cloud computing systems. The computing system can be part of or otherwise accessible by a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public or private cloud, a hybrid cloud, a server farm, any combination thereof, and the like. Any localized or distributed architecture can be used without departing from the scope of one or more embodiments herein.

[0100] It should be noted that some of the system features described in this specification have been presented as modules in order to more particularly emphasize their implementation independence. For example, a module can be implemented as a hardware circuit that includes a custom very large scale integrated circuit (VLSI) circuit or gate array, an off-the-shelf semiconductor such as a logic chip, a transistor, or other discrete components. A module can also be implemented in a programmable hardware device such as a field programmable gate array, programmable array logic, a programmable logic device, a graphics processing unit, or other device.

[0101] Modules may also be implemented at least in part in software for various types of processors to perform. The identified executable code unit may, for example, include one or more physical or logical blocks of computer instructions, which may, for example, be organized as objects, processes, or functions. However, the executable program of the identified module does not need to be physically located together, but may include different instructions stored in different locations, which include the module and realize the above-mentioned purpose of the module when logically linked together. In addition, without departing from the scope of one or more embodiments herein, the module may be stored on a computer-readable medium, which may be, for example, a hard disk driver, a flash memory device, a RAM, a magnetic tape, and / or any other such non-transient computer-readable medium for storing data.

[0102] In fact, a module of executable code can be a single instruction or multiple instructions, and can even be distributed over several different code fragments, between different programs, and between several memory devices. Similarly, operational data can be identified and shown within a module, and can be embodied in any suitable form and organized in any suitable type of data structure. The operational data can be collected as a single data set, or can be distributed over different locations, including on different storage devices, and can exist at least in part only as electronic signals on a system or network.

[0103] Various types of AI / ML models may be trained and deployed without departing from the scope of one or more embodiments herein. For example, Figure 6 An example of a neural network 600 that has been trained to recognize graphical elements in an image is shown in accordance with one or more embodiments. Here, the neural network 600 receives pixels of a screenshot image of a 1920×1080 screen (shown as column 610) as input to the input “neurons” 1 through I (shown as column 620) of the input layer. In this case, I is 2,073,600, which is the total number of pixels in the screenshot image.

[0104] Neural network 600 also includes multiple hidden layers (as shown in columns 630 and 640). Both DLNNs and shallow learning neural networks (SLNNs) typically have multiple layers, although in some cases, SLNNs may have only one or two layers, and typically fewer than DLNNs. Typically, a neural network architecture includes an input layer, multiple intermediate layers (e.g., hidden layers), and an output layer (as shown in column 650), as in the case of neural network 600.

[0105] DLNNs typically have many layers (e.g., 10, 50, 200, etc.), and subsequent layers typically reuse features from previous layers to compute more complex general functions. On the other hand, SLNNs tend to have only a few layers and are relatively fast to train because expert features are created in advance from raw data samples. However, feature extraction is laborious. On the other hand, DLNNs typically do not require expert features, but tend to take longer to train and have more layers.

[0106] For both approaches, the layers are trained simultaneously on a training set, usually with a check for overfitting on an isolated cross-validation set. Both techniques produce excellent results, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depends on the problem being solved by the neural network.

[0107] return Figure 6 , the pixels provided as the input layer are fed as inputs to the J neurons of hidden layer 1. Although all pixels are fed to each neuron in this example, various architectures may be used alone or in combination without departing from the scope of one or more embodiments herein, including but not limited to feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graph networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long / short term memory networks, gated recurrent unit networks, generative adversarial networks, liquid machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional layer networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks.

[0108] Hidden layer 2 (630) receives input from hidden layer 1 (620), hidden layer 3 receives input from hidden layer 2, and so on, until the last hidden layer (as shown by ellipse 655) provides its output as input to the output layer. It should be noted that the number of neurons I, J, K, and L is not necessarily equal, and thus, any desired number of layers may be used for a given layer of neural network 600 without departing from the scope of one or more embodiments herein. In fact, in some embodiments, the types of neurons in a given layer may not all be the same.

[0109] The neural network 600 is trained to assign confidence scores to graphical elements that are believed to have been found in the image. In order to reduce matches with unacceptably low likelihoods, in some embodiments, only those results with confidence scores that meet or exceed a confidence threshold may be provided. For example, if the confidence threshold is 80%, then the outputs with confidence scores exceeding that amount may be used, and the remaining outputs may be ignored. In this case, the output layer indicates that two text fields (as shown in outputs 661 and 662), a text label (as shown in output 663), and a submit button (as shown in output 665) were found. The neural network 600 may provide the location, size, image, and / or confidence scores of these elements without departing from the scope of one or more embodiments herein, which may then be used by an RPA robot or another process that uses the output for a given purpose.

[0110] It should be noted that neural networks are probabilistic structures that typically have confidence scores. This can be a score learned by the AI / ML model based on the frequency of correctly identifying similar inputs during training. For example, text fields typically have a rectangular shape and a white background. Neural networks can learn to identify graphical elements with these features with high confidence. Some common types of confidence scores include decimal numbers between 0 and 1 (which can also be interpreted as confidence percentages), numbers between negative ∞ and positive ∞, and a set of expressions (e.g., "low", "medium", and "high"). In order to obtain more accurate confidence scores, various post-processing calibration techniques can also be used, such as temperature scaling, batch normalization, weight decay, negative log-likelihood (NLL), etc.

[0111] A "neuron" in a neural network is a mathematical function that is generally based on the function of a biological neuron. Neurons receive weighted inputs and have a summation and activation function that governs whether they pass the output to the next layer. The activation function can be a nonlinear threshold activity function, where nothing happens if the value is below the threshold, but then the function responds linearly above the threshold (i.e., rectified linear unit (ReLU) nonlinearity). The summation function and the ReLU function are used for deep learning because real neurons can have approximately similar activity functions. Via linear transformations, information can be subtracted, added, etc. In essence, the neuron acts as a gating function that passes the output to the next layer, which is governed by its underlying mathematical function. In some embodiments, different functions may be used for at least some neurons.

[0112] exist Figure 7 An example of a neuron 700 is shown in FIG. 7 . The input x1, x2, ..., x from the previous layer n are assigned corresponding weights w1, w2, ..., w n. Therefore, the collective input from the previous neuron 1 is w1x1. These weighted inputs are used in the summation function of the neuron modified by the bias, for example as shown in Equation 1:

[0113]

[0114] This sum is compared to the activation function f(x) (as shown in block 710) to determine whether the neuron "fires". For example, f(x) may be given as shown in Equation 2:

[0115]

[0116] The output y of neuron 700 can therefore be given as shown in Equation 3:

[0117]

[0118] In this case, neuron 700 is a single layer perceptron. However, any suitable neuron type or combination of neuron types may be used without departing from the scope of one or more embodiments herein. It should also be noted that in some embodiments, the range of values ​​of the weights and / or output values ​​of the activation function may be different without departing from the scope of one or more embodiments herein.

[0119] Typically a goal or "reward function" is employed, e.g., in this case, successfully identifying a graphical element in an image. The reward function explores intermediate transitions and steps with short-term and long-term rewards to guide the search of the state space and attempt to achieve the goal (e.g., successfully identifying a graphical element, successfully identifying the next sequence of activities in an RPA workflow, etc.).

[0120] During training, various labeled data (in this case, images) are fed through the neural network 600. Successful identifications strengthen the weight of the neuron's input, while unsuccessful identifications weaken that weight. A cost function (e.g., mean squared error (MSE) or gradient descent) can be used to penalize slightly wrong predictions much less than very wrong predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, the data scientist can modify the reward function, provide an indication of where unidentified graphical elements are, provide corrections for incorrectly identified graphical elements, and so on.

[0121] Backpropagation is a technique used to optimize synaptic weights in feed-forward neural networks. Backpropagation can be used to "uncover" the hidden layers of a neural network to see how much loss each node is responsible for, and then update the weights so that the loss is minimized by giving lower weights to nodes with higher error rates, and vice versa. In other words, backpropagation allows data scientists to repeatedly adjust the weights to minimize the difference between the actual output and the desired output.

[0122] The back-propagation algorithm is built on the mathematical foundations of optimization theory. In supervised learning, training data with known outputs is passed through a neural network and the error is calculated using a cost function with known target outputs, which provides the error for back-propagation. The error is calculated at the output and transformed into a correction to the network weights to minimize the error.

[0123] In the case of supervised learning, an example of backpropagation is provided below. A series of N nonlinear activation functions f between each layer i=1, ..., N of the network i To process the column vector input x, the output of a given layer is first multiplied by the synaptic matrix W i , and add the bias vector b i . The network output o is given by Equation 4.

[0124] O=f N (W N f N-1 (W N-1 f N-2 (...f1(W1x+b1)...)+b N-1 )+b N )

[0125] Equation 4

[0126] In some embodiments, o is compared to the target output t to obtain the error that is desired to be minimized.

[0127] Optimization in the form of a gradient descent procedure can be used to modify the synaptic weights W of each layer by i to minimize the error. The gradient descent process requires computing an output o given an input x corresponding to a known target output t, and produces an error ot. This global error is then propagated backward to give a local error for the weight update, whose computation is similar but not identical to that used for the forward propagation. In particular, the backpropagation step typically requires a formula of the form p j (n j ) = f′ j (n j ) activity function, where n j is the network activity at layer j (i.e., nj =W j o j-1 +b j ), where o j =f j (n j ), and the prime symbol ' denotes the derivative of the activity function f.

[0128] The weight update can be calculated by Equation 5, Equation 6, Equation 7, Equation 8, and Equation 9:

[0129]

[0130]

[0131]

[0132]

[0133]

[0134] in represents the Hadamard product (i.e., the element-wise product of two vectors), T represents the matrix transpose, o j represents f j (W j o j-1 +b j ), where o0 = x. Here, the learning rate η is chosen based on machine learning considerations. Below, η is related to the neural Hebbian learning mechanism used in the neural implementation. Note that the synapses W and b can be combined into a large synaptic matrix, where the input vector is assumed to have one appended, and the extra columns representing the b synapses are included in W.

[0135] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., 97% or better using an F2 or F4 threshold for detection and approximately 2000 epochs). Without departing from the scope of one or more embodiments herein, in some embodiments, the level of accuracy can be determined using an F1 score, an F2 score, an F4 score, or any other suitable technique. Once trained on the training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps ensure that the AI / ML model does not "overfit," whereby it identifies graphical elements in the training data well but does not generalize well to other images.

[0136] In some embodiments, it may not be known what level of accuracy the AI / ML model can achieve. Therefore, if the accuracy of the AI / ML model begins to decline when analyzing the evaluation data (i.e., the model performs well on the training data but performs poorly on the evaluation data), the AI / ML model can go through more training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when the accuracy reaches a certain level or the accuracy of the trained AI / ML model is better than the existing deployed AI / ML model.

[0137] In certain embodiments, a collection of trained AI / ML models may be used to complete a task, e.g., employing an AI / ML model for each type of graphical element of interest, employing an AI / ML model to perform OCR, deploying yet another AI / ML model to identify proximity relationships between graphical elements, employing yet another AI / ML model to generate RPA workflows based on the output of other AI / ML models, etc. This may collectively allow AI / ML models to achieve semantic automation.

[0138] Some embodiments may use transformer networks, such as SentenceTransformers TM , a Python library for state-of-the-art sentence, text, and image embedding TM Framework. Such a transformer network learns associations of words and phrases with both high and low scores. This trains the AI / ML model to determine which are close to the input and which are not, respectively. The transformer network can also use domain length and domain type, rather than just pairs of words / phrases.

[0139] Figure 8 800 for training (multiple) AI / ML models according to one or more embodiments. Note that process 800 can also be applied to other UI learning operations, such as NLP and chatbots. The process begins with training data, such as providing Figure 8 Labeled data is shown, such as labeled screens (e.g., with identified graphical elements and text), words and phrases, a "thesaurus" of semantic associations between words and phrases so that similar words and phrases to a given word or phrase can be identified, etc. The nature of the training data provided can depend on the goals that the AI / ML model is intended to achieve. The AI / ML model is then trained over multiple epochs at box 820, and the results are reviewed at box 830.

[0140] If the AI / ML model fails to meet the desired confidence threshold at decision box 840 (process 800 proceeds according to the "no" arrow), the training data is supplemented and / or the reward function is modified at box 850 to help the AI / ML model better achieve its goal, and the process returns to box 820. If the AI / ML model meets the confidence threshold at decision box 840 (process 800 proceeds according to the "yes" arrow), the AI / ML model is tested on evaluation data at box 860 to ensure that the AI / ML model generalizes well and the AI / ML model does not overfit with respect to the training data. The evaluation data may include screens, source data, etc. that the AI / ML model has not processed before. If the confidence threshold is met at decision box 870 for the evaluation data (process 800 proceeds according to the "yes" arrow), the AI / ML model is deployed at box 880. If not (process 800 proceeds according to the "no" arrow), the process returns to box 880 and the AI / ML model is further trained.

[0141] Fig. 9 900 is a flow chart illustrating a process 900 for providing an interface engine according to one or more embodiments. In general, the process 900 provides for detecting a display (e.g., Figure 5 Example operation of an interface engine that renders a UI element that is not visible on a display 525 or a screen presenting a UI including the UI element.

[0142] According to one or more embodiments, Figure 8 The process 800 and Fig. 9 The process 900 performed in the embodiment of the present invention may be performed by a computer program that encodes instructions for a processor(s). The computer program may be embodied on a non-transitory computer readable medium. The computer readable medium may be, but is not limited to, a hard drive, a flash memory device, RAM, a magnetic tape, and / or any other such medium or combination of media for storing data. The computer program may include instructions for controlling the processor(s) of the computing system (e.g., Figure 5 The processor(s) 510 of the computing system 500 are used to implement Figure 8-Figure 9 The coded instructions for all or part of the processing steps described in the present invention may also be stored on a computer-readable medium.

[0143] Process 900 begins at block 910, where a screen presents a UI. A screen is a display as described herein. A UI may include one or more web browser windows or interface frames, as well as icons, toolbars, and the like.

[0144] The UI may include a page. The page may be a web page within one of a web browser window or an interface frame. The page may have one or more layers forming a stacked structure. The web browser window or interface frame may extend to the edge of the screen or occupy a portion of the screen. The web browser window or interface frame may include a scroll bar that enables browsing of the page. The page may be larger than the border of the web browser or interface frame. The page may be larger than the screen. Therefore, the page may be partially displayed by the screen so that the visible area of ​​the page is limited by the border of the web browser window or interface frame (e.g., whether the web browser window or interface frame is fully extended to the edge of the screen), and the invisible area of ​​the page can be seen when operating the scroll bar or screen. For ease of explanation, with respect to process 900, the border of the web browser window or interface frame and the edge of the screen are simultaneous.

[0145] A long form may be included in a page. A page and / or a long form may include UI elements. Pages, long forms, and UI elements are identified, itemized, and described by DOM. DOM is a cross-platform and language-independent mechanism that provides UI elements as a tree structure in which each node represents a portion of a page and / or a long form. For example, DOM may include an HTML structure for each UI element. UI elements may include input fields, such as such drop-down menus, check boxes, radio buttons, and other interactive elements. Each UI element may be identified by a type, such as a DOM type, a framework type, a CV type, or an element definition type. Therefore, the interface engine may find input fields and corresponding types in DOM.

[0146] At block 930, the interface engine performs DOM text extraction. DOM text extraction includes the interface engine analyzing the DOM of the page to determine which UI elements may be relevant to auto-completion or filing based on the DOM text therein.

[0147] According to one or more embodiments, the interface engine analyzes the DOM text of the DOM of the page to find all nodes for UI elements. The interface engine extracts each node and the relevant DOM text corresponding to the UI element in the DOM for possible related elements. For example, the related elements can be the input field of a long form and the corresponding label of the input field. Note that the problem with the page DOM is that "what is an input field" and "what is a container" may not be clear. The interface engine solves this problem by automatically filtering out all unnecessary context while analyzing / extracting the DOM and capturing only relevant text.

[0148] According to one or more embodiments, the interface engine generates a filtered DOM. The filtered DOM includes filtered elements, such as one or more related UI elements determined from possible related elements. For example, when the first related element is determined, the interface engine generates (i.e., creates) a filtered DOM (e.g., an element list). As each subsequent related element is determined, the interface engine builds (i.e., adds to) the element list. Therefore, the element list is an itemization of one or more related UI elements with corresponding text from the DOM. The corresponding text may include, but is not limited to, depth levels, geometric positioning, types, and text values ​​of the filtered elements.

[0149] According to one or more embodiments, the interface engine can use ML models to analyze the page, such as anchor / relationship models. For example, deep hierarchy, geometric positioning, type, and text value can be aspects of an anchor / relationship model that determines relevant DOM text (e.g., HTML structure) related to various input fields. The anchor / relationship model can determine relevant DOM text as further described herein.

[0150] At box 950, the interface engine performs a CV analysis. The CV analysis is performed on the visible area. According to one or more embodiments, the interface engine performs a CV analysis to determine the type of one or more related UI elements within the visible area. For example, to determine the type of an input field, the interface engine captures the visible area of ​​the page as a screenshot. The interface engine places the screenshot in the CV analysis to identify the types of all input fields on the visible area.

[0151] Go to Fig.10 , depicting an interface 1000 according to one or more embodiments. Interface 1000 is an example of a page shown within the boundaries of an interface frame that is concurrent with the edges of a screen. Thus, a visible area 1001 of a page of interface 1000 is within the boundaries of the interface frame and the edges of the screen, while invisible areas are not shown. Interface 1000 includes a plurality of elements, such as elements 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, and 1019. Note that element 1018 is a scroll bar that can be manipulated to move invisible areas of the page into portions of visible area 1001.

[0152] According to one or more embodiments, the interface engine identifies some of the multiple elements as unneeded context, such as element 1019 (i.e., "Assigned to Me" text) having been previously filtered. According to one or more embodiments, the interface engine identifies some of the multiple elements as related UI elements or input fields, such as element 1012 (i.e., "Question Type" field) requiring type identification. According to one or more embodiments, the interface engine can use the element list with the filtered DOM to identify the relevant UI elements of the visible area 1001.

[0153] The interface engine performs CV analysis on the visible area 1001 to identify the types of these related UI elements. For example, the interface 1000 shows that the element 1012 can be determined by the interface engine as a drop-down type according to the CV analysis. In addition, the interface engine determines the structure of the element 1012 from the filtered DOM.

[0154] return Fig. 9 , as shown in sub-block 955, the CV analysis may be stored in a cache by the interface engine. The cache may be a portion of a memory, such as in Figure 5 The interface engine can use the information in the cache when a similar long form is indicated as a destination form (e.g., on a subsequent page (e.g., a second page that may be in a different interface frame than the current page). Technical effects, advantages, and benefits of the interface engine include the interface engine using the cache to infer and identify input fields for invisible regions and other pages / destination forms without performing additional CV analysis.

[0155] At box 970, the interface engine performs extrapolation. The extrapolation of the interface engine enables the interface engine to determine (e.g., identify and understand) what type of off-screen elements exist. The interface engine extrapolates one or more target UI elements of the invisible area. The one or more target UI elements are a subset of one or more UI elements of a long form or page. The invisible area can be part of a long form or a page outside a web browser window or interface frame. The interface engine's extrapolation of one or more target UI elements includes searching for similar HTML structures from a DOM or a filtered DOM.

[0156] According to one or more embodiments, the ML model of the interface engine determines the HTML structure for the input domain of the identification of the visible area, and uses the HTML structure to find a similar HTML structure in the invisible area. In turn, using the ML model, the interface engine extrapolates the relevant UI elements of the visible area to the target UI elements in the invisible area. For example, the interface engine uses the relevant DOM extraction text and CV analysis to determine the structure of the input domain from the visible area. The structure of the input domain may include classes and attributes of the HTML part related to the domain identified from the relevant DOM extraction text and CV analysis. Then, similar structures are searched in the relevant DOM extraction text to identify other similar input domains in the invisible area of ​​the screen. For example, if the CV analysis determines that the element on the visible area of ​​the screen is a drop-down type of the input domain, the corresponding HTML structure of the input domain is determined based on the filtered DOM text. The HTML structure related to the drop-down type is searched from the DOM text related to the invisible area, and all drop-down type input domains in the invisible area are identified accordingly.

[0157] At block 980, the interface engine performs pattern extraction. According to one or more embodiments, pattern extraction can be anchoring elements of a particular type to elements that a user can interact with. In this manner, through anchoring, the interface engine links the DOM text to the corresponding type of the relevant UI element, as determined by the CV analysis that connects or creates an association between the domain and the label.

[0158] As an example of the ML model of the interface engine, the extracted DOM and types of all input domains are fed into the ML model. The ML model of the interface engine determines the relevant DOM text (e.g., HTML structure) associated with the various input domains. To determine the relevant DOM text, the interface engine embeds the information by encoding and merging the translation and scale-invariant features of the ML model (note that these features can be based on information already extracted by the interface engine). The translation and scale-invariant features are then fed through the various layers of the ML model, such as appropriate attention, convolution, full connection, etc. The predictor of the ML model calculates logical values ​​(logits) that are responsible for assigning appropriate relationships between input domains and labels.

[0159] Go to Fig.11, depicting an interface 1100 according to one or more embodiments. Interface 1100 is an example of a page shown within the boundaries of an interface frame concurrent with the edges of a screen. Thus, a visible area 1101 of a page of interface 1100 is within the boundaries of the interface frame and the edges of the screen, while invisible areas are not shown. Interface 1100 includes a plurality of elements, such as elements 1122, 1123, 1124, 1125, 1126, 1127, and 1128. Note that element 1128 is a scroll bar that can be manipulated to move invisible areas of the page into portions of visible area 1101. Note that before element 1128 is manipulated, elements 1122, 1123, 1124, 1125, 1126, 1127, and 1128 are not Fig.10 part of the visible area 1001. Therefore, the interface engine is from Fig.10 An example of an invisible area identification element (e.g., element 1140) of interface 1000 having a Fig.10 The HTML structure of element 1020 is similar to that of the scroll view. In this regard, interface 1100 shows that element 1140 is identified as a drop-down input field.

[0160] Go to Fig.12 , depicts example code 1200 according to one or more embodiments. Example code 1200 is an example of an HTML structure of a DOM for a drop-down domain in a visible area.

[0161] Problems may arise when UI elements in invisible areas do not have similar input domains in the visible area. Conventional clipboard techniques would require the user to scroll to the input domain in the invisible area to take a screenshot for CV analysis to determine the input domain type. According to one or more embodiments, the interface engine can avoid scrolling, and the extrapolation of structural elements can be done using a cache. Therefore, the interface solves the problem that UI elements in invisible areas do not have similar input domains in the visible area.

[0162] Back to Fig. 9 , at box 990, the interface engine enters data. The interface engine extracts data from a source to paste into at least one or more target UI elements. According to one or more embodiments, the interface engine enters the data from the source into one or more target UI elements. That is, once the domain is identified, the interface engine can use various methods to enter the data into the input field, such as by using RPA. Data entry can be completed by the interface engine with or without scrolling the page. The interface engine can receive an option for scrolling so that the automation of data entry can be viewed. Data can be entered into the same long form and / or a similar long form (e.g., a destination form) of a subsequent page.

[0163] Fig.13 A flow chart illustrating a process 1300 according to one or more embodiments is depicted. According to one or more embodiments, Fig.13 The process 1300 performed in the embodiment may be performed by a computer program that encodes instructions for a processor (or processors). In general, the process 1300 provides for detecting a display (e.g., Figure 5 Example operations of an interface engine for displaying a UI element that is not visible on a display 525 or a screen presenting a UI including the UI element. For ease of explaining process 1300, a page (e.g., a web page) is displayed by a UI presented on a display, and the page is larger than the UI. Therefore, the page has invisible and visible areas.

[0164] Process 1300 begins at block 1302, where the interface engine detects a framework. According to one or more embodiments, the interface engine can identify a framework by injecting "detect-Framework.ts" into a page (e.g., PageWorld). One or more examples of frameworks include, but are not limited to, SAP and Workday.

[0165] At block 1306, the interface engine starts CV analysis. According to one or more embodiments, the interface engine may initiate CV analysis of visible content of the page. CV analysis may be used to detect element types through CV.

[0166] At box 1310, the interface engine detects one or more layers. Note that a page can have multiple layers forming a stacked structure. For example, interacting with a UI element (e.g., a non-hidden element) can reveal other UI elements (e.g., hidden elements). In addition, since the UI only presents UI elements on the top layer, CV analysis cannot capture all elements (e.g., covered elements). For example, according to one or more embodiments, the interface engine can interact with UI elements (in other detected layers) that are covered by other elements. The remaining operations of process 1300 can be applied to each layer detected by the interface engine and can be repeated or looped according to the requirements of the interface engine. One or more of the remaining operations of process 1300 can be performed by the interface engine and can be performed in any order.

[0167] At box 1314, the interface engine extracts data. According to one or more embodiments, the interface engine can use "semantic-extractData.ts" to extract all non-hidden elements from the DOM. The interface engine can also create a WebElementData array that can be used to calculate the DOM tree and the pattern. According to one or more embodiments, the interface engine can extract data by sequentially refining the DOM tree of the page to reach the pattern. The pattern can include a mapping between anchor text and controls (e.g., input fields (input_field), drop-down lists (dropdown) and the like). Examples of patterns include: "First Name" -> input_field1; "Surname" -> input_field2; "Country" -> dropdown1; Note that the pattern element can include an input element, which, when interacted with, enables new features to appear within the UI.

[0168] At block 1318, the interface engine determines nodes (eg, UiNodes). According to one or more embodiments, the interface engine may obtain UiNodes for all WebElementData.

[0169] At block 1322, the interface engine determines the location. According to one or more embodiments, the interface engine may obtain the location of all WebElementData.

[0170] At block 1330, the interface engine determines the element type through the DOM. According to one or more embodiments, the interface engine can use information from the DOM to detect the element type, such as label, type, role, and aria attributes.

[0171] At box 1334, the interface engine determines the element type through the framework. According to one or more embodiments, the interface engine can use framework-specific rules to detect element types. Framework-specific rules define a set of conditions for finding or identifying elements. According to one or more embodiments, as an example, each website has a DOM tree. In addition, each element of the DOM tree has tags and attributes, and may have one or more sub-items. Framework-specific rules define a set of conditions for finding or identifying elements based on tags, attributes, or sub-items. For example, a first rule can identify elements with a specific tag (e.g., x=input), and a second rule can identify elements with specific attributes (e.g., data automation identification with a specific value, input box, date selection, etc.). Some of the framework-specific rules are managed by pages / documents. If a page / document has ten (10) different control types, the framework-specific rules can include ten (10) different rules.

[0172] At box 1338, the interface engine determines the element type through the CV. According to one or more embodiments, the interface engine can use information from the CV analysis to detect the element type. In addition, according to one or more embodiments, the interface engine can create an element definition for the CV element. The element definition is unique to the design of the page and can be composed of all tags, attributes, sub-items, and sub-tags and attributes. The interface engine can match the content in the bounding box (e.g., the border around the element that appears during the design process) with the element in the page.

[0173] At block 1342, the interface engine determines the element type through the element definition. According to one or more embodiments, the interface engine can use the element definition to detect the element type. Note that CV works at the screenshot level, but there are elements in the page that are off-screen. The interface engine works off-screen to identify other elements, such as elements in the view, based on the element's tags and attributes.

[0174] At block 1346, the interface engine populates the automatic anchor. According to one or more embodiments, the interface engine can use DOM element attributes to detect the anchor of the element and then provide the corresponding value to one or more areas of the page. DOM element attributes can include, but are not limited to, label[for], aria-labeled by, aria-label, and placeholder.

[0175] At block 1350, the interface engine populates the datetime. According to one or more embodiments, the interface engine can detect the datetime type, datetime component type, and date format based on the properties, classes, and values, and then provide the corresponding values ​​to one or more areas of the page.

[0176] At block 1358, the interface engine cleans the elements. According to one or more embodiments, the interface engine may remove duplicate elements that reference the same logical element (eg, input elements may be detected by both CV and DOM).

[0177] At block 1362, the interface engine enriches the element. According to one or more embodiments, the interface engine can promote information from any internal input element to a parent (container) element (e.g., CV can detect a DIV element representing a text box). For example, during FillElementTypeByCV, a DIV element can be promoted to an InputBox. A DIV element can include hidden <input>, with an aria-labeledby attribute. In the process of enriching the UI element, the anchor may be promoted to a container element. Note that blocks 1302 to 1362 may be performed for each layer and / or for each iframe within any layer. The resulting DOMRawUiNodeElements may then be aggregated.

[0178] At block 1366, the interface engine creates a data structure, such as a DOM tree. According to one or more embodiments, the interface engine may generate a simplified DOM tree structure from an aggregate array of DOMRawUiNodeElements.

[0179] At block 1370, the interface engine collects DOM relationships. According to one or more embodiments, the interface engine can use a DOM relationship model (e.g., a relationship anchor generated by an ML model) to detect anchors of elements. The DOM relationship model is a custom ML model of the interface engine that finds anchors of input elements from a simplified DOM tree structure. Therefore, the interface engine can infer a "schema" from a simplified DOM by using a custom ML model.

[0180] At block 1374, the interface engine collects the drop-down information. According to one or more embodiments, the interface engine can find the relative position of the drop-down arrow and populate the DropdownInfo property of the DOMRawUiNodeElement.

[0181] At box 1378, the interface engine collects geometric anchors. According to one or more embodiments, an anchor is a label for input. When the interface engine detects something in the area (e.g., an input box), the interface engine will give it (e.g., an input box) a label. Anchors (which are automatic anchors) can be inferred from the DOM because the DOM has text to help mark it. Machine learning anchors can include when the interface engine provides the DOM of the page to the model, and the model determines which anchors are used for which inputs. Geometric anchors are labels (e.g., with fixed rules) that use the DOM and the position that matches the element to determine the position and alignment. For input UI elements that do not have anchors, the interface engine assigns geometric anchors if possible. In addition, the interface engine's position-based, alignment-based, and hierarchical-based algorithms can be used to search for geometric anchors (e.g., a DOM tree can include containers and text and input boxes within the container). For example, the interface engine's position-based algorithm is a process for finding a position for an input UI element within a page so that a geometric anchor can be assigned.

[0182] At block 1382, the interface engine collects the option groups. According to one or more embodiments, the interface engine may group all radio button elements and search for labels for the groups.

[0183] At block 1386, the interface engine collects the values. According to one or more embodiments, for each input element (e.g., a drop-down list, a chat box, an option, and anything that can be interacted with and set a value), the interface engine collects the value. For example, the interface engine retrieves the content provided in the input box for the name. According to one or more embodiments, the interface engine can populate the value attribute of the DOMRawUiNodeElement.

[0184] At block 1390, the interface engine collects the date-time group. For example, the date-time group may be a set of alphanumeric characters in a specified format that represents one or more of a year, a month, a day of a month, an hour of a day, a minute of an hour, and a time zone. According to one or more embodiments, the interface engine may group all input elements of the date-time group.

[0185] At block 1394, the interface engine determines the pattern. According to one or more embodiments, the interface engine can create a pattern from DOMRawUiNodeElements. Note that the pattern can include a mapping between anchor text and controls as described herein.

[0186] According to one or more embodiments, processes 900 and 1300 herein may be implemented as activities in a designer application, such as UiPath Studio TM . This activity can be used for semantic copy and paste from a set of source domains to a set of destination domains (e.g., in visible / invisible areas of the screen). In addition, all caches captured during CV analysis can also be fed into the activity for dynamically detecting UI elements in invisible areas. In addition, processes 900 and 1300 in this article can be implemented with respect to task mining.

[0187] A computer program may be implemented in hardware, software or a hybrid implementation. A computer program may consist of modules that are in operational communication with each other and are designed to transfer information or instructions for display. A computer program may be configured to operate on a general purpose computer, an ASIC or any other suitable device.

[0188] It is easy to understand that the components of various embodiments, as generally described and illustrated in the drawings herein, can be arranged and designed in various different configurations. Therefore, as shown in the drawings, the detailed description of the embodiments is not intended to limit the scope of protection claimed, but only represents selected embodiments.

[0189] The features, structures, or characteristics described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, throughout the specification, references to "certain embodiments," "some embodiments," or similar language indicate that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, the appearance of the phrases "in certain embodiments," "in some embodiments," "in other embodiments," or similar language throughout this specification do not necessarily all refer to the same set of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0190] It should be noted that in this specification, references to features, advantages, or similar language do not imply that all features and advantages that may be realized should be present or exist in any single embodiment. Rather, language referring to features and advantages is understood to mean that a particular feature, advantage, or characteristic described in conjunction with an embodiment is included in one or more embodiments. Therefore, discussions of features and advantages and similar language throughout this specification may, but do not necessarily, refer to the same embodiment.

[0191] In addition, the features, advantages and characteristics of one or more embodiments described herein may be combined in any suitable manner. Those skilled in the relevant art will recognize that the present invention may be practiced without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages that may not be present in all embodiments may be recognized in certain embodiments.

[0192] Those skilled in the art will readily appreciate that the present disclosure may be practiced with steps in different orders and / or with hardware elements in configurations different from those disclosed. Therefore, although the present disclosure has been described based on these preferred embodiments, certain modifications, variations, and alternative configurations will be clear to those skilled in the art while remaining within the spirit and scope of the present disclosure. Therefore, in order to determine the scope and limits of the present disclosure, reference should be made to the appended claims.

Claims

1. A method for detecting and automatically filling one or more user interface (UI) elements of a page that is not visible on a screen, the method being performed by an interface engine implemented as a computer program within a computing environment, the method comprising: The interface engine analyzes the document object model DOM of the page to extract DOM text of one or more related UI elements among the one or more UI elements; Performing computer vision (CV) analysis by the interface engine to determine one or more types of the one or more related UI elements; as well as One or more target UI elements of the non-visible area of ​​the page are extrapolated by the interface engine from the one or more related UI elements. 2 . The method of claim 1 , wherein the method comprises linking the DOM text to the one or more target UI elements based on the one or more types and the one or more related UI elements determined by the CV analysis. 3 . The method of claim 1 , wherein the extrapolating of the one or more target UI elements by the interface engine comprises searching the DOM for similar Hypertext Markup Language (HTML) structures. The method of claim 1 , wherein the one or more UI elements include one or more input fields of the page. The method of claim 1 , wherein the interface engine extracts data from a source to paste into at least the one or more target UI elements. The method according to claim 1 , wherein the one or more types of the one or more related UI elements comprise a DOM type, a frame type, a CV type, or an element definition type. 7 . The method of claim 1 , wherein the executing of the CV by the interface engine comprises capturing a visible area of ​​the page to identify the one or more types of the one or more related UI elements in the visible area. The method of claim 1 , wherein the interface engine inputs data into the one or more target UI elements without scrolling the page.

9. The method of claim 1, wherein the page comprises a long form.

10. The method of claim 1, wherein the non-visible area of ​​the page includes a subsequent destination form.

11. A system comprising: A memory storing code of an interface engine for detecting and automatically filling one or more user interface (UI) elements of a page that is not visible on the screen; as well as at least one processor configured to execute the code to cause, within the system: The interface engine analyzes the document object model DOM of the page to extract DOM text of one or more related UI elements among the one or more UI elements; Performing computer vision (CV) analysis by the interface engine to determine one or more types of the one or more related UI elements; as well as One or more target UI elements of the non-visible area of ​​the page are extrapolated by the interface engine from the one or more related UI elements. 12 . The system of claim 11 , wherein the interface engine links the DOM text to the one or more target UI elements based on the one or more types and the one or more related UI elements determined by the CV analysis. 13 . The system of claim 11 , wherein the extrapolation of the one or more target UI elements by the interface engine comprises searching the DOM for similar Hypertext Markup Language (HTML) structures. The system of claim 11 , wherein the one or more UI elements include one or more input fields of the page.

15. The system of claim 11, wherein the interface engine extracts data from a source to paste into at least the one or more target UI elements. 16 . The system of claim 11 , wherein the one or more types of the one or more related UI elements comprise a DOM type, a frame type, a CV type, or an element definition type.

17. The system of claim 11, wherein the execution of the CV by the interface engine comprises capturing a visible area of ​​the page to identify the one or more types of the one or more related UI elements in the visible area.

18. The system of claim 11, wherein the interface engine inputs data into the one or more target UI elements without scrolling the page.

19. The system of claim 11, wherein the page comprises a long form.

20. The system of claim 11, wherein the non-visible area of ​​the page includes a subsequent destination form.

Citation Information

Patent Citations

  • Long running workflows for document processing using robotic process automation

    US10860905B1

  • Robot access control and governance for robotic process automation

    US20220011732A1