Graphical element detection using the combined serial and delayed parallel execution unified target technique, the default graphical element detection technique, or both
The combined serial and delayed parallel execution unified target technique in RPA dynamically switches between detection methods for improved graphical element detection, addressing inefficiencies in current techniques and ensuring reliable UI element identification.
Patent Information
- Application Number
- JP2020553455
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2020-09-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Current graphical element detection techniques for robotic process automation (RPA) are not optimal for all scenarios and often require improved approaches.
A combined serial and delayed parallel execution unified target technique is employed, utilizing multiple graphical element detection methods such as selectors, computer vision, and optical character recognition, with a default UI element detection configuration at the application or UI type level, allowing for dynamic adjustment and parallel execution of detection techniques.
Enhances the accuracy and efficiency of graphical element detection in RPA by adaptively switching between detection methods, ensuring robust identification of UI elements even in varying environments.
Smart Images

Figure 0007742069000001 
Figure 0007742069000002 
Figure 0007742069000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Non-Provisional Patent Application No. 17 / 014,171, filed September 8, 2020. The subject matter of this previously filed application is incorporated herein by reference in its entirety.
[0002] The present invention relates generally to graphical element detection, and more particularly to a combined serial and deferred parallel execution unified target technique at the application and / or UI type level, performing graphical element detection using default UI element detection technique configurations, or both. [Background technology]
[0003] For robotic process automation (RPA) in UI, graphical element detection may be performed for each UI action using selectors, computer vision (CV), or optical character recognition (OCR). However, these techniques are typically applied individually and are not optimal for all scenarios. Therefore, improved approaches may be beneficial. Summary of the Invention
[0004] Certain embodiments of the present invention may provide solutions to problems and needs in the field that have not yet been fully identified, appreciated, or solved by current graphical element detection technology. For example, some embodiments of the present invention relate to graphical element detection using a combined serial and delayed parallel execution unified target technique. Certain embodiments relate to a default UI element detection technology configuration at the application and / or UI type level. This configuration may be used to detect UI elements at runtime.
[0005] In an embodiment, a computer-implemented method for detecting graphical elements in a UI includes receiving, by a designer application, a selection of an activity in an RPA workflow configured to perform graphical element detection using a unified target technique. The computer-implemented method also includes receiving, by the designer application, changes to the activity's unified target technique and configuring, by the designer application, the activity based on the changes. The unified target technique is a combined serial and delayed parallel execution unified target technique configured to employ multiple graphical element detection techniques.
[0006] In another embodiment, a computer program is stored on a non-transitory computer-readable medium. The computer program is configured, when executed, by at least one processor to analyze a UI to identify UI element attributes and compare the UI element attributes to UI descriptor attributes for an activity of an RPA workflow using one or more initial graphical element detection techniques. If no match is found using the one or more initial graphical element detection techniques during a first time period, the computer program is configured, when executed, by the at least one processor to run one or more additional graphical element detection techniques in parallel with the one or more initial graphical element detection techniques.
[0007] In yet another embodiment, a computer program is stored on a non-transitory computer-readable medium. The computer program is configured, when executed, by at least one processor to analyze a UI to identify UI element attributes and compare the UI element attributes to UI descriptor attributes for an activity of the RPA workflow using one or more initial graphical element detection techniques. If no match is found using the one or more initial graphical element detection techniques during a first time period, the computer program is configured to cause the at least one processor to perform one or more additional graphical element detection techniques in place of the one or more initial graphical element detection techniques.
[0008] In yet another embodiment, a computer-implemented method for detecting graphical elements in a UI includes receiving, by an RPA designer application, a selection of an application or UI type. The computer-implemented method also includes receiving and saving, by the RPA designer application, a default targeting method settings configuration. The computer-implemented method further includes receiving, by the RPA designer application, an indication of a screen to be automated. The screen is associated with the selected application or UI type. Furthermore, the computer-implemented method includes automatically pre-configuring, by the RPA designer application, a default targeting method settings for the selected application or UI type.
[0009] In another embodiment, a computer program is stored on a non-transitory computer-readable medium. The computer program is configured to cause at least one processor to automatically pre-configure default targeting method settings for a selected application or UI type. The computer program is also configured to cause the at least one processor to receive changes to the default targeting method settings for the application or UI type and configure the default targeting method settings according to the changes. The computer program is further configured to cause the at least one processor to configure one or more activities in the RPA workflow using the default targeting method settings for the application or UI type.
[0010] In yet another embodiment, a computer program is stored on a non-transitory computer-readable medium. The computer program is configured to cause at least one processor to automatically pre-configure default targeting method settings for an application or UI type with a designer application and configure one or more activities in an RPA workflow using the default targeting method settings for the application or UI type. The computer program is also configured to cause the at least one processor to generate an RPA robot for performing the RPA workflow including the one or more configured activities. [Brief explanation of the drawings]
[0011] So that the advantages of particular embodiments of this invention may be readily understood, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments which are illustrated in the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope, but the invention will be described and explained with additional specificity and detail through the use of the following accompanying drawings, in which:
[0012] [Figure 1] FIG. 1 is an architectural diagram illustrating a robotic process automation (RPA) system, according to an embodiment of the present invention.
[0013] [Figure 2] FIG. 1 is an architectural diagram illustrating a deployed RPA system according to an embodiment of the present invention.
[0014] [Figure 3] FIG. 2 is an architecture diagram illustrating the relationships between designers, activities, and drivers according to an embodiment of the present invention.
[0015] [Figure 4]FIG. 1 is an architectural diagram illustrating an RPA system according to an embodiment of the present invention.
[0016] [Figure 5] FIG. 1 is an architectural diagram illustrating a computing system configured to perform graphical element detection using a combined serial and deferred parallel execution unified targeting technique and / or one or more default targeting methods configured by application or UI type, according to an embodiment of the present invention.
[0017] [Figure 6A] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6B] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6C] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6D] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6E] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6F] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention. [Figure 6G] 1 illustrates a unified target configuration interface for an RPA designer application, according to an embodiment of the present invention.
[0018] [Figure 7A] 10 illustrates a targeting method configuration interface for configuring targeting methods at the application and / or UI type level, according to an embodiment of the present invention. [Figure 7B]10 illustrates a targeting method configuration interface for configuring targeting methods at the application and / or UI type level, according to an embodiment of the present invention. [Figure 7C] 10 illustrates a targeting method configuration interface for configuring targeting methods at the application and / or UI type level, according to an embodiment of the present invention.
[0019] [Figure 8] 1 is a flowchart illustrating a process for configuring a unified target function for an activity in an RPA workflow according to an embodiment of the present invention.
[0020] [Figure 9A] 1 is a flowchart illustrating a process for graphical element detection using a combined serial and delayed parallel execution unified target technique, according to an embodiment of the present invention. [Figure 9B] 1 is a flowchart illustrating a process for graphical element detection using a combined serial and delayed parallel execution unified target technique, according to an embodiment of the present invention.
[0021] [Figure 10A] 10A-10C are flowcharts illustrating design-time and run-time portions, respectively, of a process for configuring default graphical element detection techniques at the application and / or UI type level and performing graphical element detection in accordance with an embodiment of the present invention. [Figure 10B] 10A-10C are flowcharts illustrating design-time and run-time portions, respectively, of a process for configuring default graphical element detection techniques at the application and / or UI type level and performing graphical element detection in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0022] Unless otherwise noted, like reference characters denote corresponding features consistently throughout the accompanying drawings. Detailed Description of the Embodiments
[0023] Some embodiments relate to graphical element detection using a combined serial and delayed parallel execution unified target technique that potentially uses multiple graphical element detection techniques (e.g., selector, CV, OCR, etc.). In this specification, "graphical element" and "UI element" are used interchangeably. Their core UI descriptors identify UI elements (e.g., text fields, buttons, labels, menus, checkboxes, etc.). Some types of UI descriptors include, but are not limited to, selectors, CV descriptors, image matching descriptors, OCR descriptors, unified target descriptors that may utilize multiple different types of UI descriptors in serial or parallel fashion, etc. The UI descriptors can be used to compare attributes of a given UI descriptor with UI element attributes discovered at runtime within the UI.
[0024] In some embodiments, the UI descriptor stores attributes of each UI element and its parent, for example, in an extensible markup language (XML) fragment. At runtime, UI element attributes found in the UI may be searched for matches with attributes of each RPA workflow activity, and if an exact match or a "close enough" match is found within a matching threshold, the UI element may be identified and interacted with accordingly. Attributes may include text-based identifiers (IDs), classes, roles, etc. In the case of CV, attributes may include the type of target element and its relationship to one or more anchor elements, which may be used in a multi-anchor matching approach. In the case of OCR, attributes may include, for example, text in the form of stored strings and text discovered via OCR where stored strings are fuzzy matched during runtime. Any suitable attribute and graphical element detection techniques may be used without departing from the scope of the present invention.
[0025] As used herein, a "screen" is an image of an application UI or a portion of an application UI at a point in time. In some embodiments, UI elements and screens may be further distinguished into specific types of UI elements (e.g., buttons, check boxes, text fields, etc.) and screens (e.g., top windows, modal windows, pop-up windows, etc.).
[0026] Some embodiments use UI descriptors that store attributes of UI elements and their parents in XML fragments. In modern computing systems, operating systems typically represent each user interface as a hierarchical data structure commonly referred to as a UI tree. An exemplary UI tree might include the Document Object Model (DOM) underlying a web page rendered by a web browser application.
[0027] A selector is a type for a UI descriptor that can be used to find a UI element in some embodiments. A selector, in some embodiments, has the following structure: <node_1 / ><node_2 / > ...<node_N / >
[0028] The final node represents the target GUI element, and all previous nodes represent the parents of that element.<node_1> is usually called the root node and represents the top window of the application.
[0029] Each node may have one or more attributes that assist in the correct identification of the particular level of the selected application. Each node, in some embodiments, has the following format: <ui_system attr_name_1=’attr_value_1’ ... attr_name_N=’attr_value_N’ / >
[0030] All attributes may have values assigned, and attributes with constant values may be chosen because changing the value of an attribute every time the application launches may prevent the selector from correctly identifying the associated element.
[0031] A UI descriptor is a set of instructions for locating UI elements. In some embodiments, the UI descriptor is an encapsulated data / structure format that includes UI element selector(s), anchor selector(s), CV descriptor(s), OCR descriptor(s), unified target descriptor(s) that combine two or more types of UI descriptors, screen image capture (context), element image capture, other metadata (e.g., application and application version), or a combination thereof. The encapsulated data / structure format may be extensible with future updates to the platform and is not limited to the above definition. Any suitable UI descriptor for identifying UI elements on a screen may be used without departing from the scope of the present invention. UI descriptors may be extracted from activities in an RPA workflow and added to a structured schema that groups UI descriptors by UI application, screen, and UI element.
[0032] In some embodiments, UI descriptors may work with a unified target that encompasses multiple or all UI element detection mechanisms by which image detection and definition are performed. The unified target may merge multiple techniques for identifying and automating UI elements into a single, cohesive approach. The unified target descriptor may chain multiple types of UI descriptors in series, use them in parallel, or use at least one technique (e.g., a selector) first for a certain period of time, and then, if the first technique does not find a match within that period, run at least one other technique in parallel or alternately. In some embodiments, the unified target descriptor may function like a finite state machine (FSM), applying a first UI descriptor mechanism in a first context, a second UI descriptor mechanism in a second context, and so on. The unified target prioritizes selector-based and driver-based UI detection mechanisms, and in some embodiments, may resort to CV, image matching, and / or other mechanisms to find graphical elements if the first two mechanisms are unsuccessful.
[0033] In some embodiments, fuzzy matching may be employed, where one or more attributes must match within a certain range and with a certain accuracy (e.g., 70% match, 80% match, 99% match, etc.) using a string metric (e.g., Levenshtein distance, Hamming distance, Jaro-Winkler distance, etc.), combinations thereof, etc. Those skilled in the art will appreciate that a similarity measure can quantify not only the amount of similarity but also the amount of mismatch between two attribute values. Furthermore, in various embodiments, a similarity threshold may represent a maximum amount of mismatch or a minimum amount of similarity required for a match.
[0034] Depending on the selected method for calculating the similarity measure, the similarity threshold may have various interpretations. For example, the similarity threshold may indicate the maximum character count that can differ between two strings, or the fractional degree of mismatch calculated as a percentage of the total character count (e.g., the length of the combined strings). In some embodiments, the similarity threshold may be rescaled to a predetermined interval, such as between 0 and 1, between 0 and 100, between 7 and 34, etc. In one non-limiting example, a relatively high similarity threshold (e.g., close to 1 or 100%) indicates a requirement for near-perfect match, i.e., the values of the fuzzy attributes in the run-time target are allowed to deviate only very slightly from the values of each attribute in the design-time target. On the other hand, if the similarity threshold is relatively low (e.g., close to 0), nearly all values of each fuzzy attribute are considered to match.
[0035] In certain embodiments, matching tolerances may vary by attribute criteria. For example, an exact match may be required for one or more attributes (e.g., finding a specific, exact name may be desired), while fuzzy matching may be performed for one or more other attributes. The number and / or type of attributes used from each graphical element detection technique may, in some embodiments, be custom specified by the RPA developer.
[0036] In some embodiments, attributes may be stored as attribute-value pairs and / or attribute-value-tolerance pairs (e.g., fuzzy matching). The attribute-value pairs may, in some embodiments, indicate the name and type of the UI element represented by the respective node. However, those skilled in the art will understand that there may be multiple ways to represent the location of a particular node in a UI tree other than as a list of attribute-value pairs without departing from the scope of the present invention.
[0037] These attribute-value pairs and / or attribute-value-tolerance pairs may be stored in tags in some embodiments, and each tag may include a string of characters with the sequence bookended by an implementation-specific delimiter (e.g., starting with "<" and ending with " / >"). The attribute-value pairs may, in some embodiments, indicate the name and type of the UI element represented by the respective node. However, those skilled in the art will understand that there may be multiple ways to represent the location of a particular node in a UI tree other than as a list of attribute-value pairs without departing from the scope of the present invention.
[0038] To enable successful and ideally unambiguous identification by the RPA robot, some embodiments represent each UI element using an element ID that characterizes the respective UI element. In some embodiments, the element ID indicates the location of the target node in the UI tree, where the target node represents the respective UI element. For example, the element ID may identify the target node / UI element as a member of a subset of selected nodes. The subset of selected nodes can form a line of descent through the UI tree, i.e., each node is either an ancestor or a descendant of another node.
[0039] In some embodiments, the element ID includes an ordered sequence of node indicators that trace a genealogical path through the UI tree, terminating at a respective target node / UI element. Each node indicator may represent a member of the respective UI object hierarchy and its position in the sequence that matches the respective hierarchy. For example, each member of the sequence may represent a descendant (e.g., a child node) of the previous member, which in turn is a descendant (e.g., a child node) of the next member, etc. In one HyperText Markup Language (HTML) example, element IDs representing individual form fields may indicate that each form field is a child of the HTML form, which in turn is a child of a particular section of a web page, etc. The genealogy need not be complete in some embodiments.
[0040] In some embodiments, one or more multi-anchor matching attributes may be used. Anchors are other UI elements that may be used to help uniquely identify a target UI element. For example, if a UI contains multiple text fields, searching the text fields alone is insufficient to uniquely identify a given text field. Therefore, in some embodiments, additional information is sought to uniquely identify a given UI element. Using the example of a text field, a text field for entering a first name may appear to the right of a label that reads "First Name." This first name label may be set as an "anchor" to help uniquely identify the "target" text field.
[0041] In some embodiments, various positions and / or geometric associations between the target and anchor may be used, potentially within one or more tolerances, to uniquely identify the target. For example, the center of the anchor and target's bounding box may be used to define a line segment. This line segment may then be required to have a specific length within a tolerance and / or slope within a tolerance to uniquely identify the target using the target / anchor pair. However, any desired positions for the positions associated with the target and / or anchor may be used in some embodiments without departing from the scope of the present invention. For example, the point for drawing the line segment may be the center, upper left corner, upper right corner, lower left corner, lower right corner, any other position on the boundary of the bounding box, any position within the bounding box, a position outside the bounding box, etc., as identified in connection with the bounding box characteristics. In certain embodiments, the target and one or more anchors may have different positions within or outside their bounding boxes that are used for geometric matching.
[0042] As can be seen, a single anchor may not always be sufficient to uniquely identify a target element on the screen with a certain degree of reliability. For example, consider a web form that displays two text fields for entering a first name, each to the right of a "First Name" label at different locations on the screen. In this example, one or more additional anchors may be useful in uniquely identifying a given target. Geometric characteristics between the anchor and the target (e.g., length, angle, and / or relative position of a line segment with a tolerance) may be used to uniquely identify the target. The user may be required to continue adding anchors until the match strength for the target exceeds a threshold.
[0043] As used herein, the terms “user” and “developer” are used interchangeably. A user / developer may or may not have programming and / or technical knowledge. For example, in some embodiments, a user / developer may create an RPA workflow by configuring activities within the RPA workflow without manual coding. In certain embodiments, this may be done, for example, by clicking, dragging, and dropping various functions.
[0044] In some embodiments, the default UI element detection technique (also referred to herein as a "targeting method") may be configured at the application and / or UI type level. A UI element detection technique that works well for a given application and / or UI type may not work well for another application and / or UI type. For example, a technique that works well in a Java window may not work well in a web browser window. Thus, a user may configure an RPA robot to use the most effective technique(s) for a given application and / or UI type.
[0045] Certain embodiments may be employed in robotic process automation (RPA). FIG. 1 is an architectural diagram illustrating an RPA system 100 according to an embodiment of the present invention. The RPA system 100 includes a designer 110 that enables developers to design and implement workflows. The designer 110 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 110 can facilitate the development of automation projects, which are graphical representations of business processes. Simply put, the designer 110 facilitates the development and deployment of workflows and robots.
[0046] Automation projects enable rule-based process automation by giving developers control over the execution order and relationships between a custom set of steps developed in a workflow, defined herein as "activities." One commercial example of an embodiment of the designer 110 is UiPath Studio™. Each activity may include an action such as clicking a button, reading a file, or writing to a log panel. In some embodiments, workflows may be nested or embedded.
[0047] Workflow types may include, but are not limited to, sequences, flowcharts, FSMs, and / or global exception handlers. Sequences may be particularly well-suited for linear processes, allowing the flow of one activity from another without cluttering the workflow. Flowcharts may be particularly well-suited for more complex business logic, allowing for the integration of decisions and the connection of activities in more diverse ways through multiple branching logic operators. FSMs may be particularly well-suited for large workflows. FSMs may use a finite number of states during their execution that are triggered by conditions (i.e., transitions) or activities. Global exception handlers may be particularly well-suited for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0048] Once a workflow is developed in Designer 110, the execution of the business process is orchestrated by Conductor 120, which coordinates one or more Robots 130 that execute the workflow developed in Designer 110. One commercial example of an embodiment of Conductor 120 is UiPath Orchestrator™. Conductor 120 facilitates the management of the creation, monitoring, and deployment of resources in an environment. Conductor 120 may act as, or one of the integration points with, third-party solutions and applications.
[0049] The conductor 120 may manage all robots 130, connecting and running them from a centralized point. Types of robots 130 that may be managed include, but are not limited to, attended robots 132, unattended robots 134, development robots (similar to unattended robots 134 but used for development and testing purposes), and non-production robots (similar to attended robots 132 but used for development and testing purposes). Attended robots 132 may be triggered by user events or scheduled to occur automatically, and may operate side by side with humans on the same computing system. Attended robots 132 may be used with the conductor 120 for centralized process deployment and logging media. Attended robots 132 may assist human users in accomplishing various tasks and may be triggered by user events. In some embodiments, processes cannot be initiated from the conductor 120 on this type of robot, and / or they cannot be run under a locked screen. In certain embodiments, the attended robot 132 can only be launched from the robot tray or from a command prompt. The attended robot 132 preferably operates under human supervision in some embodiments.
[0050] Unattended robots 134 operate unattended in virtual environments or on physical machines and can automate many processes. Unattended robots 134 may be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types may be performed from Designer 110 in some embodiments. Both attended and unattended robots can automate a variety of systems and applications, including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those produced by SAP®, Salesforce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0051] The conductor 120 may have various capabilities, including, but not limited to, provisioning, deployment, versioning, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining connections between robots 130 and the conductor 120 (e.g., web applications). Deployment may include ensuring the correct delivery of package versions to robots 130 assigned for execution. Versioning, in some embodiments, may include managing unique instances of some processes or configurations. Configuration may include maintaining and delivering robot environments and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include tracking robot identification data and maintaining user permissions. Logging may include saving and indexing logs in a database (e.g., an SQL database) and / or another storage mechanism (e.g., ElasticSearch®, which stores large data sets and provides the ability to quickly query them). The conductor 120 may provide interconnectivity by acting as a centralized point of communication for third-party solutions and / or applications.
[0052] Robots 130 are execution agents that execute workflows built in designer 110. One commercial example of some embodiments of robot(s) 130 is UiPath Robots™. In some embodiments, robots 130 install the Microsoft Windows Service Control Manager (SCM) management service by default. As a result, such robots 130 can open interactive Windows sessions under the local system account and may have Windows service rights.
[0053] In some embodiments, a robot 130 can be installed in user mode, meaning that for such a robot 130, the robot has the same rights as the user to whom it is installed. This feature can also be used for high-density (HD) robots, ensuring maximum utilization of each machine. In some embodiments, either type of robot 130 can be configured in an HD environment.
[0054] In some embodiments, the robot 130 is divided into multiple components, each specialized for a specific automation task. In some embodiments, the robot components include, but are not limited to, an SCM-managed robot service, a user-mode robot service, an executor, an agent, and a command line. The SCM-managed robot service manages and monitors Windows sessions and acts as a proxy between the conductor 120 and the execution host (i.e., the computing system on which the robot 130 runs). These services are responsible for managing credentials for the robot 130. A console application is launched by the SCM under Local System.
[0055] The user-mode robot service in some embodiments manages and monitors Windows sessions and acts as a proxy between the conductor 120 and the execution host. The user-mode robot service may be entrusted with and manage the credentials of the robot 130. If the SCM management robot service is not installed, a Windows application may be launched automatically.
[0056] An Executor may run a given job under a Windows session (i.e., run a workflow). An Executor may be aware of per-monitor dots-per-inch (DPI) settings. An Agent may be a Windows Presentation Foundation (WPF) application that displays available jobs in a system tray window. An Agent may be a client of a Service. An Agent may ask to start or stop a job or change settings. A Command Line is a client of a Service. A Command Line is a console application that can request the start of a job and wait for its output.
[0057] As described above, the separation of robot 130 components helps developers, support users, and computing systems more easily implement, identify, and track what each component is doing. In this way, special behaviors can be configured for each component, such as setting different firewall rules for executors and services. Executors may always be aware of per-monitor DPI settings in some embodiments. As a result, workflows may run at any DPI regardless of the configuration of the computing system on which they were created. Also, in some embodiments, projects from designer 110 may be made independent of browser zoom levels. For applications that are not DPI-aware or are intentionally marked as not-aware, some embodiments may disable DPI.
[0058] FIG. 2 is an architecture diagram illustrating a deployed RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 may be or be part of the RPA system 100 of FIG. 1. It should be noted that the client side, the server side, or both may include any desired number of computing systems without departing from the scope of the present invention. On the client side, the robot application 210 includes an executor 212, an agent 214, and a designer 216. However, in some embodiments, the designer 216 may not be running on the computing system 210. The executor 212 executes processes. As shown in FIG. 2, multiple business projects may be running simultaneously. The agent 214 (e.g., a Windows service) is a single connection point for all executors 212 in this embodiment. All messages in this embodiment are logged to the conductor 230, which further processes them via the database server 240, the indexer server 250, or both. As described above with respect to FIG. 1, the executor 212 may be a robotic component.
[0059] In some embodiments, a Robot represents an association between a machine name and a username. A Robot may manage multiple executors simultaneously. In computing systems that support multiple interactive sessions running simultaneously (such as Windows Server 2012), multiple Robots may run simultaneously, each running in a separate Windows session using a unique username. This is referred to as an HD Robot above.
[0060] The agent 214 is also responsible for transmitting the robot's status (e.g., periodically sending "heartbeat" messages to indicate that the robot is still functioning) and downloading any necessary versions of packages to be executed. Communication between the agent 214 and the conductor 230 is, in some embodiments, always initiated by the agent 214. In notification scenarios, the agent 214 may open a WebSocket channel that is later used by the conductor 230 to send commands (e.g., start, stop, etc.) to the robot.
[0061] The server side includes a presentation layer (web application 232, Open Data Protocol (OData) Representational State Transfer (REST) Application Programming Interface (API) endpoint 234, notification and monitoring 236), a service layer (API implementation / business logic 238), and a persistence layer (database server 240, indexer server 250). Conductor 230 includes web application 232, OData REST API endpoint 234, notification and monitoring 236, and API implementation / business logic 238. In some embodiments, most actions a user performs in the conductor 230 interface (e.g., via browser 220) are performed by calling various APIs. Such actions may include, but are not limited to, launching jobs on a robot, adding / removing data from a queue, scheduling jobs to run unattended, etc., without departing from the scope of the present invention. Web application 232 is the visual layer of the server platform. In this embodiment, web application 232 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of the present invention. A user interacts with web pages from web application 232, in this embodiment via browser 220, to perform various operations to control conductor 230. For example, a user may create robot groups, assign packages to robots, analyze per-robot and / or per-process logs, start and stop robots, etc.
[0062] In addition to the web application 232, the conductor 230 also includes a services layer that exposes an OData REST API endpoint 234. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 232 and the agent 214, which in this embodiment is a supervisor of one or more robots on a client computer.
[0063] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functionality. The configuration endpoint, in some embodiments, may be used to define and configure users, permissions, robots, assets, releases, and environments for an application. The logging REST endpoint may be used to log various information, such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by robots to query the version of a package that should be executed when a start job command is used in conductor 230. The queuing REST endpoint may be responsible for managing queues and queue items, such as adding data to a queue, retrieving transactions from a queue, and setting the status of transactions.
[0064] Monitoring REST endpoints may monitor the web application 232 and the agents 214. The notification and monitoring API 236 may be a REST endpoint used to register the agents 214, deliver configuration settings to the agents 214, and send and receive notifications from the server and the agents 214. The notification and monitoring API 236 may use WebSocket communication in some embodiments.
[0065] The persistence layer, in this embodiment, includes a pair of servers: database server 240 (e.g., SQL Server) and indexer server 250. Database server 240 in this embodiment stores configurations for robots, robot groups, associated processes, users, roles, schedules, etc. This information is managed, in some embodiments, via web application 232. Database server 240 may also manage queues and queue items. In some embodiments, database server 240 may store messages logged by robots (in addition to or instead of indexer server 250).
[0066] Optionally in some embodiments, indexer server 250 stores and indexes information logged by the robots. In particular embodiments, indexer server 250 may be disabled via a configuration setting. In some embodiments, indexer server 250 uses ElasticSearch®, a full-text search engine from an open source project. Messages logged by the robots (e.g., using activities such as log messages or line writes) may be sent via logging REST endpoint(s) to indexer server 250, where they are indexed for future use.
[0067] FIG. 3 is an architecture diagram illustrating the relationship 300 between a designer 310, activities 320, 330, and a driver 340, according to an embodiment of the present invention. As can be seen, a developer uses the designer 310 to develop a workflow to be executed by a robot. The workflow may include user-defined activities 320 and UI automation activities 330. In some embodiments, non-text visual components in an image can be identified, which is referred to herein as computer vision (CV). Some CV activities associated with such components include, but are not limited to, click, type, get text, hover, detect element presence, update scope, highlight, and the like. In some embodiments, click identifies an element and clicks on it, for example, using CV, optical character recognition (OCR), fuzzy text matching, and multi-anchors. Type may identify an element using the above and types within the element. Get text may locate specific text and scan it using OCR. Hover may identify an element and hover over it. Detect element presence may check for the presence or absence of an element on the screen using the techniques described above. In some embodiments, there may be hundreds or even thousands of activities that may be implemented in designer 310. However, any number and / or type of activities may be utilized without departing from the scope of the present invention.
[0068] UI automation activities 330 are a subset of specialized low-level activities that are written in low-level code (e.g., CV activities) and facilitate interaction with applications through the UI layer. In particular embodiments, UI automation activities 330 may simulate user input, for example, via window messages. UI automation activities 330 facilitate these interactions through drivers 340 that enable the robot to interact with desired software. For example, drivers 340 may include OS drivers 342, browser drivers 344, VM drivers 346, enterprise application drivers 348, etc.
[0069] Drivers 340 may interact with the OS at a low level, such as by looking for hooks, monitoring keys, etc. They may facilitate integration with Chrome®, IE®, Citrix®, SAP®, etc. For example, a "click" activity plays the same role in these different applications via drivers 340.
[0070] FIG. 4 is an architecture diagram illustrating an RPA system 400, according to an embodiment of the present invention. In some embodiments, the RPA system 400 may be or include the RPA systems 100 and / or 200 of FIGS. 1 and / or 2. The RPA system 400 includes multiple client computing systems 410 that execute robots. The computing systems 410 can communicate with a conductor computing system 420 via web applications running thereon. The conductor computing system 420 can, in turn, communicate with a database server 430 and an optional indexer server 440.
[0071] 1 and 3, it should be noted that while web applications are used in these embodiments, any suitable client and / or server software may be used without departing from the scope of the present invention. For example, a conductor may run a server-side application on a client computing system that communicates with a non-web-based client software application.
[0072] FIG. 5 is an architectural diagram illustrating a computing system 500 configured to perform graphical element detection using a combined serial and delayed parallel execution unified targeting technique and / or one or more default targeting methods configured by application or UI type, according to embodiments of the present invention. In some embodiments, computing system 500 may be one or more computing systems depicted and / or described herein. Computing system 500 includes a bus 505 or other communication mechanism for communicating information and processor(s) 510 coupled to bus 505 for processing information. Processor(s) 510 may be any type of general or application-specific processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. Processor(s) 510 may also have multiple processing cores, at least some of which may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In particular embodiments, at least one processor(s) 510 may be a neuromorphic circuit that includes processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require the typical components of a von Neumann computing architecture.
[0073] The computing system 500 further includes memory 515 for storing information and instructions executed by the processor(s) 510. The memory 515 may be comprised of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage such as a magnetic or optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media may be any available media that can be accessed by the processor(s) 510 and may include volatile media, non-volatile media, or both. Also, the media may be removable, non-removable, or both.
[0074] Additionally, the computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via wireless and / or wired connections. In some embodiments, the communications device 520 may be configured to support a variety of wireless technologies, including Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Global System for Mobile (GSM) communications, General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced (LTE-A), 802.11n, ... It may be configured to use 802.11x, Wi-Fi, Zigbee, Ultra-Wideband (UWB), 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near-Field Communications (NFC), 5th Generation (5G), New Radio (NR), any combination thereof, and / or any other currently existing or future implemented communication standards and / or protocols without departing from the scope of the present invention.In some embodiments, the communication device 520 may include one or more antennas that are a single antenna, an array of antennas, a phased antenna, a switched antenna, a beamforming antenna, a beamsteering antenna, a combination thereof, and / or any other antenna configuration without departing from the scope of the present invention.
[0075] The processor(s) 510 are further coupled via bus 505 to a display 525, such as a plasma display, a liquid crystal display (LCD), a light-emitting diode (LED) display, a field emission display (FED), an organic light-emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high-definition display, a Retina® display, an in-plane switching (IPS) display, or any other suitable display for displaying information to a user. The display 525 may be configured as a touch (haptic) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-touch display, or the like, using resistive, capacitive, surface acoustic wave (SAW) capacitive, infrared, optical imaging, dispersive signaling, acoustic pulse recognition, frustrated total internal reflection, or the like. Any suitable display device and haptic I / O may be used without departing from the scope of the invention.
[0076] A keyboard 530 and cursor control device 535, such as a computer mouse, touchpad, etc., are further coupled to bus 505 to allow a user to interface with computing system 500. However, in certain embodiments, a physical keyboard and mouse may not be present, and the user may interact with the device solely through the display 525 and / or touchpad (not shown). Any type and combination of input devices may be used as a matter of design choice. In certain embodiments, no physical input devices and / or displays are present. For example, a user may interact with computing system 500 remotely via another computing system in communication with computing system 500, or computing system 500 may operate autonomously.
[0077] The memory 515 stores software modules that provide functionality when executed by the processor(s) 510. The modules include an operating system 540 for the computing system 500. The modules further include a combined serial-and-parallel unified target / default targeting method module 545 configured to perform all or a portion of the processes described herein, or derivatives thereof. The computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0078] Those skilled in the art will appreciate that a "system" may be embodied as a server, embedded computing system, personal computer, console, personal digital assistant (PDA), mobile phone, tablet computing device, quantum computing system, or any other suitable computing device or combination of devices without departing from the scope of the present invention. Presenting the above-described functions as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. Indeed, the methods, systems, and apparatuses disclosed herein may be implemented in localized and distributed forms consistent with computing technologies, including cloud computing systems. The computing system may be part of or otherwise accessible through a local area network (LAN), a mobile communications network, a satellite communications network, the Internet, a public or private cloud, a hybrid cloud, a server farm, any combination thereof, or the like. Any localized or distributed architecture may be used without departing from the scope of the present invention.
[0079] It should be noted that some of the system features described herein are presented as modules to further emphasize implementation independence. For example, a module may be implemented as a hardware circuit comprising custom very large scale integrated (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, etc.
[0080] Modules may also be implemented at least partially in software for execution by various types of processors. For example, an identified unit of executable code may include one or more physical or logical blocks of computer instructions, which may be organized, for example, as an object, procedure, or function. Nevertheless, executable identified modules need not be physically located together; they may include separate instructions stored in different locations that, when logically combined, comprise modules to achieve the purpose stated for the modules. Furthermore, modules may be stored on computer-readable media, such as, for example, a hard disk drive, a flash device, RAM, tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of the present invention.
[0081] Indeed, a module of executable code may be a single instruction, many instructions, and even distributed across several different code segments, different programs, and multiple memory devices. Similarly, operational data may be identified and depicted herein within modules, and may be embodied and organized in any suitable form within any suitable type of data structure. Operational data may be collected as a single data set, or may be distributed in different locations across different storage devices, or may exist, at least in part, simply as electronic signals on a system or network.
[0082] 6A-G show a unified target configuration interface for an RPA designer application 600, according to an embodiment of the present invention. In this embodiment, an RPA developer can custom configure unified target functionality for an activity in an RPA workflow. The RPA designer application 600 includes an RPA workflow development pane 610 with a click activity 612. The RPA designer application 600 also includes a unified target configuration pane 620. When a user clicks an activity that interacts with a graphical element in the UI, the unified target configuration pane 620 shows the unified target options for that activity.
[0083] As shown in FIG. 6A, a serial execution selector 630 allows an RPA developer to select whether the unified target techniques for an activity 612 are executed individually serially or in parallel with at least one stage of delayed parallel execution. If serial execution is selected, the execution order of the techniques may be specified using drop-down 632. See FIG. 6B. The selected technique(s) 634 are displayed after selection, and the timeout for each technique may be specified via timeout field 636. See FIG. 6C.
[0084] If delayed parallel execution is desired, the value of the serial execution selector 630 may be set to “No,” after which the RPA developer can select whether one or more initial techniques will be used via the initial technique selector 640. See FIG. 6D. Note that if the initial technique selector 640 is set to “No,” unified target graphical element detection techniques may be executed in parallel. However, if the initial technique selector 640 is set to “Yes,” a dropdown 642 for selecting an initial technique is displayed. See FIG. 6E. However, in some embodiments, the one or more initial techniques are executed automatically and cannot be configured by the RPA developer. For example, a selector-based approach may be faster than others and may be tried first. In certain embodiments, one or more default initial techniques may be presented, and the RPA developer may change them (e.g., by adding a new technique, removing a default technique, etc.). In some embodiments, the RPA developer cannot configure the initial technique(s), but can configure the delayed parallel technique(s).
[0085] Once the initial technique 644 is selected, a parallel technique dropdown 646 appears, listing the remaining techniques. See FIG. 6F. In some embodiments, the type of graphical element detection technique is automatically selected based on the action implemented by the activity (e.g., click, text capture, hover, etc.), the type of graphical element (e.g., button, text field, etc.), and / or the specific graphical element indicated by the RPA developer (e.g., which element the user selected on the screen and what other elements exist in the application). For specific graphical elements, for example, if the RPA developer clicks an OK button, but there are two OK buttons on the screen, some attributes may be automatically added to distinguish between the two identical OK buttons. For example, when using a UI tree, the UI tree is typically constructed so that when the RPA developer displays a graphical element on the screen, at least some of the attributes of the UI tree are different for that graphical element from other graphical elements.
[0086] The RPA developer may add more initial techniques via add link 645, or the RPA developer may remove a previously selected initial technique. In some embodiments, if the RPA developer selects “No” for initial technique selector 640, parallel technique dropdown 646 remains displayed, and the RPA developer may custom select which techniques to run in parallel. After parallel technique 647 is selected, the RPA developer may add more parallel techniques via add link 648, or the RPA developer may remove a previously selected parallel technique. See FIG. 6G. A delay for how long to wait for execution of parallel technique(s) after execution of initial technique(s) begins may be specified via delay field 649. In some embodiments where a multi-anchor technique is used to identify a target and one or more anchors, unified target settings may be custom configured for the target and each anchor, or the same settings may be applied to the target and anchor(s).
[0087] In some embodiments, multiple delay periods may be used. For example, an initial technique may be used for one second, and if no match is found, one or more other techniques may be used in parallel with the initial technique, and if no match is found during a second period, still other techniques may be applied in parallel, etc. Any number of delay periods and / or techniques within each delay period may be used without departing from the scope of the invention.
[0088] In some embodiments, mutually exclusive serial stages may be employed. For example, an initial technique may be used for one second, and if no match is found, one or more other techniques may be used in place of the initial technique, and if no match is found during a second period, yet other technique(s) may be applied in place of the initial technique and the second period technique(s), etc. In this way, techniques that appear unsuccessful may be stopped, potentially reducing resource requirements.
[0089] 7A illustrates a collapsed targeting method configuration interface 700 for configuring at the application and / or UI type level, according to an embodiment of the present invention. In this embodiment, a user may configure targeting methods for web browser, Java, SAP, Microsoft UI Automation (UIA) desktop, Active Accessibility (AA) desktop, and Microsoft Win32 desktop. Each of these applications and / or display types may be individually configured by the user.
[0090] Returning to Figure 7B, the user has configured the default targeting methods for SAP® using the SAP® Targeting Methods tab 710. Because SAP® now has powerful and reliable selectors, the Execution value of the Full Selector 712 is set to "True." The Execution values of the Fuzzy Selector 714, Image Selector 716, and Enable Anchor 718 (i.e., to enable target identification using the Target / Anchor function) are set to "False."
[0091] However, this technique may not work for all applications and / or UI types. For example, selectors may not work well with many modern web browsers because attribute values tend to change dynamically. Returning to FIG. 7C , the user configured the default targeting method for the web browser using the Web Targeting Method tab 720. Input Mode 722 is set to “Simulate” via drop-down menu 723. This is a web-specific setting; other application- and / or UI-type-specific settings for web applications and / or other application / UI types may be included without departing from the scope of the present invention. In some embodiments, if the application is automated, a different mechanism for interacting with the application (e.g., providing mouse clicks, key presses, etc.) may be used. “Simulate” simulates the input the web browser would receive from the system if a user were to perform a similar interaction.
[0092] Because using the full selector is not accurate in many web browsers, the execution value of the full selector setting 724 is set to "false," while the execution values of the fuzzy selector 725, image selector 726, and enable anchor 727 are set to "true." This is the opposite configuration from the SAP® targeting methods tab 710 in Figure 7B.
[0093] Configuring the targeting method per application and / or per UI type avoids the user having to reconfigure each graphical element as he or she points to the element on the screen. For example, if a user points to a text field in a web browser, the targeting method for that text field will be as preconfigured by the user, without the user having to go in and set the execution value of a full selector to "false," a fuzzy selector to "true," etc. However, it should be noted that in some embodiments, if there are particular UI elements that may be more accurately detected using a different targeting method configuration than the default configuration, the user can change these default values.
[0094] It should be noted that other targeting methods are possible. For example, in some embodiments, CV, OCR, or a combination thereof may be used. Indeed, any suitable graphical element detection technique may be used without departing from the scope of the present invention.
[0095] FIG. 8 is a flowchart illustrating a process 800 for configuring a unified target function for an activity in an RPA workflow according to an embodiment of the present invention. In some embodiments, process 800 may be performed by the RPA designer application 600 of FIGS. 6A-6G. The process begins with receiving a selection of an activity in the RPA workflow 810 configured to perform graphical element detection using the unified target. In some embodiments, the unified target function is automatically pre-configured at 820. In particular embodiments, this pre-configuration may be based on the type of graphical element detection technique, the type of graphical element, and / or the particular graphical element indicated by the RPA developer.
[0096] The RPA designer application may receive 830 changes to the unified target function from the RPA developer to custom configure the unified target function for the activity. The RPA designer application then configures the activity based on the unified target configuration at 840. If more activities are to be configured, the RPA developer may select another activity, and the process returns to step 810. Once the desired activity(ies) are configured, the RPA designer application generates 850 an RPA robot for implementing the RPA workflow that includes the configured activity(ies). The process then ends and proceeds to FIG. 9A.
[0097] 9A and 9B are flowcharts illustrating a process 900 for graphical element detection using a combined serial and delayed parallel execution unified target technique, according to an embodiment of the present invention. In some embodiments, the process 900 may be implemented at runtime by an RPA robot created via the RPA designer application 600 of FIGS. 6A-G. The process begins by analyzing a UI (e.g., screenshots, images of application windows, etc.) to identify UI element attributes at 910. The UI element attributes may include, but are not limited to, images, text, relationships between graphical elements, hierarchical representations of graphical elements within the UI, etc. The identification may be performed via CV, OCR, API calls, analysis of text files (e.g., HTML, XML, etc.), combinations thereof, etc.
[0098] After the UI is analyzed, UI element attributes are analyzed for activity using the unified target at 920. Returning to FIG. 9B, one or more initial graphical element detection techniques are run (potentially in parallel) at 922. If a match is found for the initial technique(s) at 924 within a first period (e.g., within a tenth of a second, within a second, within 10 seconds, etc.), a result from this match is selected at 926.
[0099] If no match is found in the initial technique period at 924, one or more additional graphical element detection techniques are executed in parallel at 928. If a match is found in the second period at 929, then the results of the first technique for finding a match from all initial and subsequent parallel techniques are selected at 926, and the process proceeds to step 930. In some embodiments, finding a match for a graphical element may include finding a match between the graphical element itself as the target and its anchor(s), and a first technique for finding a match for each may be selected for that respective target / anchor. If no match is found in the second period at 929, the process also proceeds to step 930. In some embodiments, multiple stages of delayed parallel execution are executed. In certain embodiments, a different technique is executed at each stage, and the previous technique is stopped.
[0100] If a matching UI element is found via the unified target at 930, an action associated with the activity containing the UI element is performed at 940 (e.g., clicking a button, entering text, interacting with a menu, etc.), and if there are more activities at 950, the process proceeds to step 920 for the next activity. However, if no UI element is found at 930 that matches the attributes of the graphical element detection technique, an exception is thrown or the user is asked at 960 how he or she wants to proceed (e.g., whether to continue execution) and the process ends.
[0101] 10A and 10B are flowcharts illustrating design-time and run-time portions, respectively, of a process 1000 for configuring graphical element detection techniques and performing graphical element detection at the application and / or UI type level, according to an embodiment of the present invention. In some embodiments, the design-time portion of process 1000 may be performed by the targeting method configuration interface 700 of FIGS. 7A-C. The process begins by receiving at 1005 a selection of an application or UI type for a default targeting method configuration. Next, the developer sets a default targeting method setting (e.g., web browser, Win32 desktop, etc.), and this default configuration is received and saved at 1010. The default configuration may be saved in a UI object repository in some embodiments that can be accessed by multiple or multiple users. The user may then select a different application or UI type for the default targeting method configuration, if desired.
[0102] The same developer or another developer then displays the screen in the UI, and the screen display is received at 1015 (e.g., by an RPA designer application such as UiPath Studio™). In other words, different instances of the RPA designer application can be used to configure and subsequently change the default targeting method settings. Indeed, these instances may not, in some embodiments, be on the same computing system. The default targeting method settings are then pre-configured at 1020 for the UI type associated with the detected application or screen. The RPA designer application may receive the change in the default targeting method settings at 1025 and configure the settings accordingly. For example, a user may select the targeting method they want to use for a screen for which the default technique is not performing as well as desired. For example, perhaps a particular screen has visual characteristics that are substantially different from other screens within the application or desktop.
[0103] After the targeting method settings are set by default or changed from the default configuration, the user deploys an RPA workflow in which activities are configured using these targeting method settings at 1030. In some embodiments, the user may change the default targeting method settings while still designing the RPA workflow. After the user completes the RPA workflow, the RPA designer application generates an RPA robot for implementing the RPA workflow including the configured activity(ies) at 1035. The process then ends and proceeds to FIG. 10B.
[0104] Returning to Figure 10B, the UI (e.g., screenshots, images of application windows, etc.) is analyzed to identify UI element attributes at 1040. UI element attributes may include, but are not limited to, images, text, relationships between graphical elements, hierarchical representations of graphical elements within the UI, etc. Identification may be performed via CV, OCR, API calls, analysis of text files (e.g., HTML, XML, etc.), combinations thereof, etc.
[0105] After the UI is parsed, the UI element attributes are analyzed for the activity using the default targeting method configuration(s) at 1045, unless overridden for a given UI element by the user during design-time development. If a matching UI element is found using the targeting method settings at 1050, an action associated with the activity containing the UI element is performed at 1055 (e.g., clicking a button, entering text, interacting with a menu, etc.). Also, if there are more activities at 1060, the process proceeds to step 1045 for the next activity. However, if the UI element is not found using the targeting method(s) set at 1050, either an exception is thrown or the user is asked at 1065 how he or she wants to proceed (e.g., whether to continue execution) and the process ends.
[0106] The process steps performed in FIGS. 8-10B may be performed by a computer program encoding instructions to a processor(s) to perform at least a portion of the process(es) described in FIGS. 8-10B in accordance with an embodiment of the present invention. The computer program may be stored on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, tape, and / or any other such medium or combination of media used to store data. The computer program may include coded instructions for controlling a processor(s) of a computing system (e.g., processor(s) 510 of computing system 500 of FIG. 5) to implement all or a portion of the process steps described in FIGS. 8-10B, which may also be stored on a computer-readable medium.
[0107] The computer program may be implemented in hardware, software, or a hybrid implementation. The computer program may be composed of modules in operable communication with each other and designed to send information or instructions to a display. The computer program may be configured to run on a general-purpose computer, an ASIC, or any other suitable device.
[0108] It will be readily understood that the components of the various embodiments of the present invention, as generally described and illustrated herein, may be arranged and designed in a wide variety of different configurations. Thus, the detailed description of the embodiments of the present invention, as represented in the accompanying figures, is not intended to limit the scope of the invention as claimed, but is merely representative of selected embodiments of the invention.
[0109] The features, structures, or characteristics of the invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments," "some embodiments," or similar language mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, the appearances of "certain embodiments," "some embodiments," "other embodiments," or similar language throughout this specification do not necessarily refer to the same group of all embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0110] It should be noted that references to features, advantages, or similar language throughout this specification do not imply that all of the features and advantages that may be realized in the present invention are to be present in any single embodiment of the present invention, or in any embodiment of the present invention. Rather, language referring to features and advantages is understood to mean that the particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the present invention. Thus, discussions of features and advantages throughout this specification, and similar language, may, but need not, refer to the same embodiment.
[0111] Furthermore, the described features, advantages, and characteristics of the invention may be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the invention may be practiced without certain features or advantages of one or more particular embodiments. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the invention.
[0112] Those of ordinary skill in the art will readily appreciate that the invention as described above can be implemented using steps in a different order and / or with hardware elements in different configurations than those disclosed. Thus, while the invention has been described in terms of these preferred embodiments, it will be apparent to those skilled in the art that certain modifications, variations, and alternative configurations will become apparent while remaining within the spirit and scope of the invention. Accordingly, reference should be made to the appended claims to determine the scope of the invention.
Claims
1. A non-transitory computer-readable medium having stored thereon a computer program, the computer program causing at least one processor to: Analyzing a user interface (UI) at run time to identify UI element attributes; comparing the UI element attributes with UI descriptor attributes of activities in a robotic process automation (RPA) workflow using one or more early graphical element detection techniques; If no match is found using the one or more initial graphical element detection techniques during the first time period, A non-transitory computer-readable medium configured to execute one or more additional graphical element detection techniques in parallel with the one or more initial graphical element detection techniques.
2. If during a second time period no match is found using the one or more initial graphical element detection techniques and the one or more additional graphical element detection techniques, the computer process further comprises causing the at least one processor to:
10. The non-transitory computer-readable medium of claim 1, configured to perform one or more supplemental graphical element detection techniques in parallel with the one or more initial graphical element detection techniques and the one or more additional graphical element detection techniques.
3. If a match is found, the computer program further comprises causing the at least one processor to: The non-transitory computer-readable medium of claim 1 configured to take an action associated with the activity that includes the UI element.
4. 10. The non-transitory computer-readable medium of claim 1, wherein the process of claim 1 is repeated for at least one additional activity.
5. The non-transitory computer-readable medium of claim 1 , wherein the UI descriptor attributes include two or more of a selector attribute, a computer vision (CV) attribute, an image matching attribute, and an optical character recognition (OCR) attribute.
6. A non-transitory computer-readable medium having stored thereon a computer program, the computer program causing at least one processor to: Analyzing a user interface (UI) at run time to identify UI element attributes; comparing the UI element attributes with UI descriptor attributes of activities in a robotic process automation (RPA) workflow using one or more early graphical element detection techniques; If no match is found using the one or more initial graphical element detection techniques during the first time period, A non-transitory computer-readable medium configured to perform one or more additional graphical element detection techniques in place of the one or more initial graphical element detection techniques.
7. If during a second time period no match is found using the one or more additional graphical element detection techniques, the computer program further causes the at least one processor to:
7. The non-transitory computer-readable medium of claim 6, configured to perform one or more supplemental graphical element detection techniques in place of the one or more initial graphical element detection techniques and the one or more additional graphical element detection techniques.
8. If no match is found, the computer program further comprises causing the at least one processor to: The non-transitory computer-readable medium of claim 6 , configured to take an action associated with the activity that includes the UI element.
9. 7. The non-transitory computer-readable medium of claim 6, wherein the process of claim 6 is repeated for at least one additional activity.
Citation Information
Patent Citations
Robotic process automation
JP2018535459A
Systems and methods for capture and generation of process workflow
US20200050983A1