Sequence extraction using screenshot images
By using screenshot images for sequence extraction, a robot process automation (RPA) workflow is generated, which solves the differences and noise problems of UI element information platform in the prior art, and realizes efficient candidate process automation and cost reduction.
Patent Information
- Application Number
- CN202080044412.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-02
- Filing Date
- 2020-09-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-09-30
AI Technical Summary
When using user interface (UI) element information to identify and automate candidate processes, the prior art faces problems of platform differences and noise information, resulting in low efficiency and high cost of automation processes.
By using screenshot images for sequence extraction, a robot process automation (RPA) workflow is generated. The specific steps include capturing screenshots of user operation steps, storing screenshots, random clustering actions, extracting sequences and discarding subsequent events, and finally generating an automated workflow.
Automatic identification and extraction of repetitive tasks is realized, which improves the automation efficiency of candidate processes, reduces professional service fees and improves ROI.
Smart Images

Figure CN114008609B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Application No. 16 / 591,161, filed on October 2, 2019, the contents of which are incorporated herein by reference. Background Art
[0003] In order to identify candidate processes and extract action sequences, existing techniques utilize generic information about user actions (such as user clicks or keystrokes) in combination with information about user interface (UI) elements. The problem with information collected from UI elements is that it may vary from platform to platform and may contain noise because UI elements depend on application-level configuration.
[0004] Therefore, enterprises that seek to leverage Robotic Process Automation (RPA) to automate their processes struggle in identifying candidate processes that can be automated and ultimately end up with high professional service fees and / or low ROI. Summary of the invention
[0005] Disclosed are systems and methods for sequence extraction using screenshot images to generate robotic process automation workflows. The system and method relate to automatically identifying candidate tasks for robotic process automation (RPA) on desktop applications, and more specifically, to sequence extraction for identifying repetitive tasks from screenshots of user actions. The system and method include: using a processor to capture multiple screenshots of steps performed by a user on an application; storing the screenshots in a memory; determining action clusters from the captured screenshots by randomly clustering the actions into any predefined number of clusters, wherein screenshots of different variations of the same action are marked in the clusters; extracting sequences from the clusters, and discarding subsequent events on the screen from the clusters; and generating an automated workflow based on the extracted sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] A more detailed understanding may be obtained from the following description given by way of example in conjunction with the accompanying drawings, in which like reference numerals represent like elements, and in which:
[0007] Figure 1A It is a diagram of the development, design, operation or execution of Robotic Process Automation (RPA);
[0008] Figure 1B It is another illustration of RPA development, design, operation or execution;
[0009] Figure 1C is a diagram of a computing system or environment;
[0010] Figure 2 A description of the candidate identification is shown;
[0011] Figure 3 shows a set of screenshots clustered to define a template;
[0012] Figure 4 a diagram showing the flow of a sequence of actions at the screen level; and
[0013] Figure 5 A method for sequence extraction using screenshot images to generate a robotic process automation workflow is shown. DETAILED DESCRIPTION
[0014] For the methods and processes described below, the recorded steps may be performed in any order out of sequence, and sub-steps that are not explicitly described or shown may be performed. In addition, "coupled" or "operably coupled" may indicate that objects are linked but may have zero or more intermediate objects between linked objects. Moreover, any combination of the disclosed features / elements may be used in one or more embodiments. When reference to "A or B" is used, it may include A, B, or A and B, which may similarly be expanded into a longer list. When the symbol X / Y is used, it may include X or Y. Alternatively, when the symbol X / Y is used, it may include X and Y. The X / Y symbol may similarly be expanded into a longer list with the same interpretation logic.
[0015] The system and method relate to automatically identifying candidate tasks for robotic process automation (RPA) on desktop applications, and more specifically, to sequence extraction for identifying repetitive tasks from screenshots of user actions. A system and method for sequence extraction using screenshot images to generate a robotic process automation workflow is disclosed. The system and method include: using a processor to capture multiple screenshots of steps performed by a user on an application; storing the screenshots in a memory; determining action clusters from the captured screenshots by randomly clustering the actions into any predefined number of clusters, wherein screenshots of different variations of the same action are marked in the clusters; extracting sequences from the clusters, and discarding subsequent events on the screen from the clusters; and generating an automated workflow based on the extracted sequences.
[0016] Figure 1Ais an illustration of RPA development, design, operation, or execution 100. A designer 102 (sometimes referred to as a studio, development platform, development environment, etc.) can be configured to generate code, instructions, commands, etc. for a robot to execute or automate one or more workflows. Based on the (multiple) selections that a computing system can provide to the robot, the robot can determine representative data for the (multiple) areas of the visual display selected by the user or operator. As part of RPA, multi-dimensional shapes such as squares, rectangles, circles, polygons, free shapes, etc. can be used for UI robot development and runtime associated with computer vision (CV) operations or machine learning (ML) models.
[0017] Non-limiting examples of actions that can be completed by a workflow can be one or more of the following: performing a login, filling out a form, information technology (IT) management, etc. In order to run a UI automated workflow, a robot may need to uniquely identify specific screen elements such as buttons, checkboxes, text fields, labels, etc., regardless of application access or application development. Examples of application access can be local, virtual, remote, cloud, Remote desktop, virtual desktop infrastructure (VDI), etc. Examples of application development may be win32, Java, Flash, Hypertext Markup Language (HTML), HTML5, Extensible Markup Language (XML), JavaScript, C#, C++, Silverlight, etc.
[0018] Workflows may include, but are not limited to, task sequences, flow charts, finite state machines (FSMs), global exception handlers, and the like. A task sequence may be a linear process for processing linear tasks between one or more applications or windows. A flow chart may be configured to process complex business logic so that the integration of decisions and the connection of activities can be realized in a more diverse manner through multiple branching logic operators. FSMs may be configured for large workflows. FSMs may use a limited number of states in their execution, which may be triggered by conditions, transitions, activities, and the like. A global exception handler may be configured to determine workflow behavior when an execution error is encountered, and may be configured for, debugging processes, and the like.
[0019] A robot may be an application, applet, script, etc. that can automate a UI that is transparent to the underlying operating system (OS) or hardware. When deployed, one or more robots may be managed, controlled, etc. by a commander 104 (sometimes referred to as a coordinator). The commander 104 may instruct or command (multiple) robots or automation executors 106 to execute or monitor workflows in clients, applications, or programs such as mainframes, web, virtual machines, remote machines, virtual desktops, enterprise platforms, (multiple) desktop applications, browsers, etc. The commander 104 may serve as a central point or semi-central point to instruct or command multiple robots to automate a computing platform.
[0020] In certain configurations, the commander 104 may be configured to set up, deploy, configure, queue, monitor, log, and / or provide interconnectivity. Setup may include creating and maintaining connections or communications between (multiple) robots or automation executors 106 and the commander 104. Deployment may include ensuring that grouped versions are delivered to assigned robots for execution. Configuration may include maintenance and delivery of robot environment and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include tracking robot identification data and maintaining user permissions. Logging may include storing and indexing logs into a database (e.g., a SQL database) and / or other storage mechanisms (e.g., SQL Server 2003) that provide the ability to store and quickly query large data sets. ). The director 104 may provide interconnectivity by serving as a centralized communication point for third-party solutions and / or applications.
[0021] The robot or automation actuator(s) 106 may be configured to be unattended 108 or attended 110. For unattended 108 operation, the automation may be performed without the aid of third-party input or control. For attended 110 operation, the automation may be performed by receiving input, commands, instructions, guidance, etc. from a third-party component.
[0022] Robot(s) or automation executors 106 may be execution agents that run the workflows built in the designer 102. Commercial examples of robots(s) for UI or software automation are UiPath Robots TM In some embodiments, the robot(s) or automation executor 106 may be installed with Microsoft Services managed by the Service Control Manager (SCM). Therefore, such a robot can open interactive session, and has Permissions for services.
[0023] In some embodiments, (multiple) robots or automation implementers 106 can be installed in user mode. These robots can have the same permissions as the user who installed the given robot. This feature can also be used for high density (HD) robots, which ensures that each machine is fully utilized at the highest performance, such as in an HD environment.
[0024] In some configurations, the robot(s) or automation executor 106 may be split, distributed, etc. into several components, each dedicated to a specific automation task or activity. Robot components may include SCM-managed robot services, user-mode robot services, executors, agents, command lines, etc. SCM-managed robot services may manage or monitor The session acts as a proxy between the director 104 and the execution host (i.e., the computing system on which the (multiple) robots or automation executors 106 execute). These services can be trusted and manage the credentials of the (multiple) robots or automation executors 106.
[0025] User-mode robot services can be managed and monitored The user mode robot service can be trusted and manage the credentials of the robot 130. If the SCM-managed robot service is not installed, then Applications can be started automatically.
[0026] The actuator can be The executor can be aware of the dots per inch (DPI) setting of each monitor. The agent can be the one that displays available jobs in a system tray window. Presentation Foundation (WPF) applications. Agents can be clients of services. Agents can request to start or stop jobs and change settings. Command lines can be clients of services. Command lines are console applications that can request to start jobs and wait for their output.
[0027] In a configuration where the components of (multiple) robots or automated executors 106 are split as explained above, developers, support users, and computing systems are helped to more easily run, identify, and track the execution of each component. Special behaviors can be configured for each component in this way, such as setting different firewall rules for executors and services. In some embodiments, the executor can be aware of the DPI setting of each monitor. Therefore, workflows can be executed at any DPI, regardless of the configuration of the computing system on which they are created. Projects from the designer 102 can also be independent of the browser zoom level. In some embodiments, DPI can be disabled for applications that are unaware of DPI or intentionally marked as unaware.
[0028] Figure 1B is another illustration of RPA development, design, operation, or execution 120. A studio component or module 122 may be configured to generate code, instructions, commands, and the like for a robot to perform one or more activities 124. User interface (UI) automation 126 may be performed by a robot on a client using one or more driver components 128. A robot may use a computer vision (CV) activity module or engine 130 to perform activities. Other drivers 132 may be used for UI automation performed by a robot to obtain UI elements. They may include OS drivers, browser drivers, virtual machine drivers, enterprise drivers, and the like. In some configurations, a CV activity module or engine 130 may be a driver for UI automation.
[0029] Figure 1C 1 is a diagram of a computing system or environment 140, which may include a bus 142 or other communication mechanism for transmitting information or data, and one or more processors 144 coupled to the bus 142 for processing. The one or more processors 144 may be any type of general or special purpose processor, including a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a controller, a multi-core processing unit, a three-dimensional processor, a quantum computing device, or any combination thereof. The one or more processors 144 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. Multiple parallel processing may also be configured. In addition, at least one or more processors 144 may be a neuromorphic circuit including a processing element that simulates a biological neuron.
[0030] The memory 146 may be configured to store information, instructions, commands, or data to be executed or processed by the processor(s) 144. The memory 146 may include any combination of random access memory (RAM), read-only memory (ROM), flash memory, solid-state memory, cache, static storage (such as a magnetic disk or optical disk), or any other type of non-transitory computer-readable medium or combination thereof. Non-transitory computer-readable media may be any media that can be accessed by the processor(s) 144, and may include volatile media, non-volatile media, etc. The media may also be removable, non-removable, etc.
[0031] The communication device 148 may be configured as frequency division multiple access (FDMA), single carrier FDMA (SC-FDMA), time division multiple access (TDMA), code division multiple access (CDMA), orthogonal frequency division multiplexing (OFDM), orthogonal frequency division multiple access (OFDMA), global system for mobile communications (GSM), general packet radio service (GPRS), universal mobile telecommunications system (UMTS), cdma2000, wideband CDMA (W-CDMA), high speed downlink packet access (HSDPA), high speed uplink packet access (HSUP), etc. A), High Speed Packet Access (HSPA), Long Term Evolution (LTE), Advanced LTE (LTE-A), 802.11x, Wi-Fi, Zigbee, Ultra-Wideband (UWB), 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near Field Communication (NFC), Fifth Generation (5G), New Radio (NR), or any other wireless or wired device / transceiver for communicating via one or more antennas. The antennas may be single, arrayed, phased, switched, beamformed, beamsteering, etc.
[0032] The one or more processors 144 may also be coupled to a display device 150 via the bus 142, such as a plasma, a liquid crystal display (LCD), a light emitting diode (LED), a field emission display (FED), an organic light emitting diode (OLED), a flexible OLED, a flexible substrate display, a projection display, a 4K display, a high definition (HD) display, Display, in-plane switching (IPS) based display, etc. Display device 150 can be configured as a touch, three-dimensional (3D) touch, multi-input touch or multi-touch display using resistance, capacitance, surface acoustic wave (SAW) capacitance, infrared, optical imaging, dispersive signal technology, acoustic pulse recognition, frustrated total internal reflection or other technology for input / output (I / O) understood by those of ordinary skill in the art.
[0033] A keyboard 152 and control device 154 (such as a computer mouse, touch pad, etc.) may also be coupled to bus 142 for input to computing system or environment 140. In addition, input may be provided to computing system or environment 140 remotely via another computing system in communication with computing system or environment 140, or computing system or environment 140 may operate autonomously.
[0034] The memory 146 may store software components, modules, engines, etc. that provide functionality when executed or processed by one or more processors 144. This may include an OS 156 for the computing system or environment 140. Modules may also include custom modules 158 to perform application specific processes or derivatives thereof. The computing system or environment 140 may include one or more additional function modules 160 with additional functionality.
[0035] The computing system or environment 140 can be adapted or configured to perform as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing device, a cloud computing device, a mobile device, a fixed mobile device, a smart display, a wearable computer, etc.
[0036] In the examples given herein, a module may be implemented as a hardware circuit comprising a custom very large scale integrated (VLSI) circuit or gate array, an off-the-shelf semiconductor such as a logic chip, a transistor or other discrete component. A module may also be implemented in a programmable hardware device such as a field programmable gate array, programmable array logic, a programmable logic device, a graphics processing unit, etc.
[0037] A module may be implemented at least in part in software for execution by various types of processors. The identified executable code unit may include one or more physical or logical blocks of computer instructions, which may be organized, for example, as objects, procedures, routines, subroutines, or functions. Executable files of the identified modules that are co-located or stored in different locations include the module when logically linked together.
[0038] An executable code module may be a single instruction, one or more data structures, one or more data sets, multiple instructions, etc., distributed over several different code segments, between different programs, across several memory devices, etc. Operational or functional data may be identified and described herein within a module, and may be embodied in a suitable form and organized within any suitable type of data structure.
[0039] In the examples given herein, the computer program may be configured in hardware, software or a hybrid implementation. The computer program may be composed of modules that communicate effectively with each other and transfer information or instructions.
[0040] In included embodiments, screenshots of user actions are used to extract sequences of repeated actions. Sequence extraction can be performed using action clustering. Action clustering is configured to mark screenshots associated with different variations of the same action. Actions can be clustered unsupervised using a screenshot-based approach.
[0041] The disclosed embodiments relate to automatically identifying candidate tasks for RPA on desktop applications. Sequence extraction applied to screenshots of user actions can be used to extract sequences of repeated actions to identify candidate tasks. Sequence extraction can include the steps of randomly clustering actions into a predefined number of clusters, defining a template for each cluster, aggregating the features used in the template into a sparse feature space in which samples are again clustered, and introducing a sequence extraction method after unifying the samples into cluster labels. A template can be defined as a layout of a screen that is unique to that screen, although it is understood that similar screens follow the layout.
[0042] Figure 2 A description of a candidate identification 200 is shown. The candidate identification 200 identifies candidate processes that can be automated and end up with high professional service fees and / or low ROI. The candidate identification 200 includes: clustering actions 210, such as user actions and UI elements; extracting sequences 220 from the clustered actions 210; and understanding processes 230 based on the extracted sequences 220 to identify candidate processes for automation while minimizing professional fees and improving ROI. The candidate identification 200 includes action clustering or clustering actions 210, in which multiple screenshots are clustered for defining a common template. The clustering actions 210 can include templates, adaptive parameter tuning, random sampling, clustering details, and novelty, as will be described in more detail below. The candidate identification 200 includes sequence extraction or extract sequence 220, which identifies the execution sequence of the task from the cluster. The extract sequence 220 can include forward link estimation, graph representation, and action clustering, as will be described in more detail below. The candidate identification 200 includes an understanding process 230, such as a candidate process identification for RPA.
[0043] Clustering action 210 utilizes optical character recognition (OCR) data extracted from the screenshots. In an exemplary embodiment, an OCR engine is used to extract word and position data pairs from the screenshots. Using the word set and corresponding (normalized) coordinates on the screenshots, an adaptive particle-based method is implemented that iteratively extracts a sparse feature set for clustering. Clustering action 210 can be randomly clustered into any predefined number of clusters (number of clusters > 0).
[0044] The clustering action 210 iteratively utilizes a center-based clustering paradigm. For each cluster, a center is defined. In this context, the center is called a template. A template is defined as a layout of a screen that is unique to that screen, although it is understood that similar screens follow this layout. Using this assumption, the clustering action 210 determines a template for each cluster. The aggregation of features used in the template is then used as a sparse feature space in which samples are clustered again, as shown in Equation 1, given a set of N screenshots S:
[0045] S = {S1, S2, ..., SN}.
[0046] Equation 1.
[0047] In each screenshot i In the OCR engine, the image is searched for m with a corresponding position. i For simplicity, each position is normalized to the screen resolution and converted to (area, center x , center y ) format. In equation 2, in the screenshot s i The jth word seen in ij , and its corresponding position is l ij .
[0048]
[0049] Assume cluster π: S→C, where C = {c1, c2, ..., c K} is a set of K cluster labels, if π(S i )=cK, then the screenshot s i In cluster c k Templates can be created based on frequent words and positions in clusters. A list of frequently occurring words (W) can be calculated for each cluster using a frequency threshold method. A list of frequently occurring positions (L) can be calculated for each cluster based on a frequency threshold. In this frequency measure, two positions are similar if the intersection area covers more than 90% of the joint area. As will be appreciated, W and L can be calculated separately.
[0050] Using W and L, we calculate the number of times each word or position (or both) appears in the sample cluster by generating a frequency matrix F. Consider the case of infrequent words or positions, where the elements is added to W and L. The frequency matrix has an additional row and column (F |W|+1,|L|+1 ,). In this representation, F i,j shows the number of times the i-th word in W has appeared at the j-th position in L, which was generated by looking at a screenshot of the cluster. In addition, F (|w|,j) Indicates the number of times a non-frequent word appears in the jth frequent position. When various data of the data input position of the screenshot appear in the same position, the non-frequent word may appear in the jth frequent position.
[0051] To construct the template, a set of frequently occurring words and positions is selected (with a frequency greater than 70% of the maximum observed frequency in each column, excluding the last row and last column). For the last row and last column, a threshold of 70% of their maximum values is used, respectively. As will be appreciated, other thresholds may also be used, including, for example, 60%, 80%, and thresholds found incrementally between 60-80%. It is conceivable to use any threshold between 0 and 100%, although only thresholds above about 50% are most useful in this application.
[0052] Templates consist of a combination of words and locations, describing the static parts of the page, and locations with various data, which are placeholders and frequent words that appear in various locations. Figure 3 A set of screenshots 300 are shown clustered to define a template.
[0053] Adaptive parameter tuning may be employed during iterations for clustering action 210. The above template may be used to evaluate clustered samples and tune clustering parameters for future iterations. k The corresponding template t k To evaluate the clusters, we measure the percentage of template elements to non-template elements in the cluster based on Equation 3:
[0054]
[0055] In this score, is the number of infrequent words and positions. This score gives an estimate of how similar the screenshot content is to the screenshot content in the current cluster.
[0056] Screenshots from different applications lead to different template scores in the ideal clustering. This indicates that the screenshots differ in the required clustering granularity. Therefore, the variance score var(score(t k ))(in ) is used to trigger a change in the number of clusters, which can increase or decrease based on the average of the template scores.
[0057] Random sampling can be used in clustering action 210 to ensure robust clusters and for scalability proposals. A resampling method similar to traditional particle swarm optimization is used. That is, clustering is performed on small random samples of the data set, and in each iteration, weighted sampling can select R% of the previous sample, and (1-R)% is randomly sampled from the main data set. To encourage diverse samples, each time a sample is drawn from the data set, its weight can be reduced by half, or by some other amount, so as to reduce duplicate samples and increase the diversity of the samples.
[0058] In each iteration of clustering action 210, templates are extracted for each cluster. Each sample is then represented as a binary feature vector indicating the presence of any template item. The feature vectors are then clustered using a mini-batch k-means method. At the end of the iteration or at the end of a given iteration, the final set of templates is used to generate a sparse representation of the screenshots for clustering.
[0059] The clustering action 210 is performed on the screenshots of each application separately, and the data is clustered based on various granularities for each application. Sequence extraction 220 can rely on appropriate clustering of the semantic action 210. The particle-based clustering method learns a sparse representation of the screen and tunes the clustering granularity as needed, such as by processing a small subset of the entire data set to generate OCR-based features.
[0060] After the samples are unified into cluster labels produced by the clustering action 210, sequence extraction 220 can be performed. Initially, the data set is cleaned by discarding subsequent events on the same screen. Discarding helps to focus on screen-level sequence extraction 220. The discarded data can be used in subsequent detailed sequence extraction 220.
[0061] In order to extract sequence 220, the relationship between each subsequent event can be evaluated by utilizing a forward link prediction module. The forward link prediction module can consider each event individually and predict in advance to determine which future events are linked to the event under consideration. When doing so, each event can be represented as its screenshot, the OCR words and positions collected in the action cluster 210, and the screenshot cluster label. The forward link prediction module applies the link prediction method to each event in the event corresponding to the screenshot s. To this end, the list of events that will occur in the next t seconds is collected as E(s), as defined in equation 4:
[0062] E(s)={s′, time(s′)-time(s) <t}
[0063] Equation 4.
[0064] To estimate whether any event in E(s) is linked to e, use Equation 5:
[0065]
[0066] Where g(T) is a Gaussian function with zero mean, T = time(s') - time(s). The denominator in the equation is based on the co-occurrence of words that are not frequent parts of the corresponding screen, which may be an indication that the two screens are linked. Using Equation 5 provides the probability measured on a sample of clusters at normalized frequency.
[0067] A graph of action sequences can be generated. Figure 4The resulting graph 400 shown in illustrates the flow of action sequences at the screen level. By tracking the samples that contribute to the weights of the edges, the graph 400 can be used to explain the repetitive tasks performed by the user. In the graph, each node 410 corresponds to a screen type found in the action cluster 210. The edges 420 of the graph point from each event to its linked event identified in the action cluster. Each edge 420 is weighted by the corresponding p(s, s′) value.
[0068] Figure 5 A method 500 for sequence extraction using screenshot images to generate a robotic process automation workflow is shown. The method 500 includes: at step 510, using a processor to capture multiple screenshots of steps performed by a user on an application. The method 500 includes: at step 520, storing the screenshots in a memory. The method 500 includes: at step 530, determining action clusters from the captured screenshots by randomly clustering the actions into any predefined number of clusters. Screenshots of different variations of the same action can be marked in the clusters. The method 500 includes: at step 540, extracting sequences from the clusters and discarding subsequent events on the screen from the clusters. The method 500 includes: at step 550, generating an automated workflow based on the extracted sequences.
[0069] The present embodiment saves time, maintains control over shared content, and provides increased efficiency by automatically identifying repeatedly completed user tasks.
[0070] Although features and elements are described above in specific combinations, it will be understood by those of ordinary skill in the art that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented in a computer program, software, or firmware that is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as, internal hard disks and removable memory disks), magneto-optical media and optical media (such as, CD-ROM disks), and digital versatile disks (DVDs).
Claims
1. A method for using screenshot images for sequence extraction to generate a robotic process automation (RPA) workflow, the method comprising: capturing, using a processor, a plurality of screenshots of steps performed by a user on an application, wherein the capturing includes templating to find a plurality of words and a position corresponding to each of the plurality of words to form a template; storing the plurality of captured screenshots in a memory; determining action clusters from the plurality of captured screenshots by randomly clustering actions into any predefined number of clusters based on the template applied to the plurality of captured screenshots, wherein screenshots of different variations of the same action are labeled in the clusters; extracting a sequence from the cluster and discarding subsequent events on the screen from the cluster; as well as The RPA workflow is generated based on the extracted sequence.
2. The method of claim 1, wherein the capturing further comprises at least one of the following: adaptive parameter tuning to iterate the template and tune the capture for subsequent iterations; Random sampling using particle swarm optimization; clustering minutiae comprising binary feature vectors indicating the presence of template items; and by learning a sparse representation of the screen and tuning the novelty of the clustering granularity.
3. The method of claim 2, wherein the templating utilizes a threshold to indicate the plurality of words.
4. The method according to claim 3, wherein the threshold comprises: 70%。 5. The method of claim 1, wherein the extracting comprises at least one of: forward link estimation, using a forward link prediction module to consider each event and link future events to each event; and A graphical representation where each graph node corresponds to a screen type found during said clustering. The method according to claim 5 , wherein the edges of the graph represent each event and the events linked to the event.
7. The method of claim 1, wherein the clustering utilizes optical character recognition (OCR) data to extract word and position pairs.
8. A system for using screenshot images for sequence extraction to generate a robotic process automation (RPA) workflow, the system comprising: a processor configured to capture a plurality of screenshots of steps performed by a user on an application, wherein the capturing includes templating to find a plurality of words and a position corresponding to each of the plurality of words to form a template; as well as a memory module operably coupled to the processor and configured to store the plurality of captured screenshots, wherein the processor is further configured to: determining action clusters from the plurality of captured screenshots by randomly clustering actions into any predefined number of clusters based on the template applied to the plurality of captured screenshots, wherein screenshots of different variations of the same action are labeled in the clusters; extracting a sequence from the cluster and discarding subsequent events on the screen from the cluster; as well as The RPA workflow is generated based on the extracted sequence.
9. The system of claim 8, wherein the capturing further comprises: adaptive parameter tuning to iterate the template and tune the capture for subsequent iterations; clustering minutiae comprising binary feature vectors indicating the presence of template items; and by learning a sparse representation of the screen and tuning the novelty of the clustering granularity.
10. The system of claim 9, wherein the templating utilizes a threshold to indicate the plurality of words.
11. The system of claim 10, wherein the threshold comprises: 70%。 12. The system of claim 8, wherein the extracting comprises one of: forward link estimation, using a forward link prediction module to consider each event and link future events with said each event; and A graphical representation where each graph node corresponds to a screen type found during said clustering.
13. The system of claim 12, wherein an edge of the graph represents each event and events linked to the event.
14. The system of claim 8, wherein the clustering utilizes optical character recognition (OCR) data to extract word and position pairs.
15. A non-transitory computer-readable medium, comprising a computer program product recorded on the non-transitory computer-readable medium and executable by a processor, the computer program product comprising program code instructions for performing sequence extraction using screenshot images to generate a robotic process automation (RPA) workflow by implementing the following steps, the steps comprising: capturing, using a processor, a plurality of screenshots of steps performed by a user on an application, wherein the capturing includes templating to find a plurality of words and a position corresponding to each of the plurality of words to form a template; storing the plurality of captured screenshots in a memory; determining action clusters from the plurality of captured screenshots by randomly clustering actions into any predefined number of clusters based on the template applied to the plurality of captured screenshots, wherein screenshots of different variations of the same action are labeled in the clusters; extracting a sequence from the cluster and discarding subsequent events on the screen from the cluster; as well as The RPA workflow is generated based on the extracted sequence.
Citation Information
Patent Citations
System and method for capture of user actions and use of capture data in business processes
US20060184410A1
Shape Clustering in Post Optical Character Recognition Processing
US20100232719A1
Automated web task procedures based on an analysis of actions in web browsing history logs
US20140019979A1