Training of artificial intelligence / machine learning models for recognizing computer vision-based applications, screens, and user interface elements
AI/ML models trained with CV for UI element recognition address the limitations of current UI automation by enabling accurate recognition and interaction without system-level inputs, improving RPA efficiency.
Patent Information
- Application Number
- JP2023518733
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-14
- Filing Date
- 2021-10-05
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2041-10-05
AI Technical Summary
Current UI automation technologies face challenges in providing effective automation without system-level and application-level interactions, requiring extensive driver-level and application-level functionality, and lack alternative methods for recognizing applications, screens, and UI elements.
Training AI/ML models using computer vision (CV) to recognize applications, screens, and UI elements, and user interactions, utilizing recorder processes to capture screen shots or video frames, and employing feedback loops with OCR for accurate identification and interaction recognition.
Enables reliable recognition and interaction with UI elements without prior knowledge, enhancing UI automation capabilities and facilitating robotic process automation (RPA) across diverse applications and environments.
Smart Images

Figure 0007714029000001 
Figure 0007714029000002 
Figure 0007714029000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This is an international application claiming the benefit and priority of U.S. Patent Application No. 17 / 070,108, filed on October 14, 2020. The subject matter of the previously filed application is incorporated herein by reference in its entirety.
[0002] The present invention generally relates to user interface (UI) automation, and more specifically, to using computer vision (CV) to recognize applications, screens, and UI elements and to train artificial intelligence (AI) / machine learning (ML) models for recognizing user interactions with applications, screens, and UI elements.
Background Art
[0003] To perform UI automation, RPA technology can utilize driver - and / or application - level interactions to click buttons, input text, and perform other interactions with the UI. However, keystrokes, mouse clicks, and other kernel hook information may not be available at the system level in some embodiments or when building a new UI automation platform. Implementing such a UI automation platform generally requires extensive driver - level and application - level functionality. Thus, alternative technologies for providing UI automation can be beneficial.
Summary of the Invention
[0004] Certain embodiments of the present invention may provide solutions to problems and needs in the art that are not yet fully identified, evaluated, or solved by current UI automation technologies. For example, some embodiments of the present invention relate to training AI / ML models for recognizing applications, screens, and (UI) elements using CV and for recognizing user interactions with applications, screens, and UI elements.
[0005] In an embodiment, the system includes one or more user computing systems each including a respective recorder process, and a server configured to use CV to recognize applications, screens, and UI elements and to train an AI / ML model to recognize user interactions with applications, screens, and UI elements. Each recorder process is configured to record screen shots or video frames of the display associated with each of the user computing systems and other information. Each recorder process is also configured to transmit the recorded screen shots or video frames, and other information, to storage accessible by the server. The server is first configured to train an AI / ML model to recognize applications, screens, and UI elements present in the recorded screen shots or video frames using the recorded screen shots or video frames and other information. After the AI / ML model is able to reliably recognize applications, screens, and UI elements within the recorded screen shots or video frames, the server is also configured to train the AI / ML model to recognize individual user interactions with UI elements.
[0006] In another embodiment, the non-transitory computer-readable medium stores a computer program configured to use CV to recognize applications, screens, and UI elements and / or to train an AI / ML model to recognize user interactions with applications, screens, and UI elements. The computer program is configured such that at least one processor accesses recorded screenshots or video frames of a display associated with one or more computing systems and accesses other information associated with the one or more computing systems. The computer program is also configured such that at least one processor first uses the recorded screenshots or video frames and other information to train an AI / ML model to recognize applications, screens, and UI elements present in the recorded screenshots or video frames. The initial training of the AI / ML model is performed without prior knowledge of the applications, screens, and UI elements within the screenshots or video frames.
[0007] In yet another embodiment, a computer-implemented method for training an AI / ML model to recognize applications, screens, and UI elements using CV and to recognize user interactions with the applications, screens, and UI elements includes accessing a recorded screenshot or video frame of a display associated with one or more computing systems and accessing other information associated with the one or more computing systems. The computer-implemented method also includes first training an AI / ML model to recognize applications, screens, and UI elements present in the recorded screenshot or video frame using the recorded screenshot or video frame and the other information. After the AI / ML model is able to reliably recognize applications, screens, and UI elements within the recorded screenshot or video frame, the computer-implemented method further includes training the AI / ML model to recognize individual user interactions with the UI elements.
Brief Description of the Drawings
[0008] To facilitate an understanding of the advantages of particular embodiments of the present invention, a more particular description of the invention briefly described above is depicted with reference to the particular embodiments illustrated in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope, but that the invention will be described and explained with additional particularity and detail by use of the following accompanying drawings.
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3
[0012]
Figure 4
[0013]
Figure 5
[0014]
Figure 6
[0015]
Figure 7
[0016]
Figure 8
Mode for Carrying Out the Invention
[0017] Unless otherwise noted, like reference characters consistently refer to corresponding features throughout the attached drawings.
[0018] (Detailed Description of Embodiments) Some embodiments relate to training an AI / ML model for recognizing applications, screens, and (UI) elements using CV and for recognizing user interactions with applications, screens, and UI elements. In certain embodiments, optical character recognition (OCR) may also be used to assist in training the AI / ML model. In some embodiments, training of the AI / ML model can be performed without other system inputs such as system-level information (e.g., key presses, mouse clicks, position, operating system actions, etc.) or application-level information (e.g., information from an application programming interface (API) from a software application running on a computing system) as provided by a driver of UiPath Studio (trademark). However, in certain embodiments, training of the AI / ML model can be supplemented with other information such as browser history, file information, currently running applications and positions, system-level and / or application-level information.
[0019] Some embodiments start training of an AI / ML model by providing labeled screen images of an initial version of the AI / ML model as training inputs from one or more computing systems. The AI / ML model provides as output predictions such as what applications (if any) and graphical elements (if any) are present within the screen. Certain errors can be highlighted by a human reviewer (e.g., by drawing a box around a misidentified element and including the correct identification), and the AI / ML model can be trained until its accuracy is high enough to be deployed for observing applications and graphical elements present on the screens of the UI.
[0020] In some embodiments, instead of training only from images, tracking codes can also be embedded in the user's computing system. For example, a JavaScript (registered trademark) snippet can be embedded as a listener in a web browser to track what components the user interacted with, what text the user entered, which position / component the user clicked on with the mouse, what content the user scrolled past, how long the user stopped at a particular part of the content, etc. Scrolling past content may indicate that the content was somewhat close to what the user was looking for but did not exactly have it. A click may indicate success.
[0021] The listener application need not be JavaScript (registered trademark) and may be any suitable type of application and any desired programming language without departing from the scope of the present invention. This enables the "generalization" of the listener application and allows tracking of user interactions with multiple applications or any application with which multiple applications or users are interacting. Using labeled training data from scratch may enable an AI / ML model to recognize various controls, but can be difficult because it does not contain information on how any control is generally used. The listener application can be used to generate a "heatmap" and help bootstrap the training process of the AI / ML model. The heatmap can include various information such as how frequently a user used the application, how frequently the user interacted with components of the application, the location of the components, the content of the application / component, etc. In some embodiments, the heatmap can be derived from screen analysis such as detection of typed and / or pasted text, caret tracking, and active element detection of the computing system. Some embodiments recognize where on a screen associated with a computing system a user typed or pasted text that may include hotkeys or other keys where no visible text is displayed, and provide a physical location on the screen based on the current resolution (e.g., in coordinates) of the location where one or more characters were displayed, the location where the cursor was blinking, or both. The typed or pasted activity and / or the physical location of the caret can determine what field(s) the user is typing or focusing on and what application is for process discovery or other applications.
[0022] Some embodiments are implemented in a feedback loop process that continuously or periodically compares a current screenshot to a previous screenshot to identify changes. The location where a visual change has occurred on the screen can be identified, and optical character recognition (OCR) can be performed on the location where the change has occurred. Next, the results of the OCR can be compared to the content of the keyboard queue (e.g., as determined by a keyhook) to determine if a match exists. The location where the change has occurred can be determined by comparing a box of pixels from the current screenshot to a box of pixels at the same location in the previous screenshot. When a match is found, the text at the location where the change has occurred is associated with that location and can be provided as part of the listener information.
[0023] When a heatmap is generated, based on the initial heatmap information, the AI / ML model can be trained on screen images (which could be millions of images). A graphics processing unit (GPU) can process this information and train the AI / ML model relatively quickly. Once graphical elements, windows, applications, etc. can be accurately identified, the AI / ML model can be trained to recognize interactions with applications within the UI by the labeled user and understand the incremental actions the user performs. A change in one or a series of graphical elements may indicate that the user has clicked a button, entered text, interacted with a menu, closed a window, moved to another screen of an application, etc. For example, a menu item that the user has clicked can be underlined, a button can become dim while it is being pressed and then return to its original color when the user releases the mouse button, the letter 'a' can be displayed in a text field, an image can change to another image, the screen can take on a different layout when the user moves to the next screen of a series of screen application, etc.
[0024] Specific errors can be highlighted again by a human reviewer (e.g., by drawing a box around the mis-identified element and including the correct identification). The AI / ML model can then be trained until its accuracy becomes high enough to be deployed and it can understand detailed user interactions with the UI. For example, such a trained AI / ML model can then be used to observe multiple users and look for sequences of common interactions in a common application.
[0025] In some embodiments, the training of the AI / ML model is implemented via hardware or software and can be supplemented with information from an “automation box” that observes what information comes from input devices such as a mouse or keyboard. In certain embodiments, a camera can be used to track where on the screen the user is looking. Information from the automation box and / or the camera is timestamped and used in combination with the graphical elements, applications, and screens detected by the AI / ML model to assist in its training and better understand what the user is doing at that point in time.
[0026] Certain embodiments may be employed in robotic process automation (RPA). FIG. 1 is an architectural diagram showing an RPA system 100 according to an embodiment of the present invention. The RPA system 100 includes a designer 110 that enables developers to design and implement workflows. The designer 110 provides solutions for application integration and automates third-party applications, administrative information technology (IT) tasks, and business IT processes. The designer 110 can facilitate the development of automation projects that are graphical representations of business processes. Put simply, the designer 110 facilitates the development and deployment of workflows and robots.
[0027] Automation projects enable the automation of rule - based processes by giving developers control over the order of execution and relationships between a custom set of steps developed in a workflow defined herein as "activities". A commercial example of an embodiment of Designer 110 is UiPath Studio (trademark). Each activity may include actions such as clicking a button, reading a file, writing to a log panel, etc. In some embodiments, workflows may be nested or embedded.
[0028] Types of workflows may include, but are not limited to, sequences, flowcharts, FSMs, and / or global exception handlers, etc. A sequence may be particularly suitable for a linear process that enables the flow from one activity to another without cluttering the workflow. A flowchart may be particularly suitable for more complex business logic, enabling the integration of decision - making and connection of activities in more diverse ways through multiple branching logic operators. An FSM may be particularly suitable for large - scale workflows. An FSM may use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. A global exception handler may be particularly suitable for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0029] When a workflow is developed within Designer 110, the execution of the business process is coordinated by Conductor 120, which coordinates one or more robots 130 that execute the workflows developed within Designer 110. A commercial example of an embodiment of Conductor 120 is UiPath Orchestrator (trademark). Conductor 120 facilitates the management of the generation, monitoring, and deployment of resources in an environment. Conductor 120 may operate as an integration point, or one of the integration points, with third - party solutions and applications.
[0030] The conductor 120 can manage all robots 130 and connect and execute the robots 130 from a centralized point. The types of robots 130 that can be managed include, but are not limited to, attended robots 132, unattended robots 134, development robots (similar to unattended robots 134 but used for development and testing purposes), and non-production robots (similar to attended robots 132 but used for development and testing purposes). The attended robot 132 is triggered by user events and operates alongside humans on the same computing system. The attended robot 132 can be used with the conductor 120 for centralized process deployment and logging media. The attended robot 132 may assist a human user in achieving various tasks and may be triggered by user events. In some embodiments, the process cannot start from the conductor 120 on this type of robot and / or they cannot execute under a locked screen. In certain embodiments, the attended robot 132 can only be launched from a robot tray or from a command prompt. The attended robot 132 preferably operates under human supervision in some embodiments.
[0031] The unattended robot 134 can operate unmanned in a virtual environment and automate many processes. The unattended robot 134 can be responsible for providing support for remote execution, monitoring, scheduling, and work queues. Debugging for all robot types can, in some embodiments, be performed by the designer 110. Both attended and unattended robots can automate various systems and applications including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0032] The conductor 120 can have various capabilities including, but not limited to, provisioning, deployment, versioning, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning can include creating and maintaining a connection between the robot 130 and the conductor 120 (e.g., a web application). Deployment can include ensuring the correct delivery of the package version to the robot 130 assigned for execution. Versioning may, in some embodiments, include the management of unique instances of some processes or configurations. Configuration can include maintaining and delivering the robot environment and process configurations. Queuing can include providing the management of queues and queue items. Monitoring can include tracking specific data of the robot and maintaining user permissions. Logging can include the storage and indexing of logs to a database (e.g., an SQL database) and / or another storage mechanism (e.g., ElasticSearch® which provides the ability to store large datasets and execute queries quickly). The conductor 120 can provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0033] The robot 130 is an execution agent that executes the workflow built by the designer 110. One commercial example of some embodiments of the robot(s) 130 is UiPath Robots™. In some embodiments, the robot 130, by default, installs the Microsoft Windows® Service Control Manager (SCM) management service. As a result, such a robot 130 can open an interactive Windows® session under a local system account and can have the rights of a Windows® service.
[0034] In some embodiments, the robot 130 can be installed in user mode. For such a robot 130, it means having the same rights as the user in which the given robot 130 is installed. This feature may also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 130 can be configured in an HD environment.
[0035] The robot 130 in some embodiments is divided into a plurality of components, each specialized for a specific automation task. The robot components in some embodiments include, but are not limited to, SCM management robot services, user mode robot services, an executor, an agent, and a command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 120 and the execution host (i.e., the computing system on which the robot 130 is executed). These services are entrusted with managing the qualification information of the robot 130. The console application is launched by the SCM under the local system.
[0036] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 120 and the execution host. The user mode robot service may be entrusted with managing the qualification information of the robot 130. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0037] The executor can perform jobs given under a Windows® session (i.e., can execute a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows® Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0038] As described above, the fact that the components of the robot 130 are split helps developers, support users, and computing systems to more easily perform, identify, and track what each component is doing. In this way, special behavior can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be executed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 110 can be made independent of the browser zoom level. In the case of applications marked as not recognizing or intentionally not recognizing DPI, DPI can be disabled in some embodiments.
[0039] Figure 2 is an architectural diagram showing the deployed RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 may be, or may be a part of, the RPA system 100 of FIG. 1. It should be noted that either the client side, the server side, or both can include any desired number of computing systems without departing from the scope of the present invention. On the client side, the robot application 210 includes an executor 212, an agent 214, and a designer 216. However, in some embodiments, the designer 216 may not be executed on the computing system 210. The executor 212 is executing a process. As shown in FIG. 2, a plurality of business projects can be executed simultaneously. The agent 214 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 212. All messages in this embodiment are recorded in the conductor 230, which further processes them via the database server 240, the indexer server 250, or both. As described above with respect to FIG. 1, the executor 212 can be a robot component.
[0040] In some embodiments, the robot represents an association between a machine name and a user name. The robot can manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots can be executed simultaneously, each running in a separate Windows® session using a unique user name. This is called the HD robot described above.
[0041] Agent 214 is also responsible for sending the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. The communication between Agent 214 and Conducter 230 is, in some embodiments, always initiated by Agent 214. In a notification scenario, Agent 214 may open a WebSocket channel that is later used by Conducter 230 to send commands (e.g., start, stop, etc.) to the robot.
[0042] On the server side, there are a presentation layer (web application 232, Open Data Protocol (OData) Representational State Transfer (REST) Application Programming Interface (API) endpoint 234, notification and monitoring 236), a service layer (API implementation / business logic 238), and a persistence layer (database server 240, indexer server 250). The conductor 230 includes the web application 232, the OData REST API endpoint 234, notification and monitoring 236, and the API implementation / business logic 238. In some embodiments, most actions that a user performs at the interface of the conductor 230 (e.g., via the browser 220) are performed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of the present invention. The web application 232 is the visual layer of the server platform. In this embodiment, the web application 232 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of the present invention. The user interacts with the web page from the web application 232 via the browser 220 in this embodiment to perform various operations for controlling the conductor 230. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or each process, start and stop robots, etc.
[0043] In addition to the web application 232, the conductor 230 also includes a service layer that exposes an OData REST API endpoint 234. However, other endpoints may be included without departing from the scope of the present invention. The REST API is consumed by both the web application 232 and the agent 214. The agent 214 is, in this embodiment, a supervisor of one or more robots on a client computer.
[0044] The REST API of this embodiment covers configuration, logging, monitoring, and queuing functions. Configuration endpoints may be used, in some embodiments, to define and configure the users, permissions, robots, assets, releases, and environments of an application. The logging REST endpoint may be used, for example, to log various information such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by a robot to query the version of the package to be executed when a job start command is used in the conductor 230. The queuing REST endpoint may be responsible for the management of queues and queue items, such as adding data to a queue, retrieving transactions from a queue, and setting the status of a transaction.
[0045] Monitoring of the REST endpoint may monitor the web application 232 and the agent 214. The notification and monitoring API 236 may be a REST endpoint used for the registration of the agent 214, the delivery of configuration settings to the agent 214, and the sending and receiving of notifications from the server and the agent 214. The notification and monitoring API 236 may use WebSocket communication in some embodiments.
[0046] In this embodiment, the persistent layer includes a pair of server-database servers 240 (e.g., SQL servers) and an indexer server 250. The database server 240 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 232. The database server 240 may manage queues and queue items. In some embodiments, the database server 240 may store (in addition to or instead of the indexer server 250) messages recorded by robots.
[0047] Although optional in some embodiments, the indexer server 250 stores information recorded by robots and creates indexes. In certain embodiments, the indexer server 250 may be disabled via a configuration setting. In some embodiments, the indexer server 250 uses ElasticSearch (registered trademark), an open-source project full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 250 via logging REST endpoint(s), where they are indexed for future use.
[0048] Figure 3 is an architecture diagram showing the relationship 300 between the designer 310, activities 320, 330, driver 340, and AI / ML model 350 according to an embodiment of the present invention. As described above, the developer uses the designer 310 to develop a workflow to be performed by a robot. The workflow may include user-defined activities 320 and UI automation activities 330. The user-defined activities 320 and / or UI automation activities 330 may be located locally and / or remotely with respect to the computing system on which the robot operates in some embodiments and may call one or more AI / ML models 350. In some embodiments, non-text visual components in an image can be identified, which is referred to herein as computer vision (CV). Some CV activities related to such components may include, but are not limited to, click, type, get text, hover, detect presence of an element, update scope, highlight, etc. In some embodiments, click may identify an element using, for example, CV, optical character recognition (OCR), fuzzy text matching, and multi-anchor and click on it. Type may identify an element using the above and types within the element. Getting text may identify the location of specific text and scan it using OCR. Hover may identify an element and hover over it. Detecting the presence or absence of an element may confirm whether an element exists on the screen using the techniques described above. In some embodiments, there may be hundreds or thousands of activities that can be implemented in the designer 310. However, any number and / or type of activities can be utilized without departing from the scope of the present invention.
[0049] The UI automation activity 330 is a subset of special low-level activities described in low-level code (e.g., CV activities) that facilitate interaction with the screen. The UI automation activity 330 facilitates these interactions via a driver 340 and / or an AI / ML model 350 that enable the robot to interact with the desired software. For example, the driver 340 may include an OS driver 342, a browser driver 344, a VM driver 346, an enterprise application driver 348, and the like. To determine the execution of interactions with the computing system, one or more AI / ML models 350 may be used by the UI automation activity 330. In some embodiments, the AI / ML models 350 may enhance or completely replace the drivers 340. In fact, in certain embodiments, the driver 340 is not included.
[0050] The drivers 340 may interact with the OS at a low level, such as by looking for hooks and monitoring keys. They may also facilitate integration with Chrome (TM), IE (TM), Citrix (TM), SAP (TM), etc. For example, a "click" activity serves the same role in these different applications via the driver 340.
[0051] Figure 4 is an architectural diagram showing an RPA system 400 according to an embodiment of the present invention. In some embodiments, the RPA system 400 may be or may include the RPA systems 100 and / or 200 of FIGS. 1 and / or 2. The RPA system 400 includes a plurality of client computing systems 410 that execute robots. The computing system 410 can communicate with the conductor computing system 420 via a web application executed thereon. The conductor computing system 420 can in turn communicate with a database server 430 and any indexer server 440.
[0052] Regarding FIGS. 1 and 3, although web applications are used in these embodiments, it should be noted that any suitable client / server software can be used without departing from the scope of the present invention. For example, the conductor may execute a server-side application on the client computing system that communicates with a non-web-based client software application.
[0053] FIG. 5 is an architectural diagram illustrating a computing system 500 configured to train an AI / ML model to recognize applications, screens, and UI elements using CV and to recognize user interactions with the applications, screens, and UI elements, according to an embodiment of the present invention. In some embodiments, computing system 500 may be one or more computing systems depicted and / or described herein. Computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to bus 505 for processing information. The one or more processors 510 can be any type of general or special-purpose processor including a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The one or more processors 510 may also have multiple processing cores, and at least some of the cores may be configured to perform a particular function. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one of the one or more processors 510 can be a neuromorphic circuit including processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0054] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by processor(s) 510. Memory 515 may be constituted by a random access memory (RAM), a read-only memory (ROM), a flash memory, a cache, a static storage device such as a magnetic disk or an optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media may be any available media accessible by processor(s) 510 and may include volatile media, non-volatile media, or both. Also, the media may be removable, non-removable, or both.
[0055] Furthermore, computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via a wireless and / or wired connection. In some embodiments, the communication device 520 is configured to use Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), Global System for Mobile (GSM) communication, General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced (LTE-A), 802.11x, Wi-Fi, Zigbee, Ultra-WideBand (UWB) radio, 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near-Field Communications (NFC), Fifth Generation (5G), New Radio (NR), any combination thereof, and / or be configured to use any other currently existing or future implemented communication standard and / or protocol without departing from the scope of the present invention.In some embodiments, the communication device 520 may include one or more antennas that, without departing from the scope of the present invention, are a single antenna, an array antenna, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration.
[0056] The processor(s) 510 is further coupled via the bus 505 to a display 525 such as a plasma display, a liquid crystal display (LCD), a light emitting diode (LED) display, a field emission display (FED), an organic light emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high definition display, a Retina (registered trademark) display, an IPS (In-Plane Switching) display, or any other suitable display for presenting information to the user. The display 525 may be configured as a touch (haptic) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-touch display, etc., using a resistive method, a capacitive method, a surface acoustic wave (SAW) capacitive method, an infrared method, an optical imaging method, a dispersion signal method, an acoustic pulse recognition method, a frustrated total internal reflection method, etc. Without departing from the scope of the present invention, any suitable display device and haptic I / O may be used.
[0057] Keyboards 530 and cursor control devices 535, such as computer mice, touchpads, etc., are further coupled to bus 505 to enable a user to interface with computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user can interact with the device only via display 525 and / or a touchpad (not shown). Any type and combination of input devices can be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, the user may interact with computing system 500 remotely via another computing system communicating with computing system 500, or computing system 500 may operate autonomously.
[0058] Memory 515 stores software modules that provide functionality when executed by processor(s) 510. The modules include an operating system 540 for computing system 500. The modules further include an AI / ML model training module 545 configured to execute all or part of the processes described herein or derivatives thereof. Computing system 500 may include one or more additional functional modules 550 including additional functionality.
[0059] One skilled in the art will understand that the "system" can be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices, without departing from the scope of the present invention. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the present invention in any way, but rather to provide an example of many embodiments of the present invention. In fact, the methods, systems, and devices disclosed herein may be implemented in a localized and distributed form that is consistent with computing techniques including cloud computing systems. The computing system may be part of, or accessible by, a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public cloud or a private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of the present invention.
[0060] It should be noted that some of the system features described herein are presented as modules to emphasize implementation independence more. For example, a module can be implemented as a hardware circuit including off-the-shelf semiconductors such as custom very large scale integration (VLSI) circuits or gate arrays, logic chips, transistors, or other discrete components. Also, a module can be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, and graphics processing units.
[0061] The module may also be at least partially implemented in software for execution by various types of processors. For example, a specified unit of executable code may include one or more physical or logical blocks of computer instructions that may be organized as, for example, objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and may include modules when logically combined and may include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module may be stored on a non-transitory computer-readable medium such as, for example, a hard disk drive, flash device, RAM, tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of the present invention.
[0062] In fact, a module of executable code may be a single instruction, may be multiple instructions, and may even be distributed among multiple different code segments, between different programs, and among multiple memory devices. Similarly, the operational data may be specified within the module, may be shown herein, may be embodied in any suitable form, and may be organized within any suitable type of data structure. The operational data may be collected as a single data set or may be distributed in different locations across different storage devices and may exist at least partially simply as electronic signals on a system or network.
[0063] FIG. 6 is an architectural diagram illustrating a system 600 configured to train an AI / ML model to recognize applications, screens, and UI elements using CV and to recognize user interactions with the applications, screens, and UI elements, according to an embodiment of the present invention. System 600 includes user computing systems such as desktop computer 602, tablet 604, and smartphone 606. However, any desired computing system, including but not limited to smartwatches, laptop computers, etc., can be used without departing from the scope of the present invention. In some embodiments, one or more of the computing systems 602, 604, 606 may include an automation box and / or a camera. Also, although FIG. 6 shows three user computing systems, any suitable number of computing systems can be used without departing from the scope of the present invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of computing systems may be used.
[0064] Each computing system 602, 604, 606 has a recorder process 610 (i.e., a tracking application) that records screenshots and / or videos of the user's screen or a portion thereof that are executed thereon. For example, a snippet of JavaScript® can be embedded in the web browser as the recorder process 610 to track what components the user interacted with, what text the user entered, what position / component the user clicked on with the mouse, what content the user scrolled past, how long the user stopped at a particular portion of the content, etc. Scrolling past content may indicate that the content was somewhat close to what the user was looking for but did not have it exactly. A click may indicate success.
[0065] The recorder process 610 need not be JavaScript (registered trademark) and may be any suitable type of application and any desired programming language without departing from the scope of the present invention. This enables the "generalization" of the recorder process 610 and allows tracking of user interactions with multiple applications or any application with which multiple applications or users are interacting. Using labeled training data from scratch can enable an AI / ML model to recognize various controls, but it can be difficult because it does not contain information on how each control is generally used. The recorder process 610 can be used to generate a "heatmap" and help bootstrap the training process of the AI / ML model. The heatmap can include various information such as how frequently a user uses an application, how frequently a user interacts with components of an application, the position of the components, the content of the application / component, etc. In some embodiments, the heatmap can be derived from screen analysis such as detection of typed and / or pasted text, caret tracking, and detection of active elements of computing systems 602, 604, 606. Some embodiments recognize where on the screen associated with computing systems 602, 604, 606 the user typed or pasted text that may include hotkeys or other keys where no visible text is displayed, and provide a physical position on the screen based on the position where one or more characters were displayed, the position where the cursor was blinking, or both, at the current resolution (e.g., in coordinates). The typed or pasted activity and / or the physical position of the caret can determine which field(s) the user is typing or focusing on and what the application is for process discovery or other applications.
[0066] As described above, in some embodiments, the recorder process 610 may record additional data to further assist in training the AI / ML model(s), such as web browser history, heatmaps, key presses, mouse clicks, the position of mouse clicks and / or graphical elements on the screen with which the user is interacting, the position where the user was looking at the screen at any given time, timestamps associated with screenshots / video frames, etc. This can be beneficial for providing key presses and / or other user actions that may not cause a screen change. For example, some applications may not provide a visual change when the user presses CTRL+S to save a file. However, in certain embodiments, the AI / ML model(s) may be trained based only on the captured screen images. The recorder process 610 can be a robot generated via an RPA designer application, part of an operating system, a downloadable application for a personal computer (PC) or smartphone, without departing from the scope of the present invention, or any other software and / or hardware. In fact, in some embodiments, the logic of one or more recorder processes 610 is implemented partially or fully via physical hardware.
[0067] Some embodiments are implemented in a feedback loop process that continuously or periodically compares the current screenshot with previous screenshots to identify changes. The position where a visual change has occurred on the screen can be identified, and OCR can be performed on the position where the change has occurred. Next, the result of the OCR can be compared with the content of the keyboard queue (e.g., determined by key hooks) to determine whether a match exists. The position where the change has occurred can be determined by comparing the pixel box from the current screenshot with the pixel box at the same position in the previous screenshot.
[0068] The images and / or other data recorded by the recorder process 610 (e.g., web browser history, heatmaps, key presses, mouse clicks, positions of mouse clicks and / or graphical elements on the screen where the user is interacting, positions on the screen that the user looked at during a time period, screenshots / video frames, voice inputs, gestures, emotions (such as whether the user is satisfied or irritated), timestamps associated with biometrics (fingerprints, retina scans, the user's pulse, etc.), information related to periods when there is no user activity (e.g., a "dead man's switch"), haptic information from a haptic display or touchpad, heatmaps from multi-touch inputs, etc.) are sent to the server 630 via the network 620 (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). In some embodiments, the server 630 may be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the server 630 may host multiple software-based servers on a single computing system 630. In some embodiments, the server 630 may execute a conductor application, and the data from the recorder process 610 may be sent periodically as part of a heartbeat message. In certain embodiments, the data may be sent from the recorder process 610 to the server 630 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. The server 630 stores the received data from the recorder process 610 in the database 640.
[0069] In this embodiment, server 630 includes a plurality of AI layers 632 that collectively form an AI / ML model. However, in some embodiments, the AI / ML model may have only a single layer. In certain embodiments, multiple AI / ML models may be trained on server 630 and used together to collectively achieve a larger task. The AI layers 632 may employ CV techniques, may perform various functions such as statistical modeling (e.g., Hidden Markov Model (HMM)), and may utilize deep learning techniques (e.g., Long Short-Term Memory (LSTM) deep learning, encoding of previous hidden states, etc.) to identify user interactions. Initially, the AI / ML model needs to be trained so that it can perform a meaningful analysis of the data captured within database 640. In some embodiments, users of computing systems 602, 604, 606 label the images before they are sent to server 630. Additionally or alternatively, in some embodiments, the labeling occurs later, such as via application 652 running on computing system 650 that enables the user to draw a bounding box and / or other shapes around graphical elements, provide text labels for what is contained within the bounding box, etc.
[0070] The AI / ML model uses this data as input and undergoes a training phase, during which the AI / ML model is trained until it is accurate enough and does not overfit to the training data. The acceptable accuracy may depend on the application. Certain errors are highlighted by a human reviewer (e.g., by drawing a box around the mis-identified element and including the correct identification), and the AI / ML model can be retrained using this additional labeled data. Once sufficiently trained, the AI / ML model can provide as output predictions such as what applications (if any) and graphical elements (if any) are present within the screen.
[0071] However, this level of training provides information about what exists, but additional information may be needed to determine user interactions, such as comparing two or more consecutive screens to determine that typed text has appeared from one thing to another, a button has been pressed, a menu selection has occurred, etc. Thus, after the AI / ML model has been able to recognize graphical elements and applications on the screen, in some embodiments, the AI / ML model is further trained to recognize labeled user interactions with the applications within the UI to understand such incremental actions taken by the user. Certain errors can be re-emphasized by a human reviewer (e.g., by drawing a box around the misidentified element and including the correct identification), and the AI / ML model can be trained until its accuracy is high enough to be deployed to understand detailed user interactions with the UI.
[0072] Once trained to recognize user interactions, the trained AI / ML model can be used to analyze video and / or other information from the recorder process 610. This recorded information can include interactions that multiple users tend to perform. These interactions can then be analyzed for a common sequence for subsequent automation.
[0073] AI layer
[0074] In some embodiments, multiple AI layers may be used. Each AI layer is an algorithm (or model) that operates on data, and the AI model itself can be a deep learning neural network (DLNN) of artificial "neurons" trained on training data. The layers can be executed in series, parallel, or a combination thereof.
[0075] The Al layer may include, but is not limited to, a sequence extraction layer, a clustering detection layer, a visual component detection layer, a text recognition layer (e.g., OCR), a speech-text translation layer, or any combination thereof. However, without departing from the scope of the present invention, any desired number and type(s) of layers may be used. By using multiple layers, the system can develop a global image of what is occurring within the screen. For example, one AI layer may be able to perform OCR, and another may be able to detect buttons, etc.
[0076] The pattern may be determined individually by an AI layer or collectively by multiple AI layers. Probabilities or outputs regarding user actions may be used. For example, to determine the identification of a button, its text, where the user clicked, etc., the system may need to know where the button is, its text, its position on the screen, etc.
[0077] However, it should be noted that without departing from the scope of the present invention, various AI / ML models may be used. The AI / ML models may be trained using neural networks such as DLNN, recurrent neural networks (RNN), generative adversarial networks (GAN), any combination thereof, etc. in some embodiments, but other AI techniques may be used without departing from the scope of the present invention, such as deterministic models, shallow learning neural networks (SLNN), or any other suitable type of AI / ML model and training techniques.
[0078] Figure 7 is a flowchart illustrating process 700 for training an AI / ML model to recognize applications, screens, and UI elements using CV and to recognize user interactions with the applications, screens, and UI elements. The process begins at 710 by recording screenshots or video frame displays related to the user computing system and other information. In some embodiments, the recording is performed by one or more recorder processes. In certain embodiments, the recorder process is implemented as a feedback loop process that continuously or periodically compares the current screenshot or video frame to previous screenshots or video frames and identifies one or more positions where a change has occurred between the current and previous screenshots or video frames. In some embodiments, the recorder process is configured to perform OCR on the one or more positions where a change has occurred, compare the results of the OCR to the content of the keyboard queue to determine if a match exists, and if a match exists, link the text associated with the match to each position. In some embodiments, the other information includes the web browser history, one or more heatmaps, key presses, mouse clicks, the position of mouse clicks and / or graphical elements on the display with which the user is interacting, the position on the display where the user was looking, the timestamp associated with the screenshot or video frame, the text entered by the user, the content scrolled past by the user, the time the user paused on a portion of the content displayed on the display, what application the user is interacting with, or a combination thereof. In certain embodiments, at least a portion of the other information is captured using one or more automation boxes.
[0079] One or more heatmaps are generated at 720 as part of other information. In some embodiments, one or more heatmaps include the frequency with which a user uses the application, the frequency with which the user interacts with components of the application, the location of components within the application, the content of the application and / or components, or combinations thereof. In certain embodiments, one or more heatmaps are derived from display analysis including detection of typed and / or pasted text, caret tracking, detection of active elements, or combinations thereof. The recorded screenshots or video frames, and other information, are then sent at 730 to storage accessible by one or more servers.
[0080] The recorded screenshots or video frames and other information are accessed at 740 (e.g., via a server configured to train an AI / ML model). The AI / ML model is first trained at 750 to recognize the applications, screens, and UI elements present in the recorded screenshots or video frames using the recorded screenshots or video frames and other information. In some embodiments, the initial training of the AI / ML model is performed without prior knowledge of the applications, screens, and UI elements within the screenshots or video frames.
[0081] After the AI / ML model can recognize applications, screens, and UI elements within the recorded screenshots or video frames with confidence (e.g., 70%, 95%, 99.99%, etc.), at 760, the AI / ML model is trained to recognize individual user interactions with the UI elements. In some embodiments, the individual user interactions include button presses, input of a single character or string of characters, selection of an active UI element, menu selections, screen changes, or combinations thereof. In certain embodiments, training the AI / ML model to recognize individual user interactions with the UI elements includes comparing two or more consecutive screenshots or video frames and determining that typed characters have appeared from one to another, a button has been pressed, or a menu selection has occurred. Then, the AI / ML model is deployed so that it can be called and used by an invocation process (e.g., an RPA robot) at 770.
[0082] Figure 8 is an architectural diagram illustrating an automation box and an eye movement tracking system 800 according to an embodiment of the present invention. The system 800 includes a computing system 810 that includes eye tracking logic (ETL) 812 configured to process inputs from a camera 820 and automation box logic (ABL) 814 configured to process inputs from an automation box 860. In some embodiments, the computing system 810 may be or may include the computing system 500 of FIG. 5. In certain embodiments, multiple cameras may be used.
[0083] Camera 820 records the user's video while the user is interacting with computing system 810 via mouse 840 and keyboard 850. Computing system 810 converts the recorded camera video into video frames. ETL processes these frames to identify the user's eyes and interpolate the position the user is looking at to a position on display 830. Any suitable eye tracking technique(s), such as those described in U.S. Patent Application Publication No. 2018 / 0046248, U.S. Patent No. 7,682,026, etc., can be used without departing from the scope of the present invention. Timestamps can be associated with the user's video frames so that they can be made to match the screenshot frames of what is displayed on display 830 at that time.
[0084] The automation box 860 also includes automation box logic 862 in this embodiment that receives input from the mouse 840 and the keyboard 850. In some embodiments, the automation box 860 may have hardware similar to the computing system 810 (e.g., one or more processors, memory, buses, etc.). This input can then be passed to the computing system 810. Although the mouse 840 and the keyboard 850 are shown in FIG. 8, any suitable input device(s) can be used without departing from the scope of the present invention, such as a touchpad, buttons, etc. In some embodiments, only the computing system 810 or the automation box 860 includes the automation box logic. The latter reason may be to record user interactions and send them directly to a server (e.g., a cloud-based server) for subsequent processing via the network 870. In such embodiments, the screenshot frames can also be sent from the computing system 810 to the automation box 860 and then to the server via the network 870. Alternatively, the computing system 810 may send the screenshot itself via the network 870. Such embodiments may provide a plug-and-play tracking solution that plugs into the computing system 810, relays keyboard and mouse information to the computing system 810 for its operation, and also relays keyboard and mouse click information to a remote server for subsequent training of the AI / ML model.
[0085] In some embodiments, the automation box 860 may include actuation logic that performs the automation and simulates the inputs. This may enable the automation box 860 to provide the computing system 810 with simulated key presses, mouse movements and clicks, as if this information were actually coming from a human user interacting with these components. Subsequently, UI screenshots and other information may be used to train the AI / ML model. Another advantage of such embodiments is that the AI / ML model can be trained when the user is away from the computing system 810, potentially allowing for a larger amount of training information to be incorporated more quickly, and thus potentially enabling the AI / ML model to be trained more quickly.
[0086] In certain embodiments, the "information box" may be implemented as software on the computing system 610 and may function in a manner similar to the recorder process 610 of FIG. 6. Such embodiments may store screenshot frames, mouse click information, and key press information. In certain embodiments, eye-tracking information may also be tracked. This information may then be sent to a server via the network 870, and the eye tracking may potentially be performed remotely rather than on the computing system 810.
[0087] The process steps executed in FIG. 7 may be executed by a computer program that encodes instructions to a processor (s) to execute at least a part of the process(es) described in FIG. 7 according to an embodiment of the present invention. The computer program may be stored in a non-transitory computer-readable medium. The computer-readable medium may be a hard disk drive, a flash device, RAM, a tape, and / or any other such medium or combination of media used to store data, but is not limited thereto. The computer program may include encoded instructions for controlling a processor (s) of a computing system (e.g., the processor(s) 510 of the computing system 500 in FIG. 5) to implement all or part of the process steps described in FIG. 7, which may also be stored in a computer-readable medium.
[0088] The computer program may be implemented in a hardware, software, or hybrid implementation. The computer program may be composed of modules that communicate operably with each other and are designed to send information or instructions to a display. The computer program may be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0089] It will be readily understood that the components of the various embodiments of the present invention may be arranged and designed in a variety of different configurations as generally described and illustrated herein. Accordingly, the detailed description of the embodiments of the present invention as represented in the accompanying figures is not intended to limit the scope of the present invention as claimed, but merely represents selected embodiments of the present invention.
[0090] The features, structures, or characteristics of the invention described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to “certain embodiments,” “some embodiments,” or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment of the invention. Thus, appearances of the phrases “in certain embodiments,” “in some embodiments,” “in other embodiments,” or similar language throughout this specification are not necessarily all referring to the same group of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0091] It should be noted that references throughout this specification to features, advantages, or similar language do not mean that all features and advantages realizable by the invention should be in any single embodiment of the invention or in any one embodiment of the invention. Rather, language referring to features and advantages is to be understood as meaning that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, discussions of features and advantages throughout this specification, and similar language, may refer to the same embodiment, but do not necessarily have to.
[0092] Furthermore, the described features, advantages, and characteristics of the invention may be combined in any suitable manner in one or more embodiments. Those skilled in the relevant art will recognize that the invention may be practiced without one or more of the specific features or advantages of a particular one or more embodiments. In other instances, additional features and advantages may be recognized in particular embodiments that may not be present in all embodiments of the invention.
[0093] Those having ordinary skill in the art will readily understand that the present invention as described above can be practiced using steps in a different order and / or using hardware elements of a different configuration than those disclosed. Accordingly, while the invention has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations will become apparent while remaining within the spirit and scope of the invention. Therefore, reference should be made to the appended claims to determine the scope of the present invention.
Claims
1. One or more user computing systems including each recorder process, and A server configured to train an artificial intelligence (AI) / machine learning (ML) model for recognizing applications, screens, and user interface (UI) elements using computer vision (CV) and for recognizing user interactions with the applications, screens, and UI elements, the system comprising: Each of the recorder processes Records a screen shot or video frame of a display associated with each of the user computing systems and other information, And is configured to transmit the recorded screen shot or video frame and the other information to storage accessible by the server, The server First, uses the recorded screen shot or video frame and the other information to train the AI / ML model to recognize the applications, screens, and UI elements present in the recorded screen shot or video frame, After the AI / ML model is able to reliably recognize the applications, screens, and UI elements in the recorded screen shot or video frame, the AI / ML model is trained to recognize individual user interactions with the UI elements. A system configured as such.
2. The individual user interactions include button presses, input of single characters or character strings, selection of active UI elements, menu selections, screen changes, voice inputs, gestures, provision of biometric information, haptic interactions, or combinations thereof. The system according to claim 1.
3. Training of the AI / ML model for recognizing the individual user interactions with the UI elements includes comparing two or more consecutive screen shots or video frames and determining that typed characters have appeared from one screen shot to another, a button has been pressed, or a menu selection has occurred. The system according to claim 1.
4. The other information includes the history of the web browser, one or more heatmaps, key presses, mouse clicks, the position of mouse clicks and / or graphical elements on the display with which the user is interacting, the position on the display that the user was looking at, the timestamp associated with the screenshot or video frame, the text entered by the user, the content scrolled through by the user, the time the user stopped at a part of the content displayed on the display, what application the user is interacting with, voice input, gestures, sentiment information, biometric information, information regarding periods of inactivity of the user, haptic information, multi-touch input information, or a combination thereof, for the system according to claim 1.
5. The one or more user computing systems or the server are configured to generate one or more heatmaps, and the other information includes the one or more heatmaps, The one or more heatmaps include the frequency of use of an application by the user, the frequency of interaction of the user with components of the application, the position of the components within the application, the content of the application and / or components, or a combination thereof, for the system according to claim 1.
6. The one or more user computing systems or the server are configured to derive the one or more heatmaps from display analysis including detection of typed and / or pasted text, caret tracking, detection of active elements, or a combination thereof, for the system according to claim 5.
7. An automation box operably connected to one of the one or more user computing systems, the automation box receives input from one or more user input devices, associates a timestamp with the input, and further comprises an automation box configured to transmit the timestamped input to storage accessible by the server. The system according to claim 1, wherein the server is configured to use the timestamped input for initial training of the AI / ML model.
8. The system according to claim 1, wherein the server is configured to perform initial training of the AI / ML model without prior knowledge of the applications, screens, and UI elements within the screenshot or video frame.
9. A non-transitory computer-readable medium storing a computer program configured to train an artificial intelligence (AI) / machine learning (ML) model for recognizing applications, screens, and user interface (UI) elements using computer vision (CV) and / or for recognizing user interactions with the applications, screens, and UI elements, the computer program causing at least one processor to access a recorded screenshot or video frame of a display associated with the one or more computing systems and access other information associated with the one or more computing systems, first, train the AI / ML model to recognize the applications, screens, and UI elements present in the recorded screenshot or video frame using the recorded screenshot or video frame and the other information, and perform initial training of the AI / ML model without prior knowledge of the applications, screens, and UI elements within the screenshot or video frame.
10. After the AI / ML model has been able to reliably recognize the applications, screens, and UI elements within the recorded screenshot or video frame, the computer program causes the at least one processor to The non-transitory computer-readable medium according to claim 9, wherein the computer program is further configured to train the AI / ML model to recognize individual user interactions with the UI elements.
11. The training of the AI / ML model for recognizing the individual user interactions with the UI elements involves comparing two or more consecutive screenshots or video frames to determine that typed text has appeared from one thing to another, that a button has been pressed, or that a menu selection has occurred. The non - transient computer - readable medium according to claim 10.
12. The individual user interactions include button presses, input of single characters or character strings, selection of active UI elements, menu selections, screen changes, voice inputs, gestures, provision of biometric information, haptic interactions, or combinations thereof. The non - transient computer - readable medium according to claim 10.
13. The other information includes the history of the web browser, one or more heatmaps, key presses, mouse clicks, the position of the mouse clicks and / or graphical elements on the display with which the user is interacting, the position on the display that the user was looking at, the timestamp associated with the screenshot or video frame, the text input by the user, the content scrolled through by the user, the time the user paused at a part of the content displayed on the display, what application the user is interacting with, voice input, gestures, sentiment information, biometric information, information regarding periods of inactivity of the user, haptic information, multi - touch input information, or combinations thereof. The non - transient computer - readable medium according to claim 9.
14. The computer program further causes the at least one processor to generate one or more heatmaps, and the other information is configured to include the one or more heatmaps. The one or more heatmaps of claim 9, the non-transitory computer-readable medium comprising the frequency of use of one or more applications by a user, the frequency of interaction of the user with components of the one or more applications, the location of the components within the one or more applications, the content of the one or more applications and / or components, or combinations thereof.
15. The one or more heatmaps of claim 14, the non-transitory computer-readable medium derived from display analysis including detection of typed and / or pasted text, caret tracking, detection of active elements, or combinations thereof.
16. A computer-implemented method for training an artificial intelligence (AI) / machine learning (ML) model for recognizing applications, screens, and user interface (UI) elements using computer vision (CV) and for recognizing user interactions with the applications, screens, and UI elements, the method comprising: accessing recorded screenshots or video frames of a display associated with one or more computing systems and accessing other information associated with the one or more computing systems; first, training the AI / ML model to recognize the applications, screens, and UI elements present in the recorded screenshots or video frames using the recorded screenshots or video frames and the other information; after the AI / ML model is able to reliably recognize the applications, screens, and UI elements within the recorded screenshots or video frames, training the AI / ML model to recognize individual user interactions with the UI elements.
17. The computer-implemented method of claim 16, wherein initial training of the AI / ML model is performed without prior knowledge of the applications, screens, and UI elements within the screenshots or video frames.
18. The training of the AI / ML model for recognizing the individual user interactions with the UI elements involves comparing two or more consecutive screenshots or video frames to determine that typed text has appeared from one thing to another, that a button has been pressed, or that a menu selection has occurred. The computer-implemented method according to claim 16.
19. The individual user interactions include button presses, input of single characters or strings, selection of active UI elements, menu selections, screen changes, voice inputs, gestures, provision of biometric information, haptic interactions, or combinations thereof. The computer-implemented method according to claim 16.
20. The other information includes the history of the web browser, one or more heatmaps, key presses, mouse clicks, the position of the mouse click and / or graphical elements on the display with which the user is interacting, the position on the display that the user was looking at, the timestamp associated with the screenshot or video frame, the text input by the user, the content scrolled through by the user, the time the user stopped at a part of the content displayed on the display, what application the user is interacting with, voice input, gestures, emotional information, biometric information, information regarding periods of inactivity of the user, haptic information, multi-touch input information, or combinations thereof. The computer-implemented method according to claim 16.
21. Further including generating one or more heatmaps, and the other information includes the one or more heatmaps, The one or more heatmaps include the frequency with which the user has used one or more applications, the frequency with which the user has interacted with components of the one or more applications, the position of the components within the one or more applications, the content of the one or more applications and / or components, or combinations thereof. The computer-implemented method of claim 16, wherein the one or more heat maps are derived from display analysis including detection of typed and / or pasted text, caret tracking, detection of active elements, or combinations thereof.
Citation Information
Patent Citations
Image text coordinate positioning method based on deep learning
CN111723789A
Saliency prediction for a mobile user interface
US20190251707A1
Computer-vision based execution of graphical user interface (GUI) application actions
US20200012481A1
Intelligent transportation systems
WO2020069517A2