Using computer vision to train artificial intelligence / machine learning models to recognize applications, screens, and user interface elements
By using AI/ML models trained with computer vision, combined with OCR and listener applications, the functional limitations of existing UI automation technologies at the system and application levels are solved, enabling efficient recognition of UI elements and user interaction, and improving the accuracy and flexibility of UI automation.
Patent Information
- Application Number
- CN202180070042.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-14
- Filing Date
- 2021-10-05
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-10-05
AI Technical Summary
Existing UI automation technologies have limitations in terms of system-level and application-level functionality, causing operations such as key presses and mouse clicks to be unavailable in certain situations. Alternative technologies are needed to achieve effective UI automation.
Computer vision (CV) is used to train artificial intelligence/machine learning (AI/ML) models. By recording screenshots or video frames, the models are trained to recognize application, screen, and user interface elements and to identify user interactions. Heatmaps are generated by combining optical character recognition (OCR) and listener applications to assist in model training.
It enables efficient identification and interaction of UI elements without requiring system-level and application-level information, improving the accuracy and flexibility of UI automation and making it suitable for various applications and user interaction scenarios.
Smart Images

Figure CN116391174B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This is an international application that claims the benefit and priority of U.S. Patent Application No. 17 / 070,108, filed October 14, 2020. The subject matter of that earlier application is incorporated herein by reference in its entirety. Technical Field
[0003] This invention relates generally to user interface (UI) automation, and more specifically, to using computer vision (CV) to train artificial intelligence (AI) / machine learning (ML) models to identify applications, screens, and UI elements, and to identify user interactions with applications, screens, and UI elements. Background Technology
[0004] To perform UI automation, RPA technologies can leverage driver and / or application-level interactions to click buttons, input text, and perform other interactions with the UI. However, in some implementations, or when building new UI automation platforms, key presses, mouse clicks, and other kernel hook information may not be available at the system level. Implementing such UI automation platforms typically requires extensive driver-level and application-level functionality. Therefore, alternative technologies for providing UI automation may be beneficial. Summary of the Invention
[0005] Certain embodiments of the present invention may provide solutions to problems and needs that are not fully identified, understood, or resolved by current UI automation technologies in the art. For example, some embodiments of the present invention relate to using computer vision (CV) to train AI / ML models to identify applications, screens, and UI elements, and to identify user interactions with these applications, screens, and UI elements.
[0006] In one embodiment, a system includes one or more user computing systems, each including a corresponding recorder process and a server configured to train an AI / ML model to use computer vision (CV) to identify application, screen, and UI elements and to identify user interactions with these elements. The corresponding recorder process is configured to record screenshots or video frames, along with other information, of a display associated with the corresponding user computing system. The recorder process is also configured to send the recorded screenshots or video frames, along with the other information, to server-accessible storage. The server is configured to use the recorded screenshots or video frames, along with the other information, to initially train the AI / ML model to identify application, screen, and UI elements presented in the recorded screenshots or video frames. After the AI / ML model is able to reliably identify the application, screen, and UI elements in the recorded screenshots or video frames, the server is further configured to train the AI / ML model to identify individual user interactions with the UI elements.
[0007] In another embodiment, a non-transitory computer-readable medium stores a computer program configured to train an AI / ML model to use computer vision (CV) to identify application, screen, and UI elements and / or to identify user interactions with these elements. The computer program is configured to enable at least one processor to access recorded screenshots or video frames of a display associated with one or more computing systems and to access other information associated with the one or more computing systems. The computer program is also configured to enable at least one processor to use the recorded screenshots or video frames and other information to initially train the AI / ML model to identify application, screen, and UI elements presented in the recorded screenshots or video frames. This initial training of the AI / ML model is performed without prior knowledge of the application, screen, and UI elements in the screenshots or video frames.
[0008] In yet another embodiment, a computer-implemented method for training an AI / ML model to recognize application, screen, and UI elements using computer vision (CV) and to identify user interactions with these elements includes accessing recorded screenshots or video frames of a display associated with one or more computing systems and accessing other information associated with the computing systems. The computer-implemented method further includes using the recorded screenshots or video frames and other information to initially train the AI / ML model to recognize application, screen, and UI elements presented in the recorded screenshots or video frames. After the AI / ML model is able to reliably recognize the application, screen, and UI elements in the recorded screenshots or video frames, the computer-implemented method further includes training the AI / ML model to recognize individual user interactions with the UI elements. Attached Figure Description
[0009] To facilitate understanding of the advantages of certain embodiments of the invention, a more detailed description of the invention briefly described above will be presented with reference to specific embodiments illustrated in the accompanying drawings. While it should be understood that these drawings depict only exemplary embodiments of the invention and are therefore not intended to be considered as limiting its scope, the invention will be described and explained with additional specificity and detail using the drawings, in which:
[0010] Figure 1 This is an architectural diagram illustrating a robotic process automation (RPA) system according to an embodiment of the present invention.
[0011] Figure 2 This is an architectural diagram of a deployed RPA system according to an embodiment of the present invention.
[0012] Figure 3 This is an architecture diagram illustrating the relationship between the designer, activities, and drivers according to an embodiment of the present invention.
[0013] Figure 4 This is an architectural diagram of an RPA system according to an embodiment of the present invention.
[0014] Figure 5 This is an architectural diagram of a computing system according to an embodiment of the present invention, which is configured to train an AI / ML model to use CV to identify application, screen and UI elements and to identify user interactions with application, screen and UI elements.
[0015] Figure 6 This is an architectural diagram illustrating a system configured to train an AI / ML model to use computer vision (CV) to identify application, screen, and UI elements and to identify user interactions with these elements, according to an embodiment of the present invention.
[0016] Figure 7 This is a flowchart illustrating a process according to an embodiment of the present invention for training an AI / ML model using CV to identify applications, screens, and UI elements, and for identifying user interactions with applications, screens, and UI elements.
[0017] Figure 8 This is an architectural diagram illustrating an automated frame and eye motion tracking system according to an embodiment of the present invention.
[0018] Unless otherwise stated, similar reference numerals throughout the figures indicate the corresponding features. Detailed Implementation
[0019] Some embodiments involve training AI / ML models to use CV to recognize application, screen, and UI elements and to identify user interactions with these elements. In some embodiments, optical character recognition (OCR) may also be used to aid in training the AI / ML model. In some embodiments, training of the AI / ML model can be performed without other system inputs, such as system-level information (e.g., key presses, mouse clicks, location, operating system operations, etc.) or application-level information (e.g., information from the application programming interface (API) of a software application running on a computing system), such as information provided by the UiPath Studio™ driver. However, in some embodiments, training of the AI / ML model may be supplemented by other information, such as browser history, file information, currently running applications and locations, system-level and / or application-level information, etc.
[0020] Some embodiments begin training the AI / ML model by feeding an initial version of the AI / ML model labeled screen images from one or more computing systems as training input. The AI / ML model provides predictions as output, such as which application(s) and graphical elements are identified as being presented on the screen. Errors can be highlighted by a human reviewer (e.g., by drawing a box around the incorrectly labeled element and including the correct label), and the AI / ML model can be trained until its accuracy is high enough to be deployed to observe applications and graphical elements presented on UI screens.
[0021] In some embodiments, instead of training solely from images, tracking code can be embedded into the user's computing system. For example, a piece of JavaScript® can be embedded as a listener in a web browser to track which components the user interacts with, what text the user types, which locations / components the user clicks with the mouse, what content the user scrolls over, and how long the user lingers on a particular part of the content. Scrolled-over content might indicate that the content is somewhat close but not exactly what the user wants. Clicking might indicate success.
[0022] The listener application need not be JavaScript® and can be any suitable type of application and can be in any desired programming language without departing from the scope of this invention. This allows the listener application to be “generalized” so that it can track user interactions with multiple applications or any application with which the user is interacting. Using labeled training data from scratch can be difficult because while it may allow AI / ML models to learn to recognize various controls, it does not contain information about which controls are commonly used and how they are used. Using the listener application, “heatmaps” can be generated to help guide the AI / ML model training process. Heatmaps can include various information such as how frequently the user uses the application, how frequently the user interacts with components of the application, the location of the components, the content of the application / component, etc. In some embodiments, heatmaps can be derived from screen analytics, such as the computing system’s detection of typed and / or pasted text, caret tracking, and active element detection. Some embodiments identify where the user types or pastes text on the screen associated with the computing system, potentially including hotkeys or other keys that do not result in visible characters, and provide the physical location on the screen based on the location of one or more characters, the location of the cursor blinking, or both (e.g., in coordinates) at the current resolution. The physical location of the typing or pasting activity and / or caret can allow you to determine which field(s) the user is typing or focusing on, and what the application is for process discovery or other applications.
[0023] Some embodiments are implemented in a feedback loop process that continuously or periodically compares the current screenshot with previous screenshots to identify changes. Locations of visual changes on the screen can be identified, and optical character recognition (OCR) can be performed on the locations of changes. The results of the OCR can then be compared with the contents of a keyboard queue (e.g., determined by key hooks) to determine if a match exists. The location of a change can be determined by comparing pixel boxes from the current screenshot with pixel boxes at the same location from a previous screenshot. When a match is found, the text at the location of the change can be associated with that location and provided as part of the listener information.
[0024] Once a heatmap has been generated, an AI / ML model can be trained on screen images (potentially millions of images) based on the initial heatmap information. A graphics processing unit (GPU) can likely process this information and train the AI / ML model relatively quickly. Once graphical elements, windows, applications, etc., can be accurately identified, the AI / ML model can be trained to recognize labeled user interactions with the application in the UI to understand incremental actions taken by the user. A change in one or a series of graphical elements might indicate that the user clicked a button, entered text, interacted with a menu, closed a window, moved to a different screen within the application, etc. For example, a menu item clicked by the user might become underlined, a button might be colored darker when pressed and then return to its original color when the user releases the mouse button, the letter "a" might appear in a text field, an image might change to a different image, and the screen might display a different layout as the user moves to the next screen in a series of screens within the application, and so on.
[0025] Errors can be highlighted again by a human reviewer (e.g., by drawing a box around the incorrectly labeled element and including the correct label). An AI / ML model can then be trained until its accuracy is high enough to be deployed to understand fine-grained user interactions with the UI. For example, such a trained AI / ML model can then be used to observe multiple users and look for common interaction sequences in common applications.
[0026] In some embodiments, information from “automation boxes,” implemented via hardware or software, can be used to supplement the training of AI / ML models. These “automation boxes” observe what information comes from input devices such as a mouse or keyboard. In some embodiments, a camera can be used to track where the user is looking on the screen. Information from the automation boxes and / or cameras may be timestamped and used in conjunction with graphical elements, applications, and screen data detected by the AI / ML model to assist its training and better understand what the user is doing at that time.
[0027] Some implementations can be used for robotic process automation (RPA). Figure 1 This is an architectural diagram illustrating an RPA system 100 according to an embodiment of the present invention. The RPA system 100 includes a designer 110 that allows developers to design and implement workflows. The designer 110 can provide solutions for application integration and automation of third-party applications, management of information technology (IT) tasks, and business IT processes. The designer 110 can facilitate the development of automation projects, which are graphical representations of business processes. In short, the designer 110 facilitates the development and deployment of workflows and robots.
[0028] Automation projects automate rule-based processes by allowing developers to control the execution order and the relationships between custom sets of steps developed within a workflow (defined herein as "activities"). A commercial example of an embodiment of Designer 110 is UiPath Studio™. Each activity may include actions such as clicking a button, reading a file, or writing to a log panel. In some embodiments, workflows may be nested or embedded.
[0029] Some types of workflows may include, but are not limited to, sequences, flowcharts, flow charts (FSMs), and / or global exception handlers. Sequences may be particularly suitable for linear processes, implementing a flow from one activity to another without compromising the workflow. Flowcharts may be particularly suitable for more complex business logic, integrating decisions and connecting activities in more diverse ways through multiple branching logic operators. FSMs may be particularly suitable for large workflows. FSMs can use a limited number of states in their execution, triggered by conditions (i.e., transitions) or activities. Global exception handlers may be particularly suitable for determining workflow behavior when execution errors are encountered and are suitable for debugging processes.
[0030] Once a workflow is developed in Designer 110, the execution of the business process is orchestrated by Coordinator 120, which orchestrates one or more robots 130 to execute the workflow developed in Designer 110. A commercial example of an embodiment of Coordinator 120 is UiPath Orchestrator™. Coordinator 120 facilitates the management of the creation, monitoring, and deployment of resources in the environment. Coordinator 120 can act as an integration point or aggregation point with third-party solutions and applications.
[0031] Coordinator 120 can manage a queue of robots 130, connecting and executing robots 130 from a centralized point. The types of robots 130 that can be managed include, but are not limited to, manned robots 132, unattended robots 134, development robots (similar to unattended robots 134 but used for development and testing purposes), and non-production robots (similar to manned robots 132 but used for development and testing purposes). Manned robots 132 are triggered by user events and operate alongside humans on the same computing system. Manned robots 132 can be used with coordinator 120 for centralized process deployment and logging media. Manned robots 132 can assist human users in performing various tasks and can be triggered by user events. In some embodiments, processes cannot be initiated from coordinator 120 on this type of robot and / or they cannot run under a locked screen. In some embodiments, manned robots 132 can only be initiated from a robot tray or from a command prompt. In some embodiments, manned robots 132 should operate under human supervision.
[0032] Unattended robot 134 operates unattended in a virtual environment and can automate multiple processes. Unattended robot 134 can be responsible for remote execution, monitoring, scheduling, and supporting work queues. In some embodiments, debugging for all robot types can be run in designer 110. Both manned and unattended robots can automate a wide range of systems and applications, including but not limited to mainframes, web applications, virtual machines, enterprise applications (e.g., those produced by SAP®, Salesforce®, Oracle®, etc.) and computing system applications (e.g., desktop and laptop computer applications, mobile device applications, wearable computing applications, etc.).
[0033] Coordinator 120 may have various capabilities, including but not limited to provisioning, deployment, version control, configuration, queuing, monitoring, logging, and / or providing interconnectivity. Provisioning may include creating and maintaining a connection between robot 130 and coordinator 120 (e.g., a web application). Deployment may include ensuring that packaged versions are correctly delivered to assigned robots 130 for execution. In some embodiments, version control may include the management of unique instances of some processes or configurations. Configuration may include the maintenance and delivery of robot environment and process configurations. Queuing may include providing management of queues and queue items. Monitoring may include keeping track of robot identification data and maintaining user permissions. Logging may include storing logs in a database (e.g., an SQL database) and / or other storage mechanisms (e.g., ElasticSearch®) and indexing them, which provides the ability to store and quickly query large datasets. Coordinator 120 may provide interconnectivity by acting as a centralized communication point for third-party solutions and / or applications.
[0034] Robot 130 is an execution agent that runs workflows built into Designer 110. A commercial example of some embodiments of Robot 130 is UiPath Robots™. In some embodiments, Robot 130 has services managed by Microsoft Windows® Service Control Manager (SCM) installed by default. As a result, such Robot 130 can open interactive Windows® sessions under the local system account and has rights to Windows® services.
[0035] In some embodiments, the robot 130 can be installed in user mode. For such robots 130, this means they have the same rights as a user who has already installed a given robot 130. This feature can also be used for high-density (HD) robots, ensuring full utilization of the maximum potential of each machine. In some embodiments, any type of robot 130 can be configured in an HD environment.
[0036] In some embodiments, robot 130 is broken down into several components, each dedicated to a specific automation task. Robot components in some embodiments include, but are not limited to, SCM-managed robot services, user-mode robot services, actuators, agents, and command lines. SCM-managed robot services manage and monitor Windows® sessions and act as agents between coordinator 120 and the execution host (i.e., the computing system on which robot 130 executes). These services are trusted and manage credentials for robot 130. The console application is launched by the SCM on the local system.
[0037] In some embodiments, the user-mode robot service manages and monitors Windows® sessions and acts as an agent between the coordinator 120 and the execution host. The user-mode robot service can be trusted and manage credentials for robot 130. If the SCM-managed robot service is not installed, Windows® applications can be launched automatically.
[0038] Actors can run a given job (i.e., they can execute workflows) within a Windows® session. Actors can be aware of the dots per inch (DPI) setting for each monitor. Agents can be Windows® Presentation Foundation (WPF) applications that display available jobs in the system tray window. Agents can be clients of a service. Agents can request to start or stop a job and change settings. The command line is a client of a service. The command line is a console application that can request to start a job and wait for its output.
[0039] Decomposing the components of robot 130 as explained above helps developers, support users, and the computing system more easily run, identify, and track what each component is performing. This allows for configuring specific behaviors for each component, such as setting different firewall rules for executors and services. In some embodiments, the executor may always know the DPI setting for each monitor. As a result, workflows can be executed at any DPI, regardless of the configuration of the computing system on which they are created. In some embodiments, projects from designer 110 can also be independent of browser scaling levels. For applications that do not know the DPI or are intentionally marked as not knowing it, DPI can be disabled in some embodiments.
[0040] Figure 2 This is an architectural diagram illustrating a deployed RPA system 200 according to an embodiment of the present invention. In some embodiments, the RPA system 200 may be... Figure 1 The RPA system 100 may be an integral part thereof. It should be noted that the client side, server side, or both may include any desired number of computing systems without departing from the scope of the invention. On the client side, the robot application 210 includes an actuator 212, an agent 214, and a designer 216. However, in some embodiments, the designer 216 may not run on the computing system 210. The actuator 212 operates. Several business processes can run simultaneously, such as... Figure 2 As shown in the example. In this embodiment, agent 214 (e.g., Windows) ®The service is a single point of contact for all executors 212. In this embodiment, all messages are logged to the coordinator 230, which further processes them via the database server 240, the indexer server 250, or both. (See above regarding...) Figure 1 The actuator 212 discussed may be a robot component.
[0041] In some embodiments, a robot represents an association between a machine name and a username. A robot can manage multiple actuators simultaneously. On a computing system that supports multiple interactive sessions running concurrently (e.g., Windows® Server 2012), multiple robots can run simultaneously, each robot using a unique username in a separate Windows® session. This is referred to above as an HD robot.
[0042] Agent 214 is also responsible for sending the robot's status (e.g., periodically sending "heartbeat" messages indicating that the robot is still operational) and downloading the required version of the package to be executed. In some embodiments, communication between agent 214 and coordinator 230 is always initiated by agent 214. In notification scenarios, agent 214 may open a WebSocket channel, which is later used by coordinator 230 to send commands to the robot (e.g., start, stop, etc.).
[0043] On the server side, a presentation layer (web application 232, Open Data Protocol (OData) Representative State Delivery (REST) Application Programming Interface (API) endpoint 234, and notification and monitoring 236), a service layer (API implementation / business logic 238), and a persistence layer (database server 240 and indexer server 250) are included. The coordinator 230 includes the web application 232, the OData REST API endpoint 234, the notification and monitoring 236, and the API implementation / business logic 238. In some embodiments, most actions performed by the user in the interface of the coordinator 230 (e.g., via browser 220) are performed by invoking various APIs. Such actions may include, but are not limited to, starting a job on a robot, adding / removing data from a queue, scheduling a job to run unattended, etc., without departing from the scope of the invention. The web application 232 is the visual layer of the server platform. In this embodiment, the web application 232 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of the invention. In this embodiment, the user interacts with a webpage from web application 232 via browser 220 to perform various actions to control coordinator 230. For example, the user can create robot groups, assign packages to robots, analyze logs of each robot and / or each process, start and stop robots, etc.
[0044] In addition to web application 232, coordinator 230 also includes a service layer that exposes an OData REST API endpoint 234. However, other endpoints may be included without departing from the scope of the invention. The REST API is used by web application 232 and agent 214. In this embodiment, agent 214 is a manager of one or more robots on a client computer.
[0045] The REST API in this embodiment covers configuration, logging, monitoring, and queuing functionality. In some embodiments, the configuration endpoint can be used to qualify and configure application users, permissions, bots, assets, publications, and environments. The logging REST endpoint can be used to log various information, such as errors, explicit messages sent by bots, and other environment-specific information. Bots can use the deployment REST endpoint to query the wrapper version that should be executed if a start job command is used in coordinator 230. The queuing REST endpoint can be responsible for queue and queue item management, such as adding data to the queue, retrieving transactions from the queue, and setting the status of transactions.
[0046] The monitoring REST endpoint can monitor web application 232 and agent 214. The notification and monitoring API 236 can be a REST endpoint used to register agent 214, deliver configuration settings to agent 214, and send / receive notifications from the server and agent 214. In some embodiments, the notification and monitoring API 236 can also use WebSocket communication.
[0047] In this embodiment, the persistence layer includes a pair of servers—a database server 240 (e.g., an SQL server) and an indexer server 250. In this embodiment, the database server 240 stores configurations for robots, robot groups, associated processes, users, roles, schedules, etc. In some embodiments, this information is managed via a web application 232. The database server 240 can manage queues and queue items. In some embodiments, the database server 240 can store messages logged by the robot logs (in addition to or instead of the indexer server 250).
[0048] In some embodiments, the optional indexer server 250 stores and indexes information logged by the robot logs. In some embodiments, the indexer server 250 can be disabled through configuration settings. In some embodiments, the indexer server 250 uses ElasticSearch®, an open-source full-text search engine project. Messages logged by the robot logs (e.g., using activities such as log messages or write lines) can be sent to the indexer server 250 via logging REST endpoints(s), where they are indexed for future use.
[0049] Figure 3This is an architecture diagram illustrating the relationship 300 between a designer 310, activities 320, 330, a driver 340, and an AI / ML model 350 according to an embodiment of the present invention. As described above, a developer uses the designer 310 to develop a workflow executed by a robot. The workflow may include user-defined activities 320 and UI automation activities 330. In some embodiments, user-defined activities 320 and / or UI automation activities 330 may invoke one or more AI / ML models 350, which may be locally located on a computing system on which the robot is operating and / or remotely operating. Some embodiments are capable of identifying non-textual visual components in an image, referred to herein as computer vision (CV). Some CV activities associated with such components may include, but are not limited to, clicking, typing, text capture, hovering, element presence, refreshing range, highlighting, etc. In some embodiments, clicking uses, for example, CV, optical character recognition (OCR), fuzzy text matching, and multi-anchor points to identify an element and click it. Typing can use the above to identify an element and type within it. Text capture can identify the location of specific text and scan it using OCR. Hovering can identify an element and hover over it. The presence of elements can be checked using the techniques described above. In some embodiments, hundreds or even thousands of activities may be implemented in the designer 310. However, any number and / or type of activities may be available without departing from the scope of the invention.
[0050] UI automation activities 330 are a specific subset of lower-level activities written in low-level code (e.g., CV activities) and facilitating interaction with the screen. UI automation activities 330 facilitate these interactions via drivers 340 and / or AI / ML models 350 that allow the robot to interact with desired software. For example, drivers 340 may include OS drivers 342, browser drivers 344, VM drivers 346, enterprise application drivers 348, etc. UI automation activities 330 may use one or more AI / ML models 350 to determine the execution of interactions with the computing system. In some embodiments, AI / ML models 350 may enhance or completely replace drivers 340. In fact, in some embodiments, drivers 340 are not included.
[0051] Drivers 340 can interact with the OS at a low level to find hooks, monitor key presses, etc. They can facilitate integration with Chrome®, IE®, Citrix®, SAP®, and others. For example, a "click" activity can perform the same role in these different applications via driver 340.
[0052] Figure 4This is an architectural diagram illustrating an RPA system 400 according to an embodiment of the present invention. In some embodiments, the RPA system 400 may be or include Figure 1 and Figure 2 The RPA system 100 and / or 200. The RPA system 400 includes multiple client computing systems 410 that run robots. The computing systems 410 are able to communicate with the coordinator computing system 420 via a web application running thereon. The coordinator computing system 420 is then able to communicate with a database server 430 and an optional indexer server 440.
[0053] about Figure 1 and Figure 3 It should be noted that although web applications are used in these embodiments, any suitable client / server software can be used without departing from the scope of the invention. For example, the coordinator can run a server-side application that communicates with non-web-based client software applications on the client computing system.
[0054] Figure 5 This is an architectural diagram illustrating a computing system 500 according to an embodiment of the present invention, configured to train AI / ML models to recognize application, screen, and UI elements using computer vision (CV), and to recognize user interactions with these elements. In some embodiments, the computing system 500 may be one or more of the computing systems depicted and / or described herein. The computing system 500 includes a bus 505 or other communication mechanism for transmitting information, and multiple processors 510 coupled to the bus 505 for processing information. The multiple processors 510 may be any type of general-purpose or special-purpose processor, including a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The multiple processors 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multi-parallel processing may be used. In some embodiments, at least one of the multiple processors 510 may be a neuromorphic circuit, which includes processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require typical components of a von Neumann computing architecture.
[0055] The computing system 500 also includes a memory 515 for storing information and instructions to be executed by the processor(s) 510. The memory 515 may include any combination of random access memory (RAM), read-only memory (ROM), flash memory, cache, static storage such as a disk or optical disk, or any other type of non-transitory computer-readable medium or a combination thereof. The non-transitory computer-readable medium may be any available medium accessible to the processor(s) 510 and may include volatile media, non-volatile media, or both. The medium may also be removable, non-removable, or both.
[0056] Additionally, the computing system 500 includes communication devices 520, such as transceivers, to provide access to a communication network via wireless and / or wired connections. In some embodiments, the communication device 520 may be configured to use Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), and High-Speed Uplink Packet Access (HSDPA). SUPA, High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE-Advanced (LTE-A), 802.11x, Wi-Fi, Zigbee, Ultra Wideband (UWB), 802.16x, 802.15, Home Node B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near Field Communication (NFC), 5G, New Radio (NR), any combination thereof, and / or any other currently existing or future-implemented communication standards and / or protocols without departing from the scope of this invention. In some embodiments, the communication device 520 may include one or more antennas, which may be single, arrayed, phased, switched, beamforming, beamcontrolled, combinations thereof, and / or any other antenna configuration without departing from the scope of this invention.
[0057] The processors 510 are also coupled via bus 505 to a display 525, such as a plasma display, liquid crystal display (LCD), light-emitting diode (LED) display, field emission display (FED), organic light-emitting diode (OLED) display, flexible OLED display, flexible substrate display, projection display, 4K display, high-definition display, Retina® display, in-plane switching (IPS) display, or any other suitable display for displaying information to a user. The display 525 can be configured as a touch (haptic) display, 3D touch display, multi-input touch display, multi-touch display, etc., using resistive, capacitive, surface acoustic wave (SAW) capacitive, infrared, optical imaging, dispersive signal technology, acoustic pulse recognition, suppressed total internal reflection, etc., employing resistive, capacitive, surface acoustic wave (SAW) capacitive, infrared, optical imaging, dispersive signal technology, acoustic pulse recognition, suppressed total internal reflection, etc. Any suitable display device and haptic input / output can be used without departing from the scope of the invention.
[0058] Keyboard 530 and cursor control devices 535, such as a computer mouse, touchpad, etc., are also coupled to bus 505 to enable the user to interact with computing system 500. However, in some embodiments, a physical keyboard and mouse may not be present, and the user may interact with the device solely through display 525 and / or touchpad (not shown). As a design choice, any type and combination of input devices can be used. In some embodiments, no physical input devices and / or display exist. For example, the user may interact remotely with computing system 500 via another computing system with which it communicates, or computing system 500 may operate autonomously.
[0059] Memory 515 stores software modules that provide functionality when executed by processor(s) 510. The modules include an operating system 540 for computing system 500. The modules also include an AI / ML model training module 545 configured to perform all or part of the processes described herein or derivatives thereof. Computing system 500 may include one or more additional functional modules 550, which include additional functionality.
[0060] Those skilled in the art will understand that "system" can be embodied as a server, embedded computing system, personal computer, console, personal digital assistant (PDA), mobile phone, tablet computing device, quantum computing system, or any other suitable computing device or combination of devices without departing from the scope of the invention. Presenting the foregoing functionality as being performed by a "system" is not intended to limit the scope of the invention in any way, but rather to provide an example of one of many embodiments of the invention. In fact, the methods, systems, and apparatuses disclosed herein can be implemented in localized and distributed forms consistent with computing technologies including cloud computing systems. The computing system can be part of or otherwise accessible to a local area network (LAN), mobile communication network, satellite communication network, the Internet, public or private cloud, hybrid cloud, server farm, any combination thereof, etc. Any localized or distributed architecture can be used without departing from the scope of the invention.
[0061] It should be noted that some system features described in this specification have been presented as modules to more specifically emphasize their implementation independence. For example, modules can be implemented as hardware circuits, including custom-designed very large-scale integrated circuits (VLSI) or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Modules can also be implemented in programmable hardware devices such as field-programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, and so on.
[0062] Modules can also be implemented, at least partially, in software for execution by various types of processors. Identified executable code units may, for example, comprise physical or logical blocks of one or more computer instructions, which may be organized, for example, as objects, procedures, or functions. However, executable files of identified modules do not need to be physically located together, but may include different instructions stored in different locations that, when logically connected, constitute the module and achieve its intended purpose. Furthermore, modules can be stored on computer-readable media, which may be, for example, hard disk drives, flash memory devices, RAM, magnetic tape, and / or any other non-transitory computer-readable media used for storing data without departing from the scope of the invention.
[0063] In practice, a module of executable code can be a single instruction or many instructions, and can even be distributed across several different code segments, different programs, and several memory devices. Similarly, operational data can be identified and described within this module, and can be represented in any suitable form and organized within any suitable type of data structure. Operational data can be collected as a single dataset, or it can be distributed across different locations, including different storage devices, and can exist at least in part as electronic signals on a system or network.
[0064] Figure 6 This is an architectural diagram illustrating a system 600 according to an embodiment of the present invention, configured to train an AI / ML model to recognize application, screen, and UI elements using computer vision (CV) and to identify user interactions with these elements. System 600 includes user computing systems such as a desktop computer 602, a tablet computer 604, and a smartphone 606. However, any desired computing system, including but not limited to smartwatches, laptops, etc., can be used without departing from the scope of the invention. In some embodiments, one or more of computing systems 602, 604, and 606 may include an automation box and / or a camera. Furthermore, although... Figure 6 Three user computing systems are shown, but any suitable number of computing systems can be used without departing from the scope of the invention. For example, in some embodiments, dozens, hundreds, thousands, or millions of computing systems may be used.
[0065] Each computing system 602, 604, 606 has a recorder process 610 (i.e., a tracking application) running thereon, which records screenshots and / or video of the user's screen or a portion thereof. For example, a piece of JavaScript® can be embedded in a web browser as a recorder process 610 to track which components the user interacts with, what text the user types, which locations / components the user clicks with the mouse, what content the user scrolls over, how long the user lingers on a particular part of the content, etc. Scrolling over content may indicate that the content is somewhat close but not exactly what the user wants. Clicking may indicate success.
[0066] The recorder process 610 need not be JavaScript® and can be any suitable type of application and any desired programming language without departing from the scope of the invention. This allows the recorder processes 610 to be “generalized” so that they can track user interactions with multiple applications or any application with which the user is interacting. Using labeled training data from scratch can be difficult because while it may allow AI / ML models to learn to recognize various controls, it does not contain information about which controls are frequently used and how they are used. Using the recorder process 610, “heatmaps” can be generated to help guide the AI / ML model training process. Heatmaps can include various information such as the frequency with which the user uses the application, the frequency with which the user interacts with components of the application, the location of components, the content of the application / component, etc. In some embodiments, heatmaps can be derived from screen analysis, such as the detection of typed and / or pasted text, caret tracking, and active element detection by computing systems 602, 604, 606. Some embodiments identify where a user types or pastes text on a screen associated with computing systems 602, 604, 606, potentially including hotkeys or other keys that do not result in visible characters, and provide the physical location on the screen based on the location of one or more characters, the location of the cursor blinking, or both (e.g., in coordinates) at the current resolution. The physical location of typing or pasting activity and / or caret can allow determination of which field(s) the user is typing or focusing on, and what application it is for process discovery or other applications.
[0067] As described above, in some embodiments, the recorder process 610 may record additional data to further assist in training the AI / ML model(s), such as web browser history, heatmaps, key presses, mouse clicks, mouse click locations and / or graphical elements on the screen that the user is interacting with, the location the user is viewing on the screen at different times, timestamps associated with screenshots / video frames, etc. This may be beneficial for providing key presses and / or other user actions that may not cause screen changes. For example, some applications may not provide visual changes when a user presses CTRL+S to save a file. However, in some embodiments, the AI / ML model(s) may be trained solely based on the captured screen images. The recorder process 610 may be a robot generated via an RPA designer application, part of an operating system, a downloadable application for a personal computer (PC) or smartphone, or any other software and / or hardware without departing from the scope of the invention. In fact, in some embodiments, the logic of one or more recorder processes 610 is implemented partially or entirely through physical hardware.
[0068] Some embodiments are implemented in a feedback loop process that continuously or periodically compares the current screenshot with previous screenshots to identify changes. Locations of visual changes on the screen can be identified, and OCR can be performed on the changed locations. The results of the OCR can then be compared with the contents of a keyboard queue (e.g., determined by key hooks) to determine if a match exists. The changed location can be determined by comparing pixel boxes from the current screenshot with pixel boxes at the same location from a previous screenshot.
[0069] Images and / or other data recorded by the recorder process 610 (e.g., web browser history, heatmaps, key presses, mouse clicks, mouse click locations and / or graphical elements on the screen that the user is interacting with, locations the user is viewing on the screen at different times, timestamps associated with screenshots / video frames, voice input, gestures, emotions (e.g., the user is happy, frustrated, etc.), biometrics (e.g., fingerprints, retinal scans, the user's pulse, etc.), information related to periods of no user activity (e.g., "dead man switch"), haptic information from a haptic display or touchpad, heatmaps with multi-touch input, etc.) are transmitted to server 630 via network 620 (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). In some embodiments, server 630 may be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In some embodiments, server 630 may host multiple software-based servers on a single computing system 630. In some embodiments, server 630 may run a coordinator application and data from the recorder process 610 may be periodically sent as part of a heartbeat message. In some embodiments, once a predetermined amount of data has been collected, after a predetermined time period has elapsed, or both, data can be sent from the recorder process 610 to the server 630. The server 630 stores the data received from the recorder process 610 in a database 640.
[0070] In this embodiment, server 630 includes multiple AI layers 632 that collectively form an AI / ML model. However, in some embodiments, the AI / ML model may have only one layer. In some embodiments, multiple AI / ML models may be trained on server 630 and used together to accomplish a larger task. AI layer 632 may employ computer vision (CV) techniques and may perform various functions such as statistical modeling (e.g., Hidden Markov Models (HMMs)) and utilize deep learning techniques (e.g., Long Short-Term Memory (LSTM) deep learning, encoding of previously hidden states, etc.) to identify user interactions. Initially, the AI / ML model needs to be trained so that it can perform meaningful analysis on the data captured in database 640. In some embodiments, users of computing systems 602, 604, 606 tag images before they are sent to server 630. Additionally or alternatively, in some embodiments, tagging processing subsequently occurs, such as via an application 652 running on computing system 650, which allows users to draw bounding boxes and / or other shapes around graphical elements, thereby providing text labels for the content contained within the bounding boxes, etc.
[0071] The AI / ML model undergoes a training phase using this data as input and is trained until it is sufficiently accurate without overfitting the training data. Acceptable accuracy can vary depending on the application. Labeling errors can be highlighted by a human reviewer (e.g., by drawing a box around the incorrectly labeled element and including the correct label), and this additional labeled data can be used to retrain the AI / ML model. Once fully trained, the AI / ML model is able to provide predictions as output, such as which application(s) and graphical elements are identified as being presented on the screen.
[0072] However, while this level of training provides information about what is being presented, further information may be needed to determine user interactions, such as comparing two or more consecutive screens to determine when typed characters appear from one screen to another, buttons are pressed, menu selections appear, etc. Therefore, after the AI / ML model can recognize graphical elements and applications on the screen, in some embodiments, the AI / ML model is further trained to recognize labeled user interactions with applications in the UI to understand such incremental actions taken by the user. Errors can be highlighted again by a human reviewer (e.g., by drawing a box around the incorrectly labeled element and including the correct label), and the AI / ML model can be trained until its accuracy is high enough to be deployed to understand fine-grained user interactions with the UI.
[0073] Once trained to recognize user interactions, the trained AI / ML model can be used to analyze video and / or other information from the recorder process 610. This recorded information may include multiple / many interactions that users tend to perform. These interactions can then be analyzed against common sequences for use in subsequent automation.
[0074] AI layer
[0075] In some embodiments, multiple AI layers may be used. Each AI layer is an algorithm (or model) that runs on data, and the AI model itself may be a deep learning neural network (DLNN) with trained artificial “neurons” trained on training data. Layers may run in series, in parallel, or in a combination thereof.
[0076] AI layers can include, but are not limited to, sequence extraction layers, clustering detection layers, visual component detection layers, text recognition layers (e.g., OCR), audio-to-text conversion layers, or any combination thereof. However, any desired number and type of layers can be used without departing from the scope of this invention. Using multiple layers allows the system to develop a global view of what is happening on the screen. For example, one AI layer can perform OCR, while another can detect buttons, etc.
[0077] Patterns can be determined individually by an AI layer or jointly by multiple AI layers. They can be used based on the probability of a user's action or the output. For example, to determine details such as a button's location, its text, and the user's click position, the system might need to know the button's position, its text, and its location on the screen.
[0078] However, it should be noted that various AI / ML models can be used without departing from the scope of the invention. While in some embodiments, neural networks such as DLNNs, recurrent neural networks (RNNs), generative adversarial networks (GANs), any combination thereof, etc., can be used to train AI / ML models, other AI techniques, such as deterministic models, shallow learning neural networks (SLNNs), or any other suitable AI / ML model type and training technique, may also be used without departing from the scope of the invention.
[0079] Figure 7This is a flowchart illustrating a process 700 for training an AI / ML model to recognize application, screen, and UI elements using computer vision (CV) according to an embodiment of the present invention, and for recognizing user interactions with application, screen, and UI elements. At 710, the process begins by recording screenshots or video frames of a display associated with the user's computing system, as well as other information. In some embodiments, the recording is performed by one or more recorder processes. In some embodiments, the recorder process is implemented as a feedback loop process that continuously or periodically compares the current screenshot or video frame with previous screenshots or video frames and identifies one or more locations where changes have occurred between the current screenshot or video frame and the previous screenshot or video frame. In some embodiments, the recorder process is configured to perform OCR on the one or more locations where changes have occurred, compare the results of the OCR with the contents of a keyboard queue to determine if a match exists, and when a match exists, link the text associated with the match to the corresponding location. In some embodiments, the other information includes web browser history, one or more heatmaps, key presses, mouse clicks, mouse click locations, and / or graphical elements on the display that the user is interacting with, the location the user is viewing on the display, timestamps associated with screenshots or video frames, text entered by the user, content the user scrolls over, the time the user lingers on portions of content displayed on the display, what application the user is interacting with, or a combination thereof. In some embodiments, at least some of the other information is captured using one or more automated boxes.
[0080] At 720, one or more heatmaps are generated as part of other information. In some embodiments, the one or more heatmaps include the frequency of user use of the application, the frequency of user interaction with components of the application, the location of components within the application, the content of the application and / or components, or a combination thereof. In some embodiments, one or more heatmaps are derived from display analysis, which includes detection of typed and / or pasted text, caret tracking, active element detection, or a combination thereof. Recorded screenshots or video frames, along with other information, are then sent at 730 to one or more server-accessible storage devices.
[0081] At 740, the recorded screenshots or video frames, along with other information, are accessed (e.g., via a server configured to train an AI / ML model). At 750, the recorded screenshots or video frames, along with the other information, are used to initially train the AI / ML model to identify applications, screens, and UI elements presented in the recorded screenshots or video frames. In some embodiments, the initial training of the AI / ML model is performed without prior knowledge of the applications, screens, and UI elements in the screenshots or video frames.
[0082] After the AI / ML model can reliably (e.g., 70%, 95%, 99.99%, etc.) identify application, screen, and UI elements in recorded screenshots or video frames, at 760, the AI / ML model is trained to recognize individual user interactions with UI elements. In some embodiments, individual user interactions include button presses, input of a single character or sequence of characters, selection of an active UI element, menu selection, screen changes, or combinations thereof. In some embodiments, training the AI / ML model to recognize individual user interactions with UI elements includes comparing two or more consecutive screenshots or video frames and determining whether a typed character appears from one screenshot to another, a button is pressed, or a menu selection appears. At 770, the AI / ML model is then deployed so that it can be invoked and used by a process (e.g., an RPA bot).
[0083] Figure 8 This is an architectural diagram illustrating an automated frame and eye motion tracking system 800 according to an embodiment of the present invention. System 800 includes a computing system 810, which includes eye tracking logic (ETL) 812 configured to process input from a camera 820 and automated frame logic (ABL) 814 configured to process input from an automated frame 860. In some embodiments, computing system 810 may be or include... Figure 5 The computing system 500. In some embodiments, multiple cameras may be used.
[0084] When a user interacts with the computing system 810 via mouse 840 and keyboard 850, camera 820 records video of the user. The computing system 810 converts the recorded camera video into video frames. ETL processes these frames to identify the user's eyes and interpolates the location the user is looking at to a specific location on display 830. Any suitable eye-tracking technology(s) can be used without departing from the scope of the invention, such as those described in U.S. Patent Application Publication No. 2018 / 0046248, U.S. Patent No. 7,682,026, etc. Timestamps can be associated with the user's video frames so that they can be matched with screenshot frames displayed on display 830 at that time.
[0085] In this embodiment, the automation box 860 also includes automation box logic 862, which receives input from the mouse 840 and the keyboard 850. In some embodiments, the automation box 860 may have hardware similar to that of the computing system 810 (e.g., processors, memory, buses, etc.). This input can then be passed to the computing system 810. Although in Figure 8A mouse 840 and a keyboard 850 are shown, but any suitable input device(s), such as a touchpad, button, etc., can be used without departing from the scope of the invention. In some embodiments, only the computing system 810 or the automation box 860 includes automation box logic. The latter may be for recording user interactions and sending them directly to a server (e.g., a cloud-based server) via network 870 for subsequent processing. In such embodiments, screenshot frames may also be sent from the computing system 810 to the automation box 860 and then to the server via network 870. Alternatively, the computing system 810 may send the screenshot itself via network 870. Such embodiments can provide a plug-and-play tracking solution that can be plugged into the computing system 810 to relay keyboard and mouse information to the computing system 810 for its operation and also to relay keyboard and mouse click information to a remote server for subsequent training of AI / ML models.
[0086] In some embodiments, the automation box 860 may include actuation logic that runs automation and simulates input. This allows the automation box 860 to provide the computing system 810 with simulated key presses, mouse movements, and clicks, as if this information actually came from a human user interacting with these components. UI screenshots and other information can then be used to train an AI / ML model. Another advantage of such an embodiment is that the AI / ML model can be trained when the user leaves the computing system 810, potentially allowing for the capture of a larger amount of training information more quickly, and therefore potentially allowing for faster training of the AI / ML model.
[0087] In some embodiments, the "information box" can be implemented as software on the computing system 610 and can be similar to Figure 6 The recorder process 610 operates in a manner that allows it to store screenshot frames, mouse click information, and key press information. In some embodiments, eye-tracking information can also be tracked. This information can then be sent to a server via network 870, and gaze tracking can potentially be performed remotely rather than on computing system 810.
[0088] According to an embodiment of the present invention, Figure 7 The process steps executed in the process can be performed by a computer program, which encodes instructions for use by (multiple) processors to execute. Figure 7 The computer program may be embodied on a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, hard disk drives, flash memory devices, RAM, magnetic tape, and / or any other such medium or combination of media used to store data. The computer program may include processor(s) for controlling a computing system (e.g., ...). Figure 5The computing system 500 has (multiple) processors 510 to achieve Figure 7 The coded instructions for all or part of the process steps described herein may also be stored on a computer-readable medium.
[0089] Computer programs can be implemented in hardware, software, or a hybrid approach. A computer program can consist of modules that can communicate and operate with each other, and it is designed to deliver information or instructions for display. A computer program can be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0090] It will be readily understood that the components of the various embodiments of the invention as generally described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the invention as illustrated in the drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention.
[0091] The features, structures, or characteristics of the invention described throughout this specification can be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments," "some embodiments," or similar language mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the presence of phrases such as "in some embodiments," "in some embodiments," "in other embodiments," or similar language throughout this specification does not necessarily refer to the same set of embodiments, and the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0092] It should be noted that references to features, advantages, or similar language throughout this specification do not imply that all features and advantages achievable with the invention should be or are present in any single embodiment of the invention. Rather, references to features and advantages should be understood as indicating that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Therefore, the discussion of features and advantages throughout this specification, as well as similar language, may, but do not necessarily, refer to the same embodiments.
[0093] Furthermore, the features, advantages, and characteristics of the invention described herein can be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize that the invention can be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of the invention may be recognized in certain embodiments.
[0094] It will be readily understood by those skilled in the art that the invention discussed above can be practiced with steps in a different order and / or with hardware elements in a different configuration than those disclosed. Therefore, although the invention has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain modifications, variations, and alternative constructions will be readily apparent while remaining within the spirit and scope of the invention. Therefore, reference should be made to the appended claims to determine the limits and scope of the invention.
Claims
1. A system comprising: One or more user computing systems, the one or more user computing systems including corresponding recorder processes; as well as The server is configured to train an artificial intelligence (AI) / machine learning (ML) model to use computer vision (CV) to identify application, screen, and user interface (UI) elements, and to identify user interactions with said application, screen, and UI elements. The corresponding recorder process is configured as follows: Record screenshots or video frames of the monitor associated with the corresponding user's computing system, as well as other information. The recorded screenshots or video frames, along with the other information, are sent to a storage device accessible by the server. The server is configured as follows: The AI / ML model is initially trained using the recorded screenshots or video frames and the other information to identify the application, screen, and UI elements presented in the recorded screenshots or video frames. After the AI / ML model is able to identify the application, screen, and UI elements in the recorded screenshot or video frame with a certain confidence, the AI / ML model is trained to identify individual user interactions with the UI elements.
2. The system according to claim 1, wherein the individual user interaction includes button pressing, input of a single character or character sequence, selection of active UI elements, menu selection, screen changes, voice input, gestures, provision of biometric information, tactile interaction, or a combination thereof.
3. The system of claim 1, wherein training the AI / ML model to identify the individual user interactions with the UI elements comprises: Compare two or more consecutive screenshots or video frames, and determine when the typed characters appear from one screenshot to another, when a button is pressed, or when a menu selection appears.
4. The system of claim 1, wherein the other information includes web browser history, one or more heatmaps, key presses, mouse clicks, mouse click locations and / or graphical elements on the display with which the user is interacting, the location the user is viewing on the display, timestamps associated with the screenshots or video frames, text entered by the user, content the user scrolls over, the time the user remains on portions of content displayed on the display, the application the user is interacting with, voice input, gestures, emotional information, biometrics, information related to periods of no user activity, tactile information, multi-touch input information, or combinations thereof.
5. The system according to claim 1, wherein The one or more user computing systems or the server are configured to generate one or more heatmaps, and the other information includes the one or more heatmaps. The one or more heatmaps include the frequency with which a user uses the application, the frequency with which the user interacts with components of the application, the location of the components in the application, the content of the application and / or components, or a combination thereof.
6. The system of claim 5, wherein the one or more user computing systems or the server are configured to derive the one or more heatmaps from display analysis, the display analysis including detection of typed and / or pasted text, caret tracking, active element detection, or a combination thereof.
7. The system according to claim 1, further comprising: An automation box, operably connected to one or more user computing systems, is configured to: Receive input from one or more user input devices. Associate the timestamp with the input, and The timestamped input is sent to a storage device accessible by the server, wherein The server is configured to use the timestamped input for the initial training of the AI / ML model.
8. The system of claim 1, wherein the server is configured to perform the initial training of the AI / ML model without prior knowledge of the application, screen, and UI elements in the screenshot or video frame.
9. A non-transitory computer-readable medium storing a computer program configured to train an artificial intelligence (AI) / machine learning (ML) model to use computer vision (CV) to recognize application, screen, and user interface (UI) elements and / or recognize user interactions with said application, screen, and UI elements, said computer program being configured to cause at least one processor to: Access to recorded screenshots or video frames of displays associated with one or more computing systems, and access to other information associated with said one or more computing systems; and The AI / ML model is initially trained using the recorded screenshots or video frames and the other information to identify the application, screen, and UI elements presented in the recorded screenshots or video frames, wherein... The initial training of the AI / ML model is performed without prior knowledge of the application, screen, and UI elements in the screenshots or video frames.
10. The non-transitory computer-readable medium of claim 9, wherein after the AI / ML model is able to identify the application, screen, and UI elements in the recorded screenshot or video frame with a certain confidence level, the computer program is further configured to cause the at least one processor to: The AI / ML model is trained to identify individual user interactions with the UI elements.
11. The non-transitory computer-readable medium of claim 10, wherein training the AI / ML model to identify the individual user interactions with the UI elements comprises: Compare two or more consecutive screenshots or video frames, and determine when the typed characters appear from one screenshot to another, when a button is pressed, or when a menu selection appears.
12. The non-transitory computer-readable medium of claim 10, wherein the individual user interaction includes button pressing, input of a single character or sequence of characters, selection of an active UI element, menu selection, screen changing, voice input, gesture, provision of biometric information, tactile interaction, or a combination thereof.
13. The non-transitory computer-readable medium of claim 9, wherein the other information includes web browser history, one or more heatmaps, key presses, mouse clicks, mouse click locations and / or graphical elements on the display with which the user is interacting, locations viewed by the user on the display, timestamps associated with the screenshots or video frames, text entered by the user, content scrolled by the user, time the user remained on portions of content shown on the display, what application the user is interacting with, voice input, gestures, emotional information, biometrics, information related to periods of no user activity, tactile information, multi-touch input information, or combinations thereof.
14. The non-transitory computer-readable medium of claim 9, wherein the computer program is further configured to cause the at least one processor to: Generate one or more heatmaps, wherein the other information includes the one or more heatmaps, wherein The one or more heatmaps include the frequency with which a user uses one or more applications, the frequency with which the user interacts with components of the one or more applications, the location of the components in the one or more applications, the content of the one or more applications and / or components, or a combination thereof.
15. The non-transitory computer-readable medium of claim 14, wherein the one or more heatmaps are derived from display analysis, the display analysis including detection of typed and / or pasted text, caret tracking, active element detection, or a combination thereof.
16. A computer-implemented method for training an artificial intelligence (AI) / machine learning (ML) model to recognize applications, screens, and user interface (UI) elements using computer vision (CV), and to recognize user interactions with said applications, screens, and UI elements, said method comprising: Access to recorded screenshots or video frames of displays associated with one or more computing systems, and access to other information associated with said one or more computing systems; The AI / ML model is initially trained using the recorded screenshots or video frames and the other information to identify the applications, screens, and UI elements presented in the recorded screenshots or video frames. as well as After the AI / ML model is able to reliably identify the application, screen, and UI elements in the recorded screenshots or video frames, the AI / ML model is trained to identify individual user interactions with the UI elements.
17. The computer-implemented method of claim 16, wherein the initial training of the AI / ML model is performed without prior knowledge of the application, screen, and UI elements in the screenshot or video frame.
18. The computer-implemented method of claim 16, wherein training the AI / ML model to recognize the individual user interactions with the UI elements comprises: Compare two or more consecutive screenshots or video frames and determine when typed characters appear from one screenshot to another, when a button is pressed, or when a menu selection appears.
19. The computer-implemented method of claim 16, wherein the individual user interaction includes button pressing, input of a single character or sequence of characters, selection of an active UI element, menu selection, screen changing, voice input, gesture, provision of biometric information, tactile interaction, or a combination thereof.
20. The computer-implemented method of claim 16, wherein the other information includes web browser history, one or more heatmaps, key presses, mouse clicks, mouse click locations and / or graphical elements on the display with which the user is interacting, locations viewed by the user on the display, timestamps associated with the screenshots or video frames, text entered by the user, content scrolled by the user, time the user remained on portions of content displayed on the display, applications the user is interacting with, voice input, gestures, emotional information, biometrics, information related to periods of no user activity, tactile information, multi-touch input information, or combinations thereof.
21. The computer-implemented method according to claim 16, further comprising: Generate one or more heatmaps, wherein the other information includes the one or more heatmaps, wherein The one or more heatmaps include the frequency with which a user uses one or more applications, the frequency with which the user interacts with components of the one or more applications, the location of the components in the one or more applications, the content of the one or more applications and / or components, or a combination thereof, and The one or more heatmaps are derived from display analysis, which includes detection of typed and / or pasted text, caret tracking, active element detection, or a combination thereof.
Citation Information
Patent Citations
Systems and methods of eye tracking control
US20180046248A1
Eye location and gaze detection system and method
US7682026B2
Systems and methods for providing interactive streaming media
CN108337909A
Misuse index for explainable artificial intelligence in computing environments
CN111626319A