Ui task automation with guided teaching

Through multimodal interfaces and interactive context guidance, non-technical users can generate and verify UI task automation programs, solving the problem of insufficient programming skills and improving the efficiency and accuracy of generating complex UI task automation programs.

CN121970026APending Publication Date: 2026-05-01INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2024-08-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Non-technical users lack programming skills and find it difficult to create complex UI task automation programs, especially in dealing with UI component selection and DOM data representation.

Method used

It provides a multimodal interface and interactive context guidance, generates UI task automation programs by receiving instructional demonstrations and recording actions or expressions for conditional execution, and presents visual program representations for verification.

Benefits of technology

It enables the rapid generation of complex UI task automation programs without requiring programming skills, improving processing speed and accuracy, and making it easier for non-technical users to understand and verify.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970026A_ABST
    Figure CN121970026A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide methods, systems, and computer program products for implementing user interface (UI) task automation. Disclosed embodiments include receiving an automation structure and inputs and outputs of the structure to create task automation, and providing a multi-modal interface to process one or more teaching presentations for task automation, where the teaching presentations identify automation processing parameters and operations for task automation. An interactive context guide is generated to record conditional execution of one or more actions or expressions based on a state of one or more UI elements of the teaching presentation. The disclosed embodiments include recording a teaching presentation based on conditional execution of one or more actions or expressions, synthesizing a UI task automation program for task automation from the teaching presentation, and presenting the UI task automation program for verification.
Need to check novelty before this filing date? Find Prior Art

Description

Automating UI Tasks Using Guided Learning Background Technology

[0001] This invention relates to digital data processing, and more particularly, to methods, systems, and computer program products for automating user interface (UI) tasks using guided instruction to generate UI task automation programs.

[0002] UI task automation is typically complex and usually requires programmers to build UI task automation programs, also known as digital labor, software robots, digital robots, or Robotic Process Automation (RPA) robots. Non-technical users (e.g., business users) often lack an understanding of the programming concepts required to create custom UI task automation programs or digital labor. For example, non-technical users may lack the technical skills or programming abilities to tackle technical challenges such as referencing UI elements using static and dynamic selectors and defining the Document Object Model (DOM) data representation of objects. Summary of the Invention

[0003] Embodiments of this disclosure provide methods, systems, and computer program products for automating user interface (UI) tasks using guided instruction to generate UI task automation programs.

[0004] According to one embodiment of this disclosure, a non-limiting computer-implemented method is provided. The method includes receiving an automation structure and inputs and outputs of the structure to create a given task automation; providing a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generating an interactive contextual guide to record conditional execution of one or more actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification.

[0005] According to one embodiment of this disclosure, a system is provided. The system includes one or more computer processors and a memory containing a program that performs operations when executed by the one or more computer processors. The operations include: receiving an automation structure and its inputs and outputs to create a given task automation; providing a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation. The operations also include: generating interactive context guidance to record conditional execution of one or more actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification.

[0006] According to one embodiment of this disclosure, a computer program product is provided. The computer program product includes a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations. The operations include: receiving an automation structure and its inputs and outputs to create a given task automation; providing a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation. The operations further include: generating an interactive contextual guide to conditionally execute one or more actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification. Attached Figure Description

[0007] Figure 1 is a block diagram of an example computer environment used in conjunction with one or more disclosed embodiments for implementing user interface (UI) task automation through guided instruction that utilizes UI task automation or digital labor.

[0008] Figure 2 is a block diagram of an example system for implementing UI task automation of one or more of the disclosed embodiments;

[0009] Figures 3A and 3B together provide a flowchart of example operations for implementing example methods of UI task automation of one or more disclosed embodiments;

[0010] Figure 4 is a block diagram illustrating an automation structure for an example defined operation to implement one or more of the disclosed embodiments of UI task automation;

[0011] Figures 5, 6, and 7 together provide a block diagram illustrating a logical automation structure for UI task automation of one or more of the disclosed embodiments;

[0012] Figure 8 is a block diagram illustrating an automation structure for implementing an example teaching operation of UI task automation in one or more disclosed embodiments; and

[0013] Figure 9 is a flowchart of a method for automating UI tasks in one or more of the disclosed embodiments. Detailed Implementation

[0014] According to one aspect of the disclosed embodiments, a non-limiting computer-implemented method is provided. The method includes: receiving an automation structure and its inputs and outputs to create a given task automation; providing a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generating an interactive contextual guide to record conditional execution of one or more actions or expressions based on the state of one or more UI elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification. This method enables the automation of complex user interface (UI) tasks with no-code guided instruction involving digital labor, allowing non-technical users to synthesize UI task automation programs. The method improves the processing speed for generating complex UI automation logic (e.g., within minutes), including the definition, instruction, and verification of the synthesized UI task automation program. The method can efficiently and effectively generate UI task automation programs without requiring technical or programming user skills.

[0015] According to one aspect of the disclosed embodiments, a system is provided, including one or more computer processors and a memory containing a program that performs operations when executed by the one or more computer processors. The operations include: receiving an automation structure and inputs and outputs of the structure to create a given task automation; providing a multimodal UI interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generating interactive context guidance to record conditional execution of one or more actions or expressions based on the state of one or more UI elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification. This system enables the automation of complex user interface (UI) tasks with no-code guidance and instruction involving digital labor, allowing non-technical users to synthesize UI task automation programs. The system improves the processing speed for generating complex UI automation logic, including the definition, instruction, and verification of synthesized UI task automation programs. The system can efficiently and effectively generate UI task automation programs without requiring technical or programming user skills.

[0016] According to one aspect of the disclosed embodiments, a computer program product is provided. The computer program product includes a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations. The operations include: receiving an automation structure and inputs and outputs of the structure to create a given task automation; providing a multimodal UI interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generating interactive context guidance to record conditional execution of one or more actions or expressions based on the state of one or more UI elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification. This computer program product enables the automation of complex user interface (UI) tasks with code-free instruction and digital labor, synthesizing UI task automation programs for non-technical users. This computer program product improves the processing speed for generating complex UI automation logic, including the definition, instruction, and verification of synthesized UI task automation programs. This computer program product can effectively and efficiently generate UI task automation programs without requiring technical or programming skills from the user.

[0017] Embodiments of this disclosure also include: analyzing the UI task automation program; and generating user guidance and providing a multimodal interface to handle one or more additional instructional demonstrations. This embodiment enables the automatic synthesis of one or more additional instructional demonstrations into a consistent UI task automation program, which allows for the performance of additional checks and the identification of additional automated operations and parameters.

[0018] Embodiments of this disclosure further include: converting one or more instructional demonstrations into at least one logical abstract representation of task automation; and wherein the visual program representation presenting the UI task automation program is based on at least one logical abstract representation. Utilizing at least one logical abstract representation, enhanced overall processing time can be achieved for evaluating the synthesized UI task automation program.

[0019] Furthermore, embodiments of this disclosure, wherein converting one or more additional instructional demonstrations into at least one logical abstract representation of task automation further includes: combining multiple automated actions into a single logical action within the at least one logical abstract representation of task automation for presentation to a user for understanding. Enhanced overall processing time for evaluating the synthesized UI task automation program can be achieved by combining multiple automated actions into a single logical action within the at least one logical abstract representation, without displaying the calculations and details of the multiple automated actions, thereby enhancing user understanding.

[0020] Furthermore, in embodiments of this disclosure, generating interactive context guidance further includes: generating interactive context guidance to enable a user to select at least one automated action or expression on at least one UI element. This can enable the enhanced total processing time to record one or more instructional demonstrations and synthesize the UI task automation program with the user's selection of at least one automated action or expression on at least one UI element.

[0021] Additionally, in one embodiment of this disclosure, receiving the automation structure further includes: prompting and presenting a graphical visualization to the user to receive a user-selected definition for the automation structure, as well as one or more inputs and outputs of the automation structure to create task automation. The graphical visualization can be used to enable enhanced overall processing time for receiving user-selected definitions to create task automation.

[0022] Additionally, in one embodiment of this disclosure, presenting a visual program representation of a UI task automation program further includes providing at least one interactive tool having the visual program representation to enable a user to understand and verify the UI task automation program. The enhanced overall processing time for evaluating the synthesized UI task automation program can be achieved using at least one interactive tool that provides user guidance for understanding and verifying the UI task automation program.

[0023] Additionally, in one embodiment of this disclosure, the visual representation of the UI task automation program further includes enabling the user to accept and store the UI task automation program. This embodiment enables efficient and effective access to the UI task automation program while allowing the user to accept and store it.

[0024] Furthermore, in embodiments of this disclosure, recording one or more instructional demonstrations based on conditional execution of one or more actions or expressions further includes: automatically detecting at least one program parameter from the one or more instructional demonstrations, and using the one or more instructional demonstrations to record the at least one program parameter. This embodiment, by automatically detecting at least one program parameter, enables increased processing time for recording one or more instructional demonstrations for task automation, and synthesizes UI task automation programs from the one or more instructional demonstrations.

[0025] Additionally, in one embodiment of this disclosure, the visual program representation of the UI task automation program further includes: executing the UI task automation program to display UI actions and screen sequences to verify the behavior of the UI task automation program. By displaying the UI actions and screen sequences of the UI task automation program, an enhanced total processing time can be achieved for evaluating the synthesized UI task automation program.

[0026] In the disclosed embodiments, teaching digital labor or task automation includes prompting and providing guidance to the user to enable the user to select a definition of an automation type, and to receive the user's selection of a given structure, as well as the inputs and outputs of that structure, to create task automation. In one embodiment, a multimodal interface is provided to enable the user to perform one or more instructional demonstrations for task automation without requiring technical or programming user skills. The disclosed embodiments implement multimodal human-computer interaction through natural communication modes, thereby interfaceing the user with the automation system in terms of both input and output. The disclosed embodiments allow interaction through multiple input modalities such as keyboard input, mouse interaction, and natural language speech, and receive information by combining the output modalities of multiple input modalities according to time and contextual constraints, in order to analyze and derive semantic actions to update the task automation structure.

[0027] In the disclosed embodiments, one or more of the keyboard input, mouse interactions, and natural language utterances for each instructional presentation are recorded. In the disclosed embodiments, semantic actions from the instructional presentations are captured, and these semantic actions can be appended to elements of an abstract grammar or structure, where additional processing of abstract tree nodes is defined (e.g., to enable nodes to perform semantic checks and / or declare variables and variable scopes). The disclosed embodiments record conditional execution of task automation based on the state of one or more UI elements, where a user selects a UI element on the screen. The disclosed embodiments use context-aware advisory tools to provide automated user guidance to the user to define one or more actions or expressions associated with the selected UI element. In one embodiment, each instructional presentation is recorded, and one or more instructional presentations are transformed into an intermediate logical abstraction representation of task automation. The disclosed embodiments provide interactive context-aware guidance to enable the user to define one or more actions or expressions on one or more UI elements for task automation. In the disclosed embodiments, one or more instructional presentations to be processed are suggested to the user, and the user is enabled to enter another instructional mode. In the disclosed embodiments, interactive context-aware prompts include speech natural language (NL) utterances, such as enabling the user to utilize NL utterances to define business names, conditions, and decisions. In the disclosed embodiments, a task automation program is synthesized from one or more instructional demonstrations, and the task automation program is presented for user review and verification to confirm the correct execution of the task automation. In the disclosed embodiments, a program diagram is presented to provide a comprehensive view of the instructional demonstration logic. In the disclosed embodiments, a step-by-step automated view of the instructional demonstration logic is presented, enabling user verification.

[0028] The disclosed embodiments enable verification of UI task automation procedures and provide guided tools for understanding whether the task automation is being performed correctly, such as providing a flow of UI actions and screens that allows users to understand and verify the UI task automation procedure. The disclosed embodiments present users with graphical visual markers for a single instructional demonstration (e.g., a single scene) and for a complete UI task automation procedure. The disclosed embodiments enable users to verify the execution of UI task automation procedures by providing a flow or sequence of UI actions and screens, allowing users to determine whether the task automation is being performed correctly.

[0029] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies on the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0030] The following reference is made to embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specifically described embodiments. Rather, any combination of the following features and elements is contemplated for implementing and practicing the intended embodiments, regardless of whether different embodiments are involved. Furthermore, while the embodiments disclosed herein may achieve advantages over other possible solutions or prior art, whether a given embodiment achieves a particular advantage does not limit the scope of this disclosure. Therefore, the following aspects, features, embodiments, and advantages are merely illustrative and should not be considered elements or limitations of the appended claims unless expressly stated in the claims. Similarly, references to “the invention” should not be construed as a generalization of any inventive subject matter disclosed herein and should not be considered elements or limitations of the appended claims unless expressly stated in the claims.

[0031] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). Regarding any flowchart, depending on the technology involved, operations may be performed in a different order than that shown in a given flowchart. For example, again according to the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.

[0032] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in a collection of one or more storage devices, the collection of one or more storage devices collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. Computer-readable storage media, as used in this disclosure, should not be construed as storing transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data is typically moved at certain incidental points in time during the normal operation of the storage device, such as during access, defragmentation, or garbage collection; however, this does not make the storage device transient, because the data is not transient when it is stored.

[0033] Referring to Figure 1, computing environment 100 includes an example of an environment for executing at least some of the computer code involved in the method of the present invention executed in block 180, such as UI automation task control code 182. In addition to block 180, computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), communication infrastructure 111, volatile memory 112, persistent storage device 113 (including an operating system 122 and block 400, as described above), a peripheral device set 114 (including a user interface (UI) device set 123, storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes gateway 140, cloud coordination module 141, host physical machine set 142, virtual machine set 143, and container set 144.

[0034] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases such as remote database 130. As is well known in the field of computer technology, and depending on the technology, the performance of a computer-implemented method can be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, in particular computer 101, to keep the presentation as simple as possible. Computer 101 may be located in the cloud, even if it is not shown in the cloud in Figure 1; on the other hand, computer 101 does not need to be in the cloud unless it can be definitively indicated to any extent.

[0035] Processor assembly 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, such as multiple cooperating integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be readily accessible by the threads or cores running on processor assembly 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor assembly may be located “off-chip.” In some computing environments, processor assembly 110 may be designed to work with qubits and perform quantum computing.

[0036] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the invention. In computing environment 100, at least some of the instructions for performing the method of the invention may be stored in permanent storage device 113, within block 400.

[0037] Communication structure 111 is a signal transmission path that allows the various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0038] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.

[0039] The persistent storage device 113 is any form of non-volatile memory known now or developed in the future for use with a computer. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some common forms of persistent storage include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or operating systems employing an open-source portable operating system interface type with a kernel. The code included in box 400 typically includes at least some of the computer code involved in performing the methods of the present invention.

[0040] Peripheral device set 114 includes a set of peripheral devices for computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 124 can be permanent and / or volatile. In some embodiments, storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires substantial storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed for storing very large amounts of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 comprises sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.

[0041] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.

[0042] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area (e.g., a Wi-Fi network). WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0043] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives useful and available data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to an end user, these recommendations are typically transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present recommendations to the end user. In some embodiments, EUD 103 can be client equipment such as a thin client, heavy client, mainframe, desktop computer, etc.

[0044] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 can be controlled and used by the same entity operating computer 101. Remote server 104 represents a machine that collects and stores useful and available data used by other computers, such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 104.

[0045] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (particularly data storage (cloud storage) and computing power) without the need for direct, active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud coordination module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting the host physical machine set 142, which is the entirety of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud coordination module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allow public cloud 105 to communicate via WAN 102.

[0046] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.

[0047] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables coordination, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud.

[0048] The disclosed embodiments enable non-technical users to generate UI task automation programs, allowing users to execute one or more instructional demonstrations and synthesize UI task automation programs from one or more instructional demonstrations for user review and approval.

[0049] Features of the disclosed embodiments of the methods, systems, and computer program products enable users to define automation through user-selected structures or templates and one or more inputs and outputs of that structure. The disclosed embodiments enable users to record multiple instructional demonstrations or execution scenarios corresponding to different business situations via keyboard and mouse interaction and natural language speech. The disclosed embodiments achieve conditional execution of recording based on the state of UI elements by allowing users to select elements on the screen and providing automated user guidance on Boolean expressions associated with the selected elements. The disclosed embodiments automatically merge multiple scenarios (e.g., instructional demonstrations) into a consistent procedure. The disclosed embodiments demonstrate a higher-level visual program representation that is understandable to non-technical business users. The disclosed embodiments enable users to perform automated verification that allows users to check coverage and prompts users, suggests additional demonstrations to be performed, and enables users to perform debugging to check the UI task automation program in action (i.e., the constructed digital labor robot) and verify the correctness of its behavior against sample data.

[0050] Referring now to FIG2, an example system 200 for implementing UI task automation of one or more of the disclosed embodiments is shown. System 200 can be used in conjunction with computer 101 of computing environment 100 of FIG1 and a cloud environment, wherein UI automation task control code 182 is provided for implementing UI task automation of the disclosed embodiments. In the disclosed embodiments, system 200 enables UI task automation to be generated effectively and efficiently by non-technical users.

[0051] In the disclosed embodiments, system 200 includes one or more processors 202 and a memory 204 storing a UI task automation component 206. System 200 includes a definition component 208, a teaching component 210, a verification component 212, and a ready component 214, as well as the UI task automation component 206 of the disclosed embodiments. System 200 enables non-technical users to select the definition component 208 for a defined operation mode for generating UI task automation programs, allowing users to select an automation structure and its inputs and outputs to create task automation. System 200 enables non-technical users to select the teaching component 210 for a teaching operation mode for generating UI task automation programs and provides a multimodal UI interface so that users can execute one or more teaching demonstrations for task automation. System 200 enables non-technical users to select the verification component 212 for a verification operation mode for generating UI task automation programs and provides users with a flow of UI operations and screens to determine whether the complete UI task automation program is executed correctly.

[0052] In the disclosed embodiments, system 200 includes a UI task automation list 220 that stores multiple task automation applications or programs, such as multiple task automation programs identified by the user as ready for production using the ready component 214 of UI task automation component 206. For example, after the user confirms automation specifications and verifies behavior, system 200 presents the user with a given complete task automation program (e.g., digital labor or robotic program) ready for production. In the disclosed embodiments, when it is determined that a given UI task automation program fails to perform correctly, system 200 generates user guidance to fix or remove the instructional demonstration, which otherwise creates inconsistencies during program synthesis. Furthermore, for example, system 200 may generate user guidance to provide the user with suggestions for additional instructional demonstrations and provide a multi-modal interface to handle one or more additional instructional demonstrations. System 200 updates the UI task automation list 220 with the complete task automation program, for example, when identified by the user as ready for production.

[0053] Figures 3A and 3B together illustrate example operations of a method 300 for implementing enhanced UI task automation of one or more of the disclosed embodiments. In the disclosed embodiments, method 300, implemented, for example, by system 200 in conjunction with UI automation task control code 182 and computer 101 of FIG. 1, generates UI task automation programs according to one or more of the disclosed embodiments. As shown, method 300 includes example operations performed by system 200 to enable UI task automation programs to be generated by non-technical users without requiring programming skills.

[0054] As shown in block 302, system 200 displays a user interface (UI) to enable user specifications to generate UI task automation programs. In block 304, system 200 begins defining an operating mode, where system 200 prompts the user and receives a user-defined automation type, template, or structure, as well as one or more inputs and outputs of that structure to create task automation. In the disclosed embodiments, system 200 prompts and presents a graphical visualization to the user to provide a definition of an automation type structure, as well as one or more inputs and outputs of that structure, to create task automation. In the disclosed embodiments, the structure includes an automation name, an automation type, and inputs and outputs (e.g., an existing file such as an Excel file, and receiving Excel inputs and Excel outputs), the type or result of the inputs and outputs (e.g., decision, extraction), and a value space (e.g., decision A, B, C as a string). An example structure 402 of the disclosed embodiments is shown and described with reference to FIG4.

[0055] In block 306, to enable a user to provide one or more instructional demonstrations (e.g., scenarios), system 200 generates and provides a multimodal UI interface to receive input for processing one or more instructional demonstrations for task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for task automation. In the disclosed embodiments, system 200 provides a multimodal UI interface including an input multimodal interface (e.g., enabling the user to type or click and / or speak input) and an output multimodal interface or output modality (e.g., enabling the user to visualize the created steps and receive auditory feedback of the created steps with names and details). In the disclosed embodiments, the generated multimodal UI interface enables keyboard and mouse actions on UI elements, as well as object selection using nearby elements, natural language text, and voice actions for a given task automation. In the disclosed embodiments, system 200 can use natural language actions to invoke actions in one or more multimodal UI interfaces, provide documentation and description of a process, or name a business entity that the user can later refer to (e.g., "This is the search result"). In the disclosed embodiments, system 200 records one or more of the keyboard inputs, mouse interactions, and natural language speech for each instructional demonstration.

[0056] In block 308, system 200 identifies and analyzes keyboard and mouse actions on elements, object selection using nearby elements, natural language text and natural language speech actions, such as one or more actions or expressions (e.g., Boolean expressions) that automatically invoke primitive actions and multimodal primitive events, and conditional execution of tasks automated or digital labor robot operations. In one embodiment, system 200 captures semantic actions from instructional demonstrations and can attach semantic actions to elements of abstract syntax or structures, and define additional processing on abstract tree nodes (e.g., enabling nodes to perform semantic checks, identify automation parameters, declare variables and variable ranges, and include one or more multimodal primitive events including keyboard and mouse actions and natural language actions).

[0057] In block 310, system 200 provides interactive contextual guidance to the user, enabling the user to define actions and expressions associated with selected UI elements. In the disclosed embodiments, system 200 records conditional execution based on the state of one or more UI elements, where the user selects a UI element on the screen, and the system uses an automation context advisor tool to provide the user with automated user guidance to define one or more actions or expressions associated with the selected UI element. In embodiments, system 200 provides contextual information related to the automation task and provides interactive contextual guidance to help the user define actions and expressions on the UI element, such as explaining and providing one or more examples and options for the user to choose from to use the UI element for a given outcome or result. In embodiments, the user may select an object based on nearby text information, system 200 calculates possible logical actions for that object (e.g., including Boolean or other expressions), and enables the user to select one or more of the possible logical actions. In embodiments, system 200 allows natural language actions to invoke actions in the UI element and provides documentation or a description of the process, or names the entity the user will later want to use (e.g., identifying the selected search result). System 200 performs automatic computation and semantic analysis, enabling users to record multiple execution scenarios corresponding to different situations via keyboard and mouse interaction and natural language utterance. System 200 allows users to perform conditional execution based on the state of UI elements by selecting elements on a UI screen (e.g., a given multimodal UI interface), and System 200 provides automatic guidance on Boolean expressions associated with the selected elements. Example logic automation structures for implementing instructional demonstrations or scenarios using System 200 are shown in Figures 5, 6, 7, and 8.

[0058] At block 312, system 200 records and analyzes one or more instructional demonstrations and synthesizes a UI task automation program for the current automation structure of the task automation. At block 314, system 200 provides the user with guidance on handling additional instructional demonstrations or scenarios related to the current automation structure and UI task automation program. In the disclosed embodiments, system 200 suggests one or more instructional demonstrations for processing and enables another instructional mode to handle one or more additional instructional demonstrations for task automation. For example, system 200 systemically provides the user with suggestions for additional scenarios to record additional values ​​in the automation structure, or to fix or remove scenarios that otherwise generate inconsistencies during program synthesis. Operation continues at block 316 in FIG3B.

[0059] Referring to Figure 3B, at block 316, system 200 converts one or more instructional demonstrations into corresponding logical abstract representations of task automation. In the disclosed embodiments, system 200 automatically detects program parameters from the instructional demonstrations (e.g., by the similarity of input fields and UI fields, and the similarity of values), and records those program parameters in the scenario of the demonstrations as one or more instructional demonstrations are converted. In the disclosed embodiments, system 200 combines multiple automated actions into a single logical action in the logical abstract representation of each instructional demonstration to make it understandable to the user, thereby providing a logical abstract representation that is more easily understood by non-technical users. For example, the system may combine multiple calculations and actions to present a single logical action for user review without showing the details of such calculations and actions. In the disclosed embodiments, system 200 presents or renders a visual program representation to the user based on the logical abstract representation of the corresponding instructional demonstration to make it understandable to the user.

[0060] In block 318, system 200 synthesizes a UI task automation program based on one or more instructional demonstrations or scenarios. In block 320, system 200 analyzes the generated UI task automation program and provides the user with guidance and a multimodal interface to execute one or more further additional instructional demonstrations or scenarios, thereby covering further value space of the current automation structure of the UI task automation program. In block 322, system 200 synthesizes or merges one or more instructional demonstrations or scenarios into a coherent UI task automation program that includes one or more additional instructional demonstrations.

[0061] In block 324, system 200 displays or presents a visual representation of the UI task automation program based on a logical abstract representation of the program. In the disclosed embodiments, system 200 provides users with interactive tools for understanding and verifying the UI task automation program. According to the disclosed embodiments, system 200 provides users with the ability to view the operation of a UI task automation program performing defined task automation on sample test data, for example, the sample test data representing different scenarios, and system 200 provides a live graphical view of the program mapping execution for each test data. In the disclosed embodiments, system 200 presents users with graphical visual markers for individual instructional demonstrations or scenarios as well as for the complete UI task automation program. In block 326, system 200 provides users with the option to confirm the specifications and verified behavior of the UI task automation program as preparation for production, and then system 200 accepts and stores the UI task automation program, for example, for access at the UI task automation application list 220 of Figure 2.

[0062] Figure 4 illustrates an example automation structure 400 for implementing one or more of the disclosed embodiments, defining input, output, and decision values ​​for UI task automation. In the disclosed embodiments, automation structure 400 illustrates an example automation structure implemented by system 200 in conjunction with UI automation task control code 182 and computer 101 of Figure 1, according to one or more of the disclosed embodiments, for defining operating modes to build or generate UI task automation programs. The defined operating modes enabled by automation structure 400 allow for defining input, output, and decision values, as described with respect to box 304 of Figure 3. As shown, automation structure 400 includes a template or automation structure 402 of the automation type. Automation structure 402 includes one or more Decisions 404 with a [List[Decisions]], one or more Inputs 406 with a List[Entity], Data 408 with an Optional [dict], and an output or Outcome 410 with a value string (Str) for creating task automation. In the disclosed embodiments, the [List[Decisions]] of one or more Decisions 404 is represented by example Decisions 412, which includes Name 414 with value Str and Values ​​416 of List[Entity] with automation structure 400. In the disclosed embodiments, each of the List[Entity] of one or more Inputs 406 and the List[Entity] of one or more Values ​​416 is represented by example Entity 418. As shown, Entity 418 includes Name 420 with value Str and Synonyms 422 with Optional[List(Str)]. In the disclosed embodiments, the Inputs 406 and Outcome 410 of the Automation Structure 402 may include existing spreadsheet or database files, database inputs and database outputs, types of inputs and outputs (e.g., Decisions 412, Outcome 410), and Values ​​416 (e.g., decisions A, B, C as strings).

[0063] Figures 5, 6, and 7 together illustrate example logical automation structures 500, 600, and 700 for implementing UI task automation of the disclosed embodiments. In the disclosed embodiments, automation structures 500, 600, and 700 generate UI task automation programs implemented by system 200 in conjunction with UI automation task control code 182 according to one or more of the disclosed embodiments and computer 101 of FIG. 1. Logical automation structures 500, 600, and 700 include high-level logical structures that are presented to the user and are understandable to the end user, as well as low-level logical structures that are not presented to the user.

[0064] In Figure 5, the logic automation structure 500 includes a Program 502 for implementing UI task automation of the various disclosed embodiments. As shown, Program 502 includes an identifier (Id) 504 with a value string Str, and a Scenery_Trace or Scn_Trace 506 with a List[Scn_Trace], wherein Scenery or Scn is also referred to as a teaching demonstration and can be used interchangeably with the teaching demonstrations of the embodiments disclosed in the following description. After entry point A, an example Scenery_Trace 732 is shown and described with reference to Figure 7. As shown, Program 502 includes a Name 508 with a value Str, Arguments 510 with a List[Argument], and Steps 512 with a List[Union[S]]. In the disclosed embodiments, for example, S in Steps 512 is defined or represents: S = Program Click, Program Type Into, Program Extract Text, Program Extract Table, and Program Return.

[0065] The corresponding example logic automation structure of Program 502 is shown and described relative to Program Click 522 in Figure 5, Program Type Into 602 and Program Return 620 after entry points C and B in Figure 6, respectively, and Program Extract Table 702 and Program Extract Text 716 after entry points C and B in Figure 7, respectively.

[0066] As shown in Figure 5, the example logical automation structure of Arguments 510 includes Argument 514, which includes Type 516 with a value Str and Name 518 with a value Str. As shown, ProgramClick 522 includes Id 524 with a value Str and Name 526 with a value Str, as well as Scenario_Trace or Scn_Trace 526 with List[Scn_Trace], where the illustrative example Scenario_Trace 732 is illustrated and described relative to Figure 7 after entry point A. ProgramClick 522 includes type 528 with Literal[RPA_Action] and subtype 530 with Literal['Click'], where corresponding literal RPA operations and click or keyboard entries can be maintained. Program Click 522 includes Selector 532, Value 534, Label 536 and Role 538, each with Optional valueStr.

[0067] In Figure 6, the logical automation structure 600 includes illustrative examples of Program Type Into 602 and Program Return 620. As shown, Program Type Into 602 includes an Id 604 with the value Str, a Scn_Trace 606 with List[Scn_Trace], a Type 608 with Literal[RPA_Action], and a Subtype 610 with Literal['Type']. Following entry point A, an example Scenario_Trace 732 is shown and described with reference to Figure 7. Program Type Into 602 includes a Value 612 with Union[T], where T is defined or represented as: T = Variable Read, Constant, Argument Read, and Unresolved.

[0068] Illustrative examples of the logical automation structure of T include Variable Read 640, Constant 646, Argument Read 630, and Unresolved 636, as shown in Figure 6. As shown, Program Type Into 602 includes Sample Value 614 and Label 616, each with an Optional value Str, and a Selector 618 with the value Str. As shown, Argument Read 630 includes a Type 632 with Literal['Read_Argument'] and a Name 634 with the value Str, and Unresolved 636 includes a Type 638 with Literal['Unresolved'].

[0069] As shown in Figure 6, Program Return 620 includes an Id 622 with the value Str, a Scn_Trace 624 with List[Scn_Trace] (i.e., example Scenario_Trace 732 after entry point A in Figure 7), and a type 626 with Literal['Return']. Program Return 620 includes a value 628 with Union[V], where V is defined or represented by: V = Program Extract Text, Program Extract Table, and Constant and Variable Read.

[0070] Illustrative examples of the logical automation structure of V include Program Extract Text 716 and Program Extract Table 702, which are respectively following entry points B and C as illustrated and described with respect to Figure 7, and Constant 646 and Variable Read 640 as shown in Figure 6. As shown, Constant 646 includes Type 648 with Literal['Constant'] and Value 650 with the value Str, and Variable Read 640 includes Type 642 with Literal['Variable_Read'] and Name 644 with the value Str.

[0071] In Figure 7, the logic automation structure 700 includes a Program Extract Table 702 and a Program Extract Text 716 following entry points C and B as referenced in Figures 5 and 6, and a Scenario_Trace 732 following entry point A, also referenced in Figures 5 and 6. The logic automation structure 700 includes a Variable Definition Table 738 and a Variable Definition Type 746. In the disclosed embodiment, the Program Extract Table 702 includes an Id 704 with the value Str, a Scn_Trace 706 with a List[Scn_Trace] (i.e., the example Scenario_Trace 732 shown after entry point A in Figure 7), and a Type 708 with a Literal['Extract_Table']. The ProgramExtract Table 702 includes a Selector 710 with Optional[Str], an Out 712 with a Variable Definition Table, and Headers 714 with List[Str].

[0072] In the disclosed embodiments, Program Extract Text 716 includes an Id 718 with a value Str, a Scn_Trace 720 with a List[Scn_Trace] (i.e., a Scenario_Trace 732 shown after entry point A in FIG. 7), a Type 722 with a Literal['RPA_Action'], and a Subtype with a Literal['Extract_Text']. Program Extract Text 716 includes a Selector 726 with an Optional[Str], an Out 728 with a Variable Definition Type, and a Label 730 with an Optional[Str].

[0073] As shown in Figure 7, Scenario_Trace 732 includes Scenario_Name 734 with the value Str and Action_Indexes 736 with Optional[List[Str]]. As shown in the illustrative example, VariableDefinition (Defin)Table 738 includes Type 740 with Literal['VariableDefinTable'], Variable_Type 742 with Literal['Table'], and Name 744 with the value Str. For example, Variable Definition Type 746 includes Type 748 with Literal['VariableDefinType'], Variable_Type 750 with Literal['Str'], and Name 752 with the value Str.

[0074] In the disclosed embodiments, the example automation structures 500, 600, and 700 of Figures 5, 6, and 7 are used together to record multimodal raw events to generate a UI task automation program in response to receiving user input from a non-technical user, identifying raw expressions and / or actions, and generating program traces. The Program 502 in Figure 5 of the disclosed embodiment of UI task automation maintains references to raw actions, such as those demonstrated in, for example, the teaching demonstration or scenario shown and described in Figure 8.

[0075] Figure 8 illustrates an example logical automation structure 800 for automating processing parameters and operations of a teaching demonstration or scenario according to a disclosed embodiment. The logical automation structure 800 includes an example list of multimodal raw events, including keyboard / mouse actions (NL) for a teaching demonstration or scenario 802. The illustrated scenario 802 includes a Name 804 with a value Str and Actions 806 with a List[Str]. Actions 806 include one or more of Init_Action 808, NL_Action 814, or RPA_Action 824, as shown. For example, system 200 together identifies and processes the list of initialization actions, NL actions, and RPA actions to automate one or more teaching demonstration or scenario UI tasks of the disclosed embodiments.

[0076] In the disclosed embodiments, as shown, Init_Action 808 includes Program_Trace 810 with Optional[Program_Trace] and Type 812 with Literal['Initialize']. For example, NL_Action 814 includes Program_Trace 816 with Optional[Program_Trace], Type 818 with Literal['NL_Action'], Selector 820 with Optional[Str], and NL_Value with Optional[Str]. For example, RPA_Action 824 includes Program_Trace 826 with Optional[Program_Trace], Type 828 with Literal['RPA_Action'], and Subtype 830 with Optional[Literal['Click', 'Type']]. For example, RPA_Action 824 includes Semantics 832 with Optional [SemanticUnderstanding], Selector 834 with Optional [Str], and Value 836 with Optional [Str].

[0077] In the disclosed embodiments, the instructional demonstration or scenario 802 includes an automated structure Program Trace 840 for defining Program Trace 810 for Init_Action 808, Program Trace 816 for NL_Action 814, and Program Trace 826 for RPA_Action 824. Program Trace 840 includes Program_Name 842 with Optional[Str] and Program_Step_Index 844 with Optional[Str].

[0078] In the disclosed embodiments, Scenario 802 includes an automated structure Semantic Understanding 846 for the original actions demonstrated in the scenario of the disclosed embodiments. Semantic Understanding 846 840 includes Is_input 848 with Optional[Str], Is_button 850 with Optional[Str], Nearby_label 852 with Optional[Str], and Is_navigation 854 with Optional[Str].

[0079] Figure 9 illustrates a method 900 for automating UI tasks according to the disclosed embodiments. Method 900 illustrates the features and operations of method 300, and the automation structures 400, 500, 600, 700, and 800 of the disclosed embodiments. For example, method 900 is implemented by system 200 in conjunction with computer 101 of Figure 1 using UI automation task control code 182.

[0080] At block 902, system 200 receives an automation structure or template, as well as one or more inputs and outputs to the automation structure to create task automation. In the disclosed embodiments, system 200 prompts the user and presents graphical visualizations to provide a definition of the automation structure and one or more inputs and outputs to the automation structure to create task automation. At block 904, system 200 provides a multimodal interface to receive input and process one or more instructional demonstrations for task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations of the disclosed embodiments. The disclosed multimodal interface includes keyboard and mouse actions on UI elements of task automation, object selection using nearby UI elements, natural language text, and voice actions. For example, in the disclosed embodiments, system 200 enables the use of natural language actions to invoke actions with a multimodal interface, provides documentation and descriptions for actions or decision-making processes, or named entities for later user reference.

[0081] In block 906, system 200 generates an interactive context guide to enable a user to record conditional execution of one or more actions or expressions based on the state of one or more UI elements for task automation. The disclosed multiple interfaces, for example, allow the user to select an object based on nearby text information, system 200 calculates possible logical processing actions (including expressions) for that object, sends them for display, and the user can select one of the logical processing actions. In block 908, system 200 records one or more instructional demonstrations for task automation based on the conditional execution of one or more actions or expressions. In block 910, system 200 synthesizes a UI task automation program from one or more instructional demonstrations. In the disclosed embodiments, system 200 analyzes the instructional demonstrations and may automatically synthesize or merge multiple instructional demonstrations into a consistent UI task automation program. In block 912, system 200 presents a visual program representation of the UI task automation program for verification. For example, system 200 executes the UI task automation program to display a sequence of UI actions and screens for the user to view and verify the behavior of the UI task automation program.

[0082] While the foregoing relates to embodiments of the present invention, other and further embodiments of the present invention may be designed without departing from the basic scope of the present invention, and the scope of the present invention is defined by the appended claims.

Claims

1. A method comprising: Receive an automation structure and one or more inputs and outputs of the automation structure to create task automation; Provide a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generate interactive context guidance to record conditional execution of one or more automated actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; record the one or more instructional demonstrations based on the conditional execution of the one or more actions or expressions; synthesize a UI task automation program from the one or more instructional demonstrations; and present a visual program representation of the UI task automation program for verification.

2. The method according to claim 1, further comprising: Analyze the UI task automation program; It also generates user guidance and provides the multimodal interface to handle one or more additional instructional demonstrations.

3. The method according to claim 1, further comprising: Transform the one or more instructional demonstrations into at least one logical abstract representation of the task automation; Furthermore, the visual program representation that presents the UI task automation program is based on the at least one logical abstract representation.

4. The method according to claim 3, wherein, Converting the one or more additional instructional demonstrations into the at least one logical abstract representation of the task automation further includes: combining multiple automated actions into a logical action in the at least one logical abstract representation of the task automation to present it for user understanding.

5. The method according to claim 1, wherein, Generating the interactive context guide further includes: generating the interactive context guide to enable the user to select at least one automated action or expression on at least one UI element.

6. The method according to claim 1, wherein, Receiving the automation structure further includes: prompting and presenting a graphical visual sign to receive a definition of user selection for the automation structure, as well as the one or more inputs and outputs of the automation structure to create the task automation.

7. The method according to claim 1, wherein, The visual program representation of the UI task automation program further includes providing at least one interactive tool having the visual program representation to enable a user to understand and verify the UI task automation program.

8. The method according to claim 1, wherein, The visual representation of the UI task automation program further includes enabling the user to accept and store the UI task automation program.

9. The method according to claim 1, wherein, Based on the conditions of the one or more actions or expressions, recording the one or more instructional demonstrations further includes: automatically detecting at least one program parameter from the one or more instructional demonstrations, and using the one or more instructional demonstrations to record the at least one program parameter.

10. The method according to claim 1, wherein, The visual program representation of the UI task automation program further includes: executing the UI task automation program to display a sequence of UI operations and screens to verify the behavior of the UI task automation program.

11. A system comprising: One or more computer processors; The system also includes a memory containing a program that, when executed by the one or more computer processors, performs operations including: receiving an automation structure and one or more inputs and outputs of the automation structure to create a task automation; providing a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generating an interactive context guide to record conditional execution of one or more actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; recording the one or more instructional demonstrations based on the conditional execution of the one or more actions or expressions; synthesizing a UI task automation program from the one or more instructional demonstrations; and presenting a visual program representation of the UI task automation program for verification.

12. The system of claim 11, further comprising: Analyze the UI task automation program; It also generates user guidance and provides the multimodal interface to handle one or more additional instructional demonstrations.

13. The system of claim 11, further comprising: The one or more instructional demonstrations are converted into at least one logical abstract representation of the task automation; and wherein the visual program representation presenting the UI task automation program is based on the at least one logical abstract representation.

14. The system according to claim 11, wherein, The visual program representation of the UI task automation program further includes: executing the UI task automation program to display a sequence of UI operations and screens to verify the behavior of the UI task automation program.

15. The system according to claim 11, wherein, The visual program representation of the UI task automation program further includes providing at least one interactive tool having the visual program representation to enable a user to understand and verify the UI task automation program.

16. A computer program product comprising a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations, the operations including: Receive an automation structure and one or more inputs and outputs of the automation structure to create task automation; Provide a multimodal interface to receive input and process one or more instructional demonstrations for the task automation, wherein the one or more instructional demonstrations identify automation processing parameters and operations for the task automation; generate interactive context guidance to record conditional execution of one or more actions or expressions based on the state of one or more user interface (UI) elements of the one or more instructional demonstrations; record the one or more instructional demonstrations based on the conditional execution of the one or more actions or expressions; synthesize a UI task automation program from the one or more instructional demonstrations; and present a visual program representation of the UI task automation program for verification.

17. The computer program product of claim 16, further comprising: Analyze the UI task automation program; It also generates user guidance and provides the multimodal interface to handle one or more additional instructional demonstrations.

18. The computer program product of claim 16, further comprising: The one or more instructional demonstrations are converted into at least one logical abstract representation of the task automation; and wherein the visual program representation presenting the UI task automation program is based on the at least one logical abstract representation.

19. The computer program product according to claim 16, wherein, The visual program representation of the UI task automation program further includes: executing the UI task automation program to display a sequence of UI operations and screens to verify the behavior of the UI task automation program.

20. The computer program product according to claim 16, wherein, The visual program representation of the UI task automation program further includes providing at least one interactive tool having the visual program representation to enable a user to understand and verify the UI task automation program.