Method and system for learning and enabling commands via user demonstration

Through user demonstrations to generate task commands, virtual assistants can execute tasks in mobile applications without APIs, solving the problem that virtual assistants have difficulty in executing user requests and improving the flexibility and efficiency of task execution.

CN113646746BActive Publication Date: 2025-05-09SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080025796.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-29
Filing Date
2020-03-18
Publication Date
2025-05-09
Estimated Expiration
2040-03-18

AI Technical Summary

Technical Problem

In the prior art, virtual assistants and smart agents in mobile applications have difficulty performing user requests because most applications lack the provision of APIs or interfaces to implement specific tasks.

Method used

By using user demonstrations to generate task commands, the virtual assistant can learn and execute tasks in the application. The method includes obtaining information associated with the application, recording a sequence of user interface interactions, extracting information, filtering events or actions, generating semantic ontology, and generating task commands based on semantic ontology and filtered information.

Benefits of technology

This method allows virtual assistants to perform user requests in the absence of API, reducing the workload and time of developers and improving the task execution capabilities of virtual assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113646746B_ABST
    Figure CN113646746B_ABST
Patent Text Reader

Abstract

A method for learning a task includes obtaining first information associated with at least one application executed by an electronic device. Recording a sequence of user interface interactions for the at least one application. Extracting second information from the sequence of the user interface interactions. Filtering at least one of an event or an action from the second information based on the first information. Performing discrimination on each element included in the first information to generate a semantic ontology. Generating a task command for the sequence of the user interface interactions based on the semantic ontology and the filtered second information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments generally relate to task learning for electronic devices, and more particularly to task learning for a virtual assistant or voice assistant using user demonstrations for generating task commands for at least one application. Background Art

[0002] Personal assistants (PA) and intelligent agents are widely present in mobile devices, TV devices, home speakers, consumer electronics, etc. to implement user tasks as skills on multi-modal devices. In mobile devices, most PAs and intelligent agents work with applications to perform specific tasks in response to user requests implemented in voice commands, text input, shortcut buttons and / or gestures. In order to perform these tasks, the platform developer or application developer of the PA or intelligent agent needs to implement 'task execution' by calling the application programming interface (API) provided by the app developer. For example, if there is no API or interface in a specific application to implement an action to complete a task, the user request may not be executed.

[0003] On a popular mobile app platform ”, as of the first quarter of 2018, there are approximately 3.8 million apps in the app store. In this app store, less than 0.5% of the apps have APIs provided to developers to complete tasks when used in PA and intelligent agents. Summary of the invention

[0004] Solution to the problem

[0005] One or more embodiments generally relate to performing task learning for a virtual assistant using user demonstrations for generating task commands for at least one application. In one embodiment, a method for learning a task includes obtaining first information associated with at least one application executed by an electronic device. Recording a sequence of user interface interactions for the at least one application. Extracting second information from the sequence of user interface interactions. Filtering at least one of an event or an action from the second information based on the first information. Performing discrimination on each element included in the first information to generate a semantic ontology. Generating a task command for the sequence of user interface interactions based on the semantic ontology and the filtered second information.

[0006] In some embodiments, an electronic device includes a memory storing instructions. At least one processor executes the instructions, the instructions including a process configured to: obtain first information associated with at least one application executed by the electronic device; record a sequence of user interface interactions about the at least one application; extract second information from the sequence of user interface interactions; use the first information to filter at least one of an event or an action from the second information; perform discrimination on each element included in the first information to generate a semantic ontology; and generate a task command for the sequence of user interface interactions based on the semantic ontology and the filtered second information.

[0007] In one or more embodiments, a non-transitory processor-readable medium includes a program that, when executed by a processor, performs a method, the method comprising obtaining first information associated with at least one application executed by an electronic device. Recording a sequence of user interface interactions about the at least one application. Extracting second information from the sequence of user interface interactions. Filtering at least one of an event or an action from the second information based on the first information. Performing discrimination on each element included in the first information to generate a semantic ontology. Generating a task command for the sequence of user interface interactions based on the semantic ontology and the filtered second information.

[0008] These and other aspects and advantages of one or more embodiments will become apparent from the following detailed description which, when taken in conjunction with the accompanying drawings, illustrates by way of example the principles of the one or more embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] For a more complete understanding of the nature and advantages of the embodiments and the preferred mode of use, reference should be made to the following detailed description read in conjunction with the accompanying drawings, in which:

[0010] Figure 1 shows a schematic diagram of a communication system according to some embodiments;

[0011] Figure 2 A block diagram illustrating the architecture of a system for performing intelligent learning processing, either alone or in combination, including an electronic device and a cloud or server environment, according to some embodiments;

[0012] Figure 3 shows a high-level block diagram of an Intelligent Learning System (ILS) process according to some embodiments;

[0013] Figure 4 shows an example flow of an application for user task demonstration (using simple steps) according to some embodiments;

[0014] Figure 5 illustrates a high-level component flow of an ILS for filtering unwanted events from systems and services according to some embodiments;

[0015] Figure 6 shows a high-level flow of an ILS process according to some embodiments that performs dynamic event sorting by filtering events based on packages, data and semantics such as icons, text, and descriptions, and prioritizing events and text / image semantic ontologies;

[0016] Figure 7 shows an example process flow diagram for performing ILS task learning on a single application according to some embodiments;

[0017] Figure 8 Another example process flow diagram for performing ILS task learning on more than one application according to some embodiments is shown;

[0018] Fig.9A , Fig. 9B , Fig. 9C and Fig.9D According to some embodiments, Examples of learning to post on the app;

[0019] Fig. 10A , Fig. 10B and Fig. 10C According to some embodiments, Example of learning flight booking on the app;

[0020] Fig.11A , Fig. 11B and Fig. 11C According to some embodiments, Example of finding a restaurant on the app;

[0021] Fig. 12A , Fig. 12B and Fig. 12C shows an example of sharing travel pictures from a photo album to a cloud service according to some embodiments;

[0022] Fig.13 shows a block diagram of a process for task learning according to some embodiments; and

[0023] Fig.14 is a high-level block diagram illustrating an information handling system including a computing system that implements one or more embodiments. DETAILED DESCRIPTION

[0024] The following description is made to illustrate the general principles of one or more embodiments and is not meant to limit the inventive concepts claimed herein. In addition, the specific features described herein may be used in combination with other described features in each of various possible combinations and permutations. Unless otherwise expressly defined herein, all terms will be given their broadest possible interpretation, including the meaning implied from the specification and the meaning understood by those skilled in the art and / or the meaning defined in dictionaries, monographs, etc.

[0025] It should be noted that the term "at least one of" refers to one or more of the following elements. For example, "at least one of a, b, c, or a combination thereof" can be interpreted as "a", "b", or "c" alone; or "a" and "b" combined together, "b" and "c" combined together, "a" and "c" combined together; or "a", "b", and "c" combined together.

[0026] In one or more embodiments, a "task" may refer to a user task that includes one or more steps for completing a goal. Some tasks may be implemented via interaction with one or more apps, such as adding a calendar event, sending a message, booking a flight, making a travel reservation, etc. Some tasks may involve one or more devices or equipment and one or more actions to control one or more devices or equipment, such as setting up a nighttime entertainment environment, arranging a morning routine, etc.

[0027] One or more embodiments provide for task learning for a virtual assistant using user demonstrations for generating task commands for at least one application. In some embodiments, a method for learning a task includes obtaining first information associated with at least one application executed by an electronic device. Recording a sequence of user interface interactions for the at least one application. Extracting second information from the sequence of the user interface interactions. Filtering at least one of an event or an action from the second information based on the first information. Performing discrimination on each element included in the first information to generate a semantic ontology. Generating a task command for the sequence of user interface interactions based on the semantic ontology and the filtered second information.

[0028] Figure 1 1 is a schematic diagram of a communication system 10 according to one or more embodiments. The communication system 10 may include a communication device (transmitting device 12) that initiates an outgoing communication operation and a communication network 110, and the transmitting device 12 may use the communication network to initiate and implement communication operations with other communication devices within the communication network 110. For example, the communication system 10 may include a communication device (receiving device 11) that receives communication operations from the transmitting device 12. Although the communication system 10 may include multiple transmitting devices 12 and receiving devices 11, Figure 1Only one of each is shown in order to simplify the drawing.

[0029] The communication network 110 may be created using any suitable circuits, devices, systems, or combinations thereof operable to create a communication network (e.g., a wireless communication infrastructure including communication towers and telecommunication servers). The communication network 110 may be capable of providing communications using any suitable communication protocol. In some embodiments, the communication network 110 may support, for example, traditional telephone lines, cable television, Wi-Fi (e.g., IEEE 802.11 protocol), High frequency systems (e.g., 900 MHz, 2.4 GHz, and 5.6 GHz communication systems), infrared, other relatively localized wireless communication protocols, or any combination thereof. In some embodiments, the communication network 110 may support wireless and cellular telephones and personal email devices (e.g., ) protocol used. Such protocols may include, for example, GSM, GSM+EDGE, CDMA, quadband, and other cellular protocols. In another example, the remote communication protocol may include Wi-Fi and protocols for making or receiving calls using VOIP, LAN, WAN, or other TCP-IP based communication protocols. When located within the communication network 110, the transmitting device 12 and the receiving device 11 may communicate via a two-way communication path (such as path 13) or via two one-way communication paths. Both the transmitting device 12 and the receiving device 11 may be capable of initiating communication operations and receiving initiated communication operations.

[0030] The transmitting device 12 and the receiving device 11 may include any suitable device for sending and receiving communication operations. For example, the transmitting device 12 and the receiving device 11 may include, but are not limited to, devices including voice assistants (personal assistants, virtual assistants, etc.), such as mobile phone devices, television (TV) systems, smart TV systems, cameras, video cameras, devices with audio and video capabilities, tablet computers, wearable devices, smart home appliances, smart photo frames, and any other device capable of communicating wirelessly (with or without the help of wireless-enabled accessory systems) or via a wired path (e.g., using traditional telephone lines). The communication operations may include any suitable form of communication, including, for example, voice communication (e.g., telephone calls), data communication (e.g., data and control messages, emails, text messages, media messages), video communication, or a combination of these (e.g., video conferencing).

[0031] Figure 2A block diagram of an architecture for a system 100 is shown that is capable of performing task learning for a virtual assistant or intelligent agent using user demonstrations for generating task commands for at least one application using: an electronic device 120 (e.g., a mobile phone device, a TV system, a camera, a camcorder, a device with audio-video capabilities, a tablet computer, a tablet device, a wearable device, a smart home appliance, a smart photo frame, smart lighting, etc.), a cloud or server 140, or a combination of an electronic device 120 and a cloud computing (e.g., a shared pool of configurable computing system resources and higher-level services, etc.) or server (e.g., a computer, device, or program that manages network resources, etc.) 140. A transmitting device 120 ( Figure 1 ) and receiving device 11 may both include some or all of the features of electronic device 120. In some embodiments, electronic device 120 may include display 121, microphone 122, audio output 123, input mechanism 124, communication circuit 125, control circuit 126, camera 128, processing and memory 129, intelligent learning (e.g., Figure 3 Processing 130 and / or 131 (for processing on the electronic device 120, on the cloud / server 140, on a combination of the electronic device 120 and the cloud / server 140, communicating with the communication circuit 125 to obtain information / provide its information to the cloud or server 140; and may include any processing for, but not limited to, the examples described below) and any other suitable components. Applications 1 to N 127 are provided and can be accessed from the cloud or server 140, the communication network 110 ( Figure 1 ) etc. to obtain the application, where N is a positive integer equal to or greater than 1.

[0032] In some embodiments, all applications employed by audio output 123, display 121, input mechanism 124, communication circuitry 125, and microphone 122 may be interconnected and managed by control circuitry 126. In one example, a handheld music player capable of streaming music to other tuning devices may be incorporated into electronic device 120.

[0033] In some embodiments, the audio output 123 may include any suitable audio component for providing audio to a user of the electronic device 120. For example, the audio output 123 may include one or more speakers (e.g., mono or stereo speakers) built into the electronic device 120. In some embodiments, the audio output 123 may include an audio component that is remotely coupled to the electronic device 120. For example, the audio output 123 may include a speaker that may be wired (e.g., connected to the electronic device 120 with a jack) or wirelessly (e.g., Headphones or A headset, earphone or earbud coupled to a communication device.

[0034] In some embodiments, display 121 may include any suitable screen or projection system for providing a display visible to a user. For example, display 121 may include a screen (e.g., an LCD screen, an LED screen, an OLED screen, etc.) incorporated into electronic device 120. As another example, display 121 may include a removable display or projection system (e.g., a video projector) for providing a display of content on a surface remote from electronic device 120. Display 121 may be operable to display content (e.g., information about communication operations or information about available media selections) under the direction of control circuitry 126.

[0035] In some embodiments, input mechanism 124 may be any suitable mechanism or user interface for providing user input or instructions to electronic device 120. Input mechanism 124 may take a variety of forms, such as buttons, a keypad, a dial, a click wheel, a mouse, a visual indicator, a remote control, one or more sensors (e.g., a camera or visual sensor, a light sensor, a proximity sensor, etc.), a touch screen, gesture recognition, voice recognition, etc. Input mechanism 124 may include a multi-point touch screen.

[0036] In some embodiments, the communication circuit 125 may be operable to connect to a communication network (e.g., Figure 1 The communication circuit 125 may be operable to interface with the communication network using any suitable communication protocol, such as Wi-Fi (e.g., IEEE 802.11 protocol), High frequency systems (e.g., 900 MHz, 2.4 GHz, and 5.6 GHz communication systems), infrared, GSM, GSM+EDGE, CDMA, quadband and other cellular protocols, VOIP, TCP-IP, or any other suitable protocol.

[0037] In some embodiments, the communication circuit 125 can be operated to create a communication network using any suitable communication protocol. For example, the communication circuit 125 can use a short-range communication protocol to create a short-range communication network to connect to other communication devices. For example, the communication circuit 125 can be operated to use The protocol creates a local communication network to connect the electronic device 120 with Headset coupling.

[0038] In some embodiments, the control circuit 126 can be operated to control the operation and performance of the electronic device 120. The control circuit 126 may include, for example, a processor, a bus (e.g., for sending instructions to other components of the electronic device 120), a memory, a storage body, or any other suitable component for controlling the operation of the electronic device 120. In some embodiments, one or more processors (e.g., in the processing and memory 129) can drive the display and process inputs received from the user interface. The memory and storage body may include, for example, cache, flash memory, ROM and / or RAM / DRAM. In some embodiments, the memory can be specifically used to store firmware (e.g., for device applications, such as operating systems, user interface functions, and processor functions). In some embodiments, the memory can be operated to store information related to other devices with which the electronic device 120 performs communication operations (e.g., saving contact information related to communication operations or storing information related to different media types and media items selected by the user).

[0039] In some embodiments, the control circuit 126 may be operable to perform operations of one or more applications implemented on the electronic device 120. Any suitable number or type of applications may be implemented. Although the following discussion will enumerate different applications, it will be understood that some or all of the applications may be combined into one or more applications. For example, the electronic device 120 may include applications 1 to N 127, including, but not limited to: an automatic speech recognition (ASR) application, an OCR application, a conversation application, a PA (or intelligent agent) app, a map application, a media application (e.g., a photo album app, QuickTime, Mobile Music app, or Mobile Video app), a social networking application (e.g., The electronic device 120 may include an application that is operable to perform communication operations. For example, the electronic device 120 may include a messaging application, an email application, a voicemail application, an instant messaging application (e.g., for chatting), a video conferencing application, a fax application, or any other suitable application for performing any suitable communication operation.

[0040] In some embodiments, the electronic device 120 may include a microphone 122. For example, the electronic device 120 may include a microphone 122 to allow a user to transmit audio (e.g., voice audio) for verbal control and navigation of the applications 1 to N 127 during a communication operation or as a means of establishing a communication operation or as an alternative to using a physical user interface. The microphone 122 may be incorporated into the electronic device 120, or may be remotely coupled to the electronic device 120. For example, the microphone 122 may be incorporated into a wired headset, the microphone 122 may be incorporated into a wireless headset, the microphone 122 may be incorporated into a remote control device, etc.

[0041] In some embodiments, camera module 128 includes one or more camera devices that include functionality for capturing still and video images, editing functionality, communication interoperability for sending, sharing, etc., photos / videos, etc.

[0042] In some embodiments, the electronic device 120 may include any other components suitable for performing communication operations. For example, the electronic device 120 may include a power supply, a port, or an interface for coupling to a host device, an auxiliary input mechanism (e.g., an ON / OFF switch), or any other suitable component.

[0043] Figure 3 A high-level block diagram of an ILS 300 process according to some embodiments is shown. In some embodiments, the high-level block diagram of the ILS 300 process includes an app (or application) 310, a user action 315 (e.g., an interaction with the app 310), an ILS service 320, an event validator 340, an app UI screen image capture service 330, a UI element tree semantic mining tool 360, an event metadata extractor 345, a task event queue 350, and command processing 370, which includes task queue command 371 processing, semantic ontology processing 372, and event handlers / filters 373. In some embodiments, the result of the ILS 300 process is a task command 380.

[0044] In order to Figure 2 In one or more embodiments, the ILS 300 process (task learning system process) is used to enable voice commands, gestures, etc. on applications 1 to N 127, etc.). In one or more embodiments, in the ILS 300 process, a user demonstrates a task on an application (e.g., Figure 2Applications 1 to N 127, etc.). The ILS 300 process repeatedly performs tasks on the application using different sets of input values ​​as needed based on a one-time or one-time process. The input values ​​can come from user voice commands, text input, gestures, conversations and / or contexts. In some embodiments, under the ILS 300 process, voice commands can be implemented by the application even if the application does not have voice assistance capabilities. Learning from demonstrations significantly reduces the developer's work, time, cost and resources to implement specific tasks to support the PA's capabilities. Any user can easily use the mobile application to generate tasks through demonstrations as normal, without the need for a learning curve.

[0045] In some embodiments, the ILS 300 process generates task commands to implement tasks in response to a series of user actions demonstrated by a user on one or more applications. Each user gesture or action generates one or more events, which may contain and / or reflect various data, states (e.g., specific conditions at a specific time) and / or contexts (e.g., time, date, interaction or use of electronic devices / applications, places, calendars, events / appointments, activities, etc.) in the task execution sequence. In one or more embodiments, the mobile application (e.g., smart learning process 131, smart learning process 130, or a combination of smart learning processes 130 and 131, etc.) used for the ILS 300 process is highly dynamic, state representative, and context-aware. Not only may one or more user actions generate one or more events, but also devices and / or reliable services may generate event sequences as well. In some embodiments, the task execution sequence is determined by filtering unwanted events based on a semantic understanding of one or more actions. The ILS 300 process is highly context-aware in learning tasks, where context helps to understand states and event priorities to generate task commands that can be executed with high accuracy.

[0046] In some embodiments, the ILS 300 process captures events (e.g., events generated in conjunction with user actions 315) based on application user interface (UI) interactions (e.g., using appUI screen image capture service 330), where the ILS 300 process can obtain detailed information for one or more UI elements. System, app UI screen image capture service 330 is implemented based on accessibility services, which enables listeners of UI actions as background services and configures which user actions can be captured or ignored in response to user interactions with applications. When a user performs a UI action, a callback method is generated for the app UI screen image capture service 330, where information about the user action is captured as an event. The event contains detailed information such as user action type (e.g., single click, long click, selection, scroll, etc.), UI element, UI element type, position (x, y), text, resource id and / or UI parent-child elements, etc. The app UI screen image capture service 330 returns any special metadata assigned to the UI element by the application to access specific application packages, configuration elements such as enablement, export and processing, etc. Each user interaction can be associated with different amounts of systems / services, such as notifications, batteries, calls, messages, etc., and the processing of the systems / services has a high system priority. The ILS 300 process will collect information such as application status, permissions, services and events that can be generated by the application based on various actions to determine the context. Some of the events are filtered (eg, via event handler / filter 373 ) or the sequence of the events is altered to generate task commands 380 .

[0047] In some embodiments, the application state is based on the current application screen (e.g., The current application screen is loaded from multiple states such as create, restart or start. The application can start in any state as described above, depending on the previous user interaction with the application. If the application is not executing in the background, the application is called as a new application process, that is, the 'create' state, which is the default state of the application, in which all UI elements and resources (databases, services, receivers) are instantiated into memory. If the user interacts with the application before presenting a new task, the application remains executed in the background, which is the 'active saved instance' state or the 'persistent' state. In some applications (for example, In an application loaded from a saved state, it means that the user previously operated via some actions, but the application was not closed. Since the application is still in the memory, the screen data and / or application data due to user actions or system responses are temporarily saved. When the user calls or opens the application next time, the application opens the previously closed state screen, which has all the saved application data, instead of the default screen activity. The data can be visible input implemented by the user or invisible data due to user actions (e.g., click, select, etc.). When the application is opened from a saved state, the application triggers an event called 'onActivitySaveInstanceState', which is important for identifying the application state while handling the window screen state. In one or more embodiments, the ILS 300 process learns new tasks via demonstration to complete user specific commands, generates executable task commands 380, filters unwanted actions / events generated by the system / service relative to context and state, provides dynamic sequencing of actions and partial task execution through state and context prediction, provides version compatibility of tasks that can be executed on different versions of applications, etc. In some embodiments, the learned tasks are not limited to a single application, and the ILS 300 process can be applied to multiple applications to implement a single task.

[0048] In some embodiments, the ILS 300 process functions to generate a task event queue 350 in the form of task commands 380 for the system (e.g., Figure 2 System 100, Fig.14 In one or more embodiments, the task command 380 includes a sequence of semantic events, where each semantic event includes data that provides a target UI element, an input value, an event order priority, whether it is a system feedback event, etc. There are challenges in constructing the task command 380 from an application because sufficient information is required for each user action such as a voice, gesture, or input event so that the task command 380 can be executed on one or more devices (e.g., Figure 2 electronic device 120) and various versions of applications.

[0049] Some challenges in generating task commands 380 are dynamic state and event handling. Mobile system applications (developed on mobile OS platforms such as In some embodiments, the ILS 300 process can understand each system state and perceive the context when implementing task learning. All events generated by user actions are verified based on the application package of the event, the type of event and / or the context of the event (which is dynamic based on multiple data elements). When an event is generated by the currently active application package of the learning task, the event is processed as valid, and other events will be filtered out as irrelevant events, or system or external application events running in the background to generate task commands 380.

[0050] In some embodiments, the UI of mobile application 310 is constructed in a manner similar to a document object model (DOM) tree. Due to the small screen size (e.g., mobile device), a task may span several UI screens, and each screen contains tree-like UI elements. UI elements are classified based on their functions, visibility, and accessibility, such as buttons, text views, text editing, list views, etc. In order to build a single task, fewer or more UI operations are included in a single screen or in multiple screens. The user action 315 that generates an event is processed in multiple stages to generate a task command 380. For example, based on a package, data, and semantics, an event is verified by an event validator 340, and the event is filtered by an event processor / filter 373. Events are prioritized to build executable actions and feedback events. In some embodiments, such stages are very important for processing event filtering.

[0051] In some embodiments, the ILS 300 process generates executable task commands 380 without following conventional development processes, such as coding, compiling, and debugging. The ILS 300 process sequentially understands each user UI action with the help of available data about each user action 315. The ILS 300 process pre-processes the user action 315 or the generated event by semantically classifying a series of UI elements, such as text and / or icon images, and data of such UI elements, such as type, description, possible actions, etc., based on the context of the system, service, and task. The ILS 300 process generates task commands 380, which are a collection of actions similar to user actions and feedback events that can be automatically executed, which helps to ensure that a completion state is achieved for each action running on the application 310 when the user initiates a task with multiple data values ​​using a voice command.

[0052] In some embodiments, ILS 300 processes learning tasks as follows. After understanding that the user will demonstrate the task by interacting with application 310, ILS service 320 begins learning the task by capturing application package information associated with application 310 including data of the package such as name, license, data provider, services, activities (UI screens), hardware features, etc. In the OS, such package information can be obtained using platform APIs such as PackageManager().getPackageInfo). The app UI screen image capture service 330 captures the current active screen of visible UI elements visible to the user, and extracts the text and icon images of such UI elements. The user action 315 records the user UI action, extracts the user action 315 information (e.g., events, states, text, images or icon data, etc.) and the UI element tree semantic mining tool 360 generates a hierarchical tree of UI elements based on how the elements are placed in each screen by extracting each element data into text, resource id, location, description and parent-child elements. The event metadata extractor 345 extracts metadata from the user action 315 information, such as the UI element type, the location of the UI element, the text or image that appears, the hidden description, the timestamp, etc. The task event queue 350 queues the user action event information for command processing 370. The command processing 370 obtains information from the task event queue 350 and the UI element tree semantic mining tool 360. The event handler / filter 373 cleans up system, service and out-of-context events to make the task more precise and accurate. System events include notifications, warnings, hardware state change events such as Wi-Fi connection changes, and service events of other applications running in the background. Most non-contextual events are inactive package events that are not the direct result of one or more user actions performed on the application, generated by the background services of the system or other packages / applications. For example, when learning the "Send SMS" task, SMS is received when the user demonstrates the action. The received SMS is marked as an event out of context. The semantic ontology 372 performs discrimination on each data element of the event generated from the user action, such as the text, image and type of the UI element, and possible input data to generate a semantic ontology of all user actions or events. If some of the data elements do not have enough information for semantic understanding, the task queue command processing 371 adds additional data / descriptions based on voice, gesture or text input about such specific data elements. The event validator 340 verifies each user action 315 based on the application package generated by each event, the type of event, and the current context event (which is dynamic based on data from user actions or system feedback events). If the ILS 300 process does not understand a specific user action, the user is prompted to request help. In one embodiment, command processing 370 generates subsequent event command executable tasks in a 'protocol buffer' format, which is a language-independent and platform-independent mechanism for serializing structured data that can be run on applications. Protocol buffers can be encoded and decoded in byte code using a variety of computer languages. It helps learn tasks from an application on one device and execute the tasks on another device.In some embodiments, no voice data or text data is required to be input by the user when training the task. Almost all data is collected from the source application of the demonstration task to understand and semantically classify each user action and element performed by the user.

[0053] Figure 4 An example process 400 for an application for user task demonstration according to some embodiments is shown. The example process includes a series of screens 410, 411, 412 and UI element activities 440, 441, 442, 443, 444, 445 and 446. According to some embodiments, the process 400 uses the ILS 300 to process ( Figure 3 ) to learn user tasks. Task learning generates task commands 380 based on a series of actions demonstrated by the user on a single application or multiple applications. Each user interaction (e.g., Figure 3 The user action 315 of the embodiment of the present invention generates one or more events containing various data, states and contexts, which helps to classify the user action in order to generate a task execution sequence. UI element activities 440, 441, 442, 443, 444, 445 and 446 are demonstrated on each screen 410, 411 and 412 via user actions. For these user actions, the user enters some text (UI element activities 440, 441, 443 and 445), and clicks a button (UI element activities 442, 444 and 446) to navigate from one screen to another. In the background, the ILS 300 processes the data captured by the user action and processes the data semantically. If some of the user action or UI element activity does not have enough information to process semantically, the screen 420 prompts the user to enter via voice and / or text. After the user completes the interaction with the screen 420, the events generated by the user actions 430, 431 and 432 are processed by the ILS 300 process to form a task command 380.

[0054] Figure 5 A high-level component flow 500 is shown that illustrates the ILS 300 process using the event handler / filter 373 to filter unwanted system and service events according to some embodiments. Figure 3 ) environment is highly dynamic and context-aware. When a user performs a specific action (e.g., interacting with app 310, typing, voice command, etc.), the ILS 300 process generates a lot of information corresponding to the user action (e.g., response to the action, information prompt, etc.). Each user action can generate one or more events. The ILS 300 process effectively works to clear one or more events or reorder the events based on priority. Some of the events are used to verify the action through the ILS service 320 process.

[0055] In some embodiments, process 500 includes extracted event metadata information 510, user actions 315, and system generated events. As the user demonstrates a task, information including active window state 520, text editing focus 521 (when the user attempts to enter some text), and saved state (such as navigating to the next screen or submitting input data via a button click 523) is obtained. The ILS 300 processing in process 500 includes event pre-processing 530, which is a component that checks all events for captured necessary data (such as text and images), and saves all events in the original event queue in memory to create a temporary copy for subsequent processing. If any information is missing or insufficient, process 500 utilizes specific elements (e.g., in Figure 4 The ILS 300 processing in process 500 also includes: event filtering 540 for systems and services, which processes and filters unwanted service and system events; text and image recognition 550, which generates a semantic ontology for each event generated by a user action; and context / state prediction 560, which recognizes the context / state prediction once it is formed into a task command 380 ( Figure 3 and Figure 4 ) what kind of context is needed to perform those actions, such as permissions, hardware requirements, other service interfaces, etc. In one embodiment, task command generator 570 combines information from event pre-processing 530, event filtering for systems and services 540, text and image recognition 550, and context / state prediction 560 with user actions 315 and converts the information into a protocol buffer format (which can be executed and deployed to any system to generate task commands 380).

[0056] Figure 6A high-level flow 600 of an ILS 300 process according to some embodiments is shown, which performs dynamic event sorting by filtering events based on packages, data such as icons, text, and descriptions, and semantics, and prioritizing events and text / image semantic ontologies. When a task is learned on one version of an application using a specific amount of mandatory data to complete the task, the learned task may not be implemented on other versions of the application that are unknown to the ILS 300. The ILS 300 works based on a semantic understanding of the task and the resource ID of the task. For example, if the application screen changes or adds more UI elements, the ILS 300 uses a "resource ID" that is unique to the UI element to implement the task. Even if the UI element changes, as long as the resource ID does not change, the task can still be executed. If the resource ID of the UI element also changes, the ILS 300 matches the semantically close UI element to the learning element to implement the action. For example, a task of sending a message is learned, and the message has a "contact", "text" data element, and a "send" action. In the next version of the application, the UI elements change and are labeled "Recipient," "Message," and an icon button representing a send action. The domain text classifier that generates the semantic ontology 372 detects that the closest element to "Contact" is "Recipient" and the closest element to "Text" is "Message" in the updated version of the application. The ILS 300 uses the semantic ontology to generate task commands 380, which helps perform the task for the previous and next versions.

[0057] In some embodiments, the ILS 300 processes events generated by semantically understanding the priority. Process 600 obtains user 605 actions 315 and uses event handlers / filters 373 ( Figure 3 ) to filter system / service events 610. The event handler / filter 373 works to identify events generated by user actions. The event handler / filter 373 processes events at multiple validation and filtering levels based on event information, such as event type 620, event priority 621, and packet type. Some similar events can be generated based on a single action. For example, when a user touches a text editing UI control, the system will generate two events, such as "event_click" and "event_focus". However, only one event is needed to perform the user touch action again. Therefore, the event handler / filter 373 filters out "event_focus" and retains "event_click". In some applications (for example, and ), there may be a lot of dynamic data loaded when the application starts. Dynamic data such as 'event_scroll' and 'event_select' are generated without any user action. These events are generated by the application package, but are not context-related. Therefore, these events are filtered out and subsequent events change their position in the task queue. Context prediction is the state of the application process when the application is active and running. When the application is running, not all resources used to complete the task will run in the same context. For example, databases, hardware resources, access to application controls / UI elements may have their own application context. Context prediction 560 identifies whether a specific resource is running in the application context or in the system context, which helps to keep the task execution complete by obtaining the appropriate context. Semantic ontology processing 372 is performed to mine all elements present in the current active screen of the application. Semantic ontology processing 372 generates complete semantic data for all elements from text, resource ID, description, location, and actions that can be performed on the screen text and icon image of the UI element.

[0058] Figure 7 An example process flow diagram 700 for performing ILS task learning on a single application according to some embodiments is shown. In some embodiments, the ILS 300 processes ( Figure 3 ) learns how to perform a task via demonstration; thereby enabling an average end user to teach an intelligent agent to perform a task he / she wants / preferred. In some embodiments, the voice command utterance 705 provided by the user is processed by a natural language understanding (NLU) engine 710 (e.g., of a PA, virtual assistant, etc.). The output from the NLU engine 710 provides the intent and slot information of the voice command in the intent slot 715. An intent is an intention conveyed by the end user to perform some specific task. These intents help identify applications that can be used to perform the task for which the intent was given, such as "book a flight," "order a pizza," "play a video," etc. To implement each intent, an electronic device (e.g., Figure 2 electronic device 120) from available applications (e.g., Figure 2 A slot is a variable value that performs a task that the user intends. In the example utterance of "book a flight to New York on January 12", the part "book a flight" is the intent and "New York" and "January 12" are slots.

[0059] In some embodiments, the user (task learning) determination 720 provides information about which application among the available applications in the electronic device is suitable for performing the task. If the specific application is not installed in the electronic device, the user is prompted to install the application, or if more than one application exists, the user is prompted to select an appropriate application for the application 310 (see also Figure 3 ) learning task. It is also possible to select a specific application for the user based on application package information, usage history, ranking of applications, etc. Various screens 730 and user actions 315 of the electronic device are used to determine user events 716 caused only by user actions, and system events 717 that occur together with user actions. For example, system events 717 may include timely system events such as notifications, or events from other applications / services, low battery, keyboard input, etc. All of these events are added to the semantic event queue or raw event queue 350 together. Event processor / filter 373 filters and prioritizes events in the semantic event queue 350, performs context / state prediction and adds the output to the main task queue 371. Semantic ontology processing 372 is semantically verified based on a set of predefined values ​​constructed manually. Possible values ​​are extracted from various contexts for UI elements in multiple applications. In one example, the send button has possible semantic meanings, such as posting, tweeting, submitting, completing, etc. These values ​​are constructed manually or algorithmically by collecting data from a large group of applications. The constructed data values ​​are used to validate the slot values ​​given by the NLU engine 710. In one example, in the voice command "Send SMS to Mom that I'll be late," there are two slot values ​​required to complete the task as Mom (a phone number from a contact) and the text "late." The ILS action validator 340 verifies whether the user demonstrates the slots required to perform the task at any time or whether the system processes the slots correctly. At box 740, the user is prompted to obtain additional data or perform user verification. Once the task is verified, the task command generator 570 generates a task command 380. In some embodiments, the ILS 300 processing for a single application can be used for various tasks, such as booking a flight, finding a place (e.g., a restaurant, a place, a hotel, a store, etc.), etc. For example, Different tasks may be implemented, such as "book a hotel," "book a flight," "book a car," and "find nearby attractions." If multiple users perform different tasks on the same application, additional user authentication may be required to implement the tasks to avoid ambiguity in task execution.

[0060] Figure 8Another example process flow diagram 800 for performing ILS task learning on more than one application according to some embodiments is shown. In some embodiments, the voice command utterance 705 is generated by the NLU engine 710 (see also Figure 7 ) processing. As described above, the output from the NLU engine 710 provides intent and slot information in the intent slot 715. As described above, the user (task learning) determines 720 the recognition application 310 (see also Figure 3 ). Various screens 730 and user actions 315 of the electronic device are used to determine user events 716 including events caused only by user actions, and system events 717 occurring together with user actions. All of these events are added to the semantic event queue or raw event queue 350 collectively. The event processor / filter 373 processes the events, filters, prioritizes, and performs context / state prediction. An event queue 871 with subtasks (e.g., 2 or more) is formed based on the applications involved in the entire task demonstration process. If, for example, the user interacts with application 1 and application 2 during the demonstration process, the action from application 1 is formed as a subtask, and the action from application 2 is formed as a second subtask. These actions can be classified based on the application package ID using a package manager (as described above), and the output is added to the main task queue 371. As described above, the semantic ontology processing 372 semantically verifies elements based on a predefined set of values. The ILS action validator 372 verifies that the user demonstrates the slots required to implement the task at any time or that the electronic device correctly handles the slots. At block 740, the user is prompted to obtain additional data or perform user verification. Once the task is verified, the task command generator 570 generates a task command 380. In some embodiments, the ILS 300 process for multiple applications can be used for various tasks, such as finding a specific place in a specific area (e.g., a specific type of place in a specific location (e.g., finding Korean BBQ restaurants in Palo Alto, California), sharing photos of a specific trip in the cloud, etc.).

[0061] Fig.9A , Fig. 9B , Fig. 9C and Fig.9D According to some embodiments, Example of learning posting on the application. In this example embodiment, the ILS 300 processes ( Figure 3 ) are learning about The "Add New Tweet" task of the application. Screen 1 910 shows "Create Tweet" and Screen 2 920 is used to add messages and tweets. This task is learned in two screens in the application, the first screen (Screen 1 910) records the user launching a 'new tweet', and the second screen (Screen 2 920) captures the user input, and the action 'button click' to perform the 'tweet' is captured in Screen 3 930. User actions are tracked programmatically through each event, data, and UI element. These corresponding events are dynamic events and state change events, which result in the generation of multiple supplementary events and notifications. Typically, for each task, the total number of events generated is almost three times the actions performed by the user, which gives the event in-depth insights into how it interacts with UI elements. The ILS 300 process captures UI element categories, events generated by user actions, text, element coordinates, and each active screen, each UI element image to semantically understand each user action. The ILS 300 process understands each active screen that is called as part of the task learning, and mines the data of the UI element tree for each active screen. The semantic tree data is stored in the task event queue 350 ( Figure 3 ) in command processing 370( Figure 3 ) for further semantic ontology processing. In the example, screen 3 930 provides recipient information of the tweet, and screen 4 940 displays the tweet.

[0062] In some embodiments, a temporary input prompt is issued to obtain additional data. Natural language instructions are a common means for humans to teach each other new tasks. In order for a PA or intelligent agent to learn from such instructions, a significant challenge is basic training - the agent needs to extract semantic meaning from the teachings and associate the semantic meaning with actions and perceptions. If a particular UI element does not have enough information, the ILS 300 process further inputs simple natural language instructions to infer data for classifying the semantic meaning. In some embodiments, keyword-based instructions are used to interpret actions, terms, or descriptions to more accurately infer data. As a reminder, the ILS 300 process provides prompts to the screen of the electronic device to obtain information for possible semantic classifications.

[0063] In some embodiments, the ILS 300 process powerfully processes and filters events by understanding the context. One embodiment uses a keymap event classifier, which is a predefined set of values ​​that is manually constructed. The possible values ​​are extracted from various contexts for UI elements in multiple applications. In one example, a send button has multiple possible semantic meanings, such as post, tweet, submit, complete, etc. These values ​​are constructed manually or algorithmically by collecting data from a large set of applications. App UI Screen Image Capture Service 330 ( Figure 3) and UI element tree semantic mining tool 360 are used to understand each of the UI elements and their corresponding actions demonstrated by the user. Event handler / filter 373 ( Figure 3 ) from task event queue 350( Figure 3 ) and uses the keymap event classifier to verify and filter unwanted events. The keymap event classifier classifies system events based on data such as application package, application type and context, and filters out-of-context information about the task learning application. Events such as 'window change', notifications, announcements, etc. are filtered. However, other events generated by the system such as 'window state change' are used in task learning for activity change state or context change state during the task learning process. After processing the events, the next is the semantic ontology 372 ( Figure 3 ) to generate a semantic ontology based on text and images of UI elements captured during user demonstrations.

[0064] In some embodiments, the event-oriented ILS 300 processes various metadata for each different task (by the event metadata extractor 345 ( Figure 3 ) Extraction) works. For example, to clear all system events from the task event queue 350, the ILS 300 processes the package name that identifies the source of the event. For example, in In mobile devices, package names other than "com.samsung.android.app.aodservice", "com.sec.android.app.launcher", and "com.samsung.android.MtpApplication" are classified as system packages and filtered. Example event captured when a tweet is posted on the application.

[0065] ∑1: Event type: TYPE_VIEW_FOCUSED; Event time: 310865; Package name: com.twitter.android; Move granularity: 0; Actions: 0 [Category name: android.support.v7.widget.RecyclerView; Text: [●]; Content description: empty; Item count: 1; Current item index: -1; IsEnabled: true; IsPassword: false; IsChecked: false; IsFullScreen: false; Scrollable: false; Before text: empty; Start index: 0; End index: 0; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: empty]; Record count: 0

[0066] ∑2: Event type: TYPE_VIEW_FOCUSED; Event time: 310916; Package name: com.twitter.android; Mobile granularity: 0; Actions: 0 [Category name: android.support.v7.widget.RecyclerView; Text: [●]; Content description: Home timeline list; Item count: 6; Current item index: 0; IsEnabled: True; IsPassword: False; IsChecked: False; IsFullScreen: False; Scrollable: False; Before text: Empty; Start index: -1; End index: -1; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: Empty]; Record count: 0

[0067] ∑3: Event type: TYPE_VIEW_SCROLLED; Event time: 311238; Package name: com.twitter.android; Move granularity: 0; Actions: 0 [Category name: android.support.v7.widget.RecyclerView; Text: []; Content description: empty; Item count: 1; Current item index: -1; IsEnabled: true; IsPassword: false; IsChecked: false; IsFullScreen: false; Scrollable: false; Before text: empty; Start index: 0; End index: 0; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: empty]; Record count: 0

[0068] ∑4: Event type: TYPE_WINDOW_STATE_CHANGED; Event time: 312329; Package name: com.samsung.android.MtpApplication Mobile granularity: 0; Actions: 0 [Category name: com.samsung.android.MtpApplication.USBConnection; Text: [MTP application]; Content description: empty; Item count: -1; Current item index: -1; IsEnabled: true; IsPassword: false; IsChecked: false; IsFullScreen: false; Scrollable: false; Before text: empty; Start index: -1; End index: -1; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: empty]; Record count: 0

[0069] ∑5: Event type: TYPE_VIEW_CLICKED; Event time: 322105; Package name: com.samsung.android.MtpApplication Mobile granularity: 0; Action: 0 [Category name: android.widget.Button; Text: [OK]; Content description: empty; Item count: -1; Current item index: -1; IsEnabled: true; IsPassword: false; IsChecked: false; IsFullScreen: false; Scrollable: false; Before text: empty; Start index: -1; End index: -1; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: empty]; Record count: 0

[0070] ∑6: Event type: TYPE_WINDOW_STATE_CHANGED; Event time: 322170; Package name: com.twitter.android; Mobile granularity: 0; Actions: 0 [Category name: com.twitter.app.main.MainActivity; Text: [Home]; Content description: empty; Item count: -1; Current item index: -1; IsEnabled: true; IsPassword: false; IsChecked: false; IsFullScreen: true; Scrollable: false; Before text: empty; Start index: -1; End index: -1; ScrollX: -1: ScrollY: -1; MaxScrollX: -1; MaxScrollY; -1; Add count: -1; Remove count: -1; Package data: empty]; Record count: 0

[0071] In this example, event ∑2 will be filtered because ∑1 and ∑2 are duplicates, which is determined based on the following parameters: the type of event, package name, and category name, etc. (for example, event type: TYPE_VIEW_FOCUSED; package name: com.twitter.android; category name: android.support.v7.widget.RecyclerView). Both events are generated by the system when the user opens the 'twitter' application. Event ∑3 will not be filtered, which is determined based on the package name and the fact that there are no duplicates. Events ∑4 and ∑5 are system events, which is determined by the package name (for example, package name: com.samung.android.MtpApplication, which is a system event with USB connectivity). ∑6 is a master event and is not filtered. In this example, it should be noted that a new screen is opened from an application with the event type: TYPE_WINDOW_STATE_CHANGED according to the package name: com.twitter.android.

[0072] In some embodiments, the event handler / filter 373 may be system specific, such as and The OS is customized to specifically identify various system events and application events for processing and filtering based on metadata information.

[0073] In some embodiments, the task file (task command 380) is a 'byte code file', which includes a series of user actions, metadata information, and semantic data for understanding the classification of each UI element. Metadata information helps to implement tasks on the same or different versions of applications. When the application version changes, the UI elements may also change. In one example, a task is learned for an old version of the application, and the user demonstrates an action on, for example, a 'button'. In the new version of the application, the button UI element is replaced by the 'image button' UI element, so that the source and target UI elements are not the same. However, the system can still implement the button click action by mapping semantic information using text, description, and icon images (if available). In some embodiments, the task file can be a 'json' data file including each action and the feedback action that acts on the application. Typically, JSON files are open formats, and anyone can use malware or robots (bot) to change the action sequence. In this case, the result of the task execution may be unpredictable and harmful to the user because of unpredictable and unwanted actions. In some embodiments, the JSON file is converted to a protocol buffer (with a .pb file extension) file format, which is platform independent, uses less memory, is secure, and is faster than using a JSON file. Protocol buffers are a way to serialize structured data in an arbitrary format. It is useful for communicating or storing data with each other wired when developing programs.

[0074] Fig. 10A , Fig. 10B and Fig. 10C According to some embodiments, Let’s learn an example of flight booking on an app. Booking flight tickets is a popular activity that can be found on various apps such as In one embodiment, ILS 300 processes ( Figure 3 ) learns tasks from applications in response to voice commands, and executes the same application when the same voice command is received. Airline ticket booking can be a repetitive task. Typically, each user searches different applications, adjusting dates and times to get the best deal. Each user performs the task of comparing data for flights with the same date and destination using various applications. In another embodiment, the learned tasks can be applied to different applications in the same category. For example, the first screen 1010 uses user input to start a search. The second screen 1020 shows the user starting to type in a destination, with temporary results shown for selection. The third screen 1030 shows that the data typed by the user will be used to search for flights. The ILS 300 processes ( Figure 3 ) Repeat tasks with multiple dates and varying parameters.

[0075] Fig.11A , Fig. 11B and Fig. 11C According to some embodiments, Example of searching for restaurants on an application. When a user is driving or out and about, searching for restaurants / venues / hotels / stores, etc. may occur frequently. In some embodiments, the ILS 300 processes ( Figure 3 ) can be obtained from an application (e.g. ) learns the task of finding restaurants, and can later perform the learned task while the user is in a moving car. The ILS 300 process will not only perform and find the results, but will also read them to the user based on the selection (e.g., through the car audio system, through the electronic device 120 ( Figure 2 ) etc.). In some embodiments, this task can be repeatedly performed by changing the type of cuisine for nearby locations to find the best restaurants and / or hotels. The ILS 300 process effectively learns this task to perform on various combinations of cuisines and locations. In some embodiments, when learning a task, the ILS 300 process determines the total amount of input values ​​required to complete the task based on how many input UI elements are present in a given sequence of actions. Input elements may include text edits, list boxes, drop-down lists, etc. Therefore, when performing a task in the absence of any such input elements, the ILS 300 process prompts (e.g., text prompts, voice or sound prompts, etc.) the user to obtain the corresponding input to perform the task. When the user directs the performance of the task with voice commands, the NLU engine (e.g., Figures 7 and 8 The NLU engine 710 of the embodiment will identify how many input values ​​(such values ​​are called 'slots') are present in the voice command. For example, in the above description In the application example, the user voice command is "post a tweet that I'm on vacation". This voice command has an input value (slot) "I'm on vacation". If the user voice command is "post a tweet", there is no input value posted in text. Therefore, when implementing such a task, the system will prompt to obtain the input value to be posted.

[0076] In this example, the first screen 1110 is the starting point for using the application. The second screen 1120 shows a search page using Korean BBQ as the cuisine type and Palo Alto, California as the desired location. The third screen 1130 shows the results of the search criteria in the second screen 1120.

[0077] Fig. 12A , Fig. 12B and Fig. 12CAn example of sharing travel pictures from an album to a cloud service according to some embodiments is shown. In this example, two users are participating in a conversation using a messaging (or chat) application, in which the first user sends a request to the second user to "share Los Angeles travel pictures in the cloud" (i.e., share photos of the Los Angeles trip from the album app to the cloud environment or app). For the second user to perform the task, she must leave the messaging application, search for pictures from the album application, and then use a method such as link, Share pictures with cloud services such as Drive and restart the session. ILS 300 Processing ( Figure 3 ) learns the task of sharing a picture from the photo album to the cloud service, and can automatically perform the task next time without leaving the session. The first screen 1210 shows that the first user and the second user are using the messaging application. The second screen 1220 shows the photo album application opened for selecting a picture / photo. The third screen 1230 shows the cloud application for sharing the selected picture.

[0078] Fig.13 1300 for learning a task according to some embodiments. In some embodiments, in block 1310, the process 1300 obtains information related to a task performed by an electronic device (e.g., Figure 2 The electronic device 120, Fig.14 at least one application (e.g., Figure 3 In block 1320, process 1300 records a user interface (e.g., using a user interface) for at least one application. Figure 2 124) interaction (e.g., Figure 3 In block 1330, process 1300 extracts second information (e.g., at least one of system information or service information) from the sequence of user interface interactions. In block 1340, process 1300 (e.g., using Figure 3 The event processor / filter 373 of the embodiment filters at least one of (unwanted) events or actions from the second information based on the first information obtained for each detail. In block 1350, the process 1300 performs a discrimination (e.g., Figure 5 550) to (e.g., via Figure 3 In block 1360, process 1300 generates a task command (e.g., Figure 3The task command 380) may be an executable sequential event task command of at least one application.

[0079] In some embodiments, the second information is processed to understand the context and state. In some embodiments, process 1300 may include extracting metadata from the user interface interaction (e.g., using Figure 3 The metadata and task information can be retrieved from the task event queue (e.g., Figure 3 The task command may be queued in a task event queue 350 of the embodiment of the present invention. Events, actions or combinations thereof may be further filtered from the second information based on the task event queue. In some embodiments, the task command may be configured to be executed on at least one application to perform the task.

[0080] In one or more embodiments, process 1300 may include adding additional voice or text data to each element of the sequence of user interface interactions (e.g., Figure 4 Additional data 420) to obtain additional metadata (e.g., using event validator 340).

[0081] In some embodiments, process 1300 may include pre-processing user interface interactions by classifying text, images, and user interface elements based on systems, services, and data for the services out of context (e.g., using Figure 5 Event preprocessing 530).

[0082] In some embodiments, process 1300 may also include a task command being formulated to be repeatedly executed on at least one application based on a user voice command, and the task command includes a semantic event sequence using logical data, which implements the task on the at least one application when executed.

[0083] Fig.141400 includes one or more processors 1411 (e.g., ASICs, CPUs, etc.), and may also include an electronic display device 1412 (for displaying graphics, text, and other data), a main memory 1413 (e.g., random access memory (RAM), cache devices, etc.), a storage device 1414 (e.g., a hard drive), a removable storage device 1415 (e.g., a removable storage drive, a removable memory, a tape drive, an optical drive, a computer-readable medium storing computer software and / or data), a user interface device 1416 (e.g., a keyboard, a touch screen, a keypad, a pointing device), and a communication interface 1417 (e.g., a modem, a wireless transceiver (such as Wi-Fi, a cellular system), a network interface (such as an Ethernet card), a communication port or a PCMCIA slot and card).

[0084] The communication interface 1417 allows software and data to be transferred between the computer system and external devices via the Internet 1450, mobile electronic devices 1451, servers 1452, networks 1453, etc. The system 1400 also includes a communication infrastructure 1418 (e.g., a communication bus, crossbar, or network) to which the aforementioned devices 1411 to 1417 are connected.

[0085] The information transmitted via the communication interface 1417 may be in the form of signals, such as electronic, electromagnetic, optical or other signals capable of being received by the communication interface 1417 via a communication link, which carries the signals and may be implemented using wire or cable, optical fiber, a telephone line, a cellular telephone link, a radio frequency (RF) link and / or other communication channels.

[0086] In electronic devices (e.g. Figure 2 In one embodiment of one or more embodiments of the electronic device 120), the system 1400 also includes an image capture device 1420, such as a camera 128 ( Figure 2 ), and an audio capture device 1419, such as a microphone 122 ( Figure 2 ). The system 1400 may additionally include application processing or processors such as MMS 1421, SMS 1422, email 1423, social network interface (SNI) 1424, audio / video (AV) player 1425, web browser 1426, image capture 1427, etc.

[0087] In some embodiments, as described above, system 1400 includes intelligent learning processing 1430, which can implement processing similar to that described with respect to the following: ILS 300 processing ( Figure 3), process 400 processing ( Figure 4 ), process 500 processing ( Figure 5 ), process 600 processing ( Figure 6 ), process 700 processing ( Figure 7 ), process 800 processing ( Figure 8 ) and process 1300( Fig.13 ). In one embodiment, the intelligent learning process 1430 and the operating system (O / S) 1429 may be implemented as executable code residing in a memory of the system 1400. In another embodiment, the intelligent learning process 1430 may be provided in hardware, firmware, etc.

[0088] In one embodiment, main memory 1413 , storage device 1414 , and removable storage device 1415 , each by themselves or in any combination, may store instructions for the above-described embodiments, which may be executed by one or more processors 1411 .

[0089] As known to those skilled in the art, the aforementioned example architecture described above (according to the architecture) can be implemented in a variety of ways, such as as program instructions for execution by a processor, as a software module, microcode, as a computer program product on a computer-readable medium, as an analog / logic circuit, as an application-specific integrated circuit, as firmware, as a consumer electronic device, an AV device, a wireless / wired transmitter, a wireless / wired receiver, a network, a multimedia device, etc. In addition, an embodiment of the architecture can take the form of a pure hardware embodiment, a pure software embodiment, or an embodiment containing both hardware elements and software elements.

[0090] One or more embodiments have been described with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to one or more embodiments. Each frame of such diagram / chart or its combination can be implemented by computer program instructions. The computer program instructions produce a machine when provided to a processor so that the instructions executed by the processor create a tool for implementing the function / operation specified in the flowchart and / or block diagram. Each frame in the flowchart / block diagram can represent the hardware and / or software module or logic implementing one or more embodiments. In an alternative embodiment, the functions marked in the frame may not occur in the order marked in the figure, occur simultaneously, etc.

[0091] The terms "computer program medium", "computer usable medium", "computer readable medium" and "computer program product" are generally used to refer to media, such as main memory, auxiliary memory, removable storage drive, hard disk installed in a hard drive. These computer program products are tools for providing software to computer systems. Computer readable media allow a computer system to read data, instructions, messages or message packets, and other computer readable information from a computer readable medium. For example, a computer readable medium may include non-volatile memory, such as a floppy disk, ROM, flash memory, disk drive memory, CD-ROM, and other permanent storage bodies. For example, this is useful for transferring information (such as data and computer instructions) between computer systems. Computer program instructions can be stored in a computer readable medium, which can instruct a computer, other programmable data processing device, or other device to function in a specific manner so that the instructions stored in the computer readable medium produce an article of manufacture, which includes instructions that implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0092] The computer program instructions representing the block diagrams and / or flow charts of this article can be loaded onto a computer, a programmable data processing device, or a processing device to perform a series of operations thereon, thereby generating a computer-implemented process. The computer program (i.e., computer control logic) is stored in a main memory and / or an auxiliary memory. The computer program can also be received via a communication interface. Such computer programs enable a computer system to perform the features of the embodiments as discussed herein when executed. In particular, the computer program enables a processor and / or a multi-core processor to perform the features of a computer system when executed. Such computer programs represent the controller of a computer system. A computer program product includes a tangible storage medium that can be read by a computer system and stores instructions for executing the method for performing one or more embodiments for a computer system to perform.

[0093] Although the embodiments have been described with reference to certain versions thereof; however, other versions are possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the preferred versions contained herein.

Claims

1. A method for learning a task, the method comprising: obtaining first information associated with at least one application executed by the electronic device; recording a sequence of user interface interactions with respect to the at least one application; extracting second information from the sequence of user interface interactions; filtering out at least one of an event or an action from the second information based on the first information; performing identification on each element included in the first information; as well as An executable sequential event task command is generated based on each element and the filtered second information.

2. The method of claim 1, wherein: The first information includes at least one of: a voice command utterance, information from a conversation, or contextual information; and The executable sequence of events task commands are configured to be executed on the at least one application program to perform a task.

3. The method of claim 1, further comprising: extracting metadata from the user interface interactions; Enqueuing the metadata and task information in a task event queue; as well as At least one of events or actions is further filtered from the second information based on the task event queue.

4. The method of claim 3, further comprising: Additional voice or text data is added to each element of the sequence of user interface interactions to obtain additional metadata.

5. The method of claim 1, further comprising: The sequence of user interface interactions is pre-processed by classifying text, images, and user interface elements based on systems, services, and decontextualized data for the task.

6. The method of claim 1, wherein the executable sequential event task command is formulated to be repeatedly executed on the at least one application based on a user voice command.

7. An electronic device, comprising: at least one processor; as well as a memory storing instructions that, when executed by a processor, cause the electronic device to: obtaining first information associated with at least one application executed by the electronic device; recording a sequence of user interface interactions with respect to the at least one application; extracting second information from the sequence of user interface interactions; filtering out at least one of an event or an action from the second information based on the first information; performing identification on each element included in the first information; as well as An executable sequential event task command is generated based on each element and the filtered second information.

8. The electronic device as claimed in claim 7, wherein: The first information includes at least one of: a voice command utterance, information from a conversation, or contextual information; and The executable sequence of events task commands are configured to be executed on the at least one application program to perform a task.

9. The electronic device of claim 7, wherein when the instructions are executed by a processor, the electronic device is further caused to: extracting metadata from the user interface interactions; Enqueuing the metadata and task information in a task event queue; and At least one of events or actions is further filtered from the second information based on the task event queue.

10. The electronic device of claim 9, wherein when the instructions are executed by a processor, the electronic device is further caused to: Additional voice or text data is added to each element of the sequence of user interface interactions to obtain additional metadata.

11. The electronic device of claim 7, wherein when the instructions are executed by a processor, the electronic device is further caused to: The sequence of user interface interactions is pre-processed by classifying text, images, and user interface elements based on systems, services, and decontextualized data for the task.

12. The electronic device of claim 7, wherein the executable sequential event task command is formulated to be repeatedly executed on the at least one application based on a user voice command.

13. A non-transitory computer-readable storage medium storing instructions which, when executed by at least one processor of an electronic device, cause the electronic device to perform the method of one of claims 1 to 6.

Citation Information

Patent Citations

  • Devices, methods, and graphical user interfaces for processing intensity information associated with touch inputs

    CN108351750A