Intelligent Auto Assistant

An intelligent automated assistant simplifies user interaction by integrating software components and invoking services through natural language dialogs, addressing the challenge of inconsistent interfaces in electronic devices and enhancing usability for diverse user groups.

JP2026090448APending Publication Date: 2026-06-02APPLE INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2026-02-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing electronic devices present users with inconsistent and numerous interfaces, making it difficult for novice users, the elderly, and busy or distracted individuals to effectively utilize their diverse functions and online services.

Method used

An intelligent automated assistant that integrates through natural language dialogs, coordinates various software components, and invokes external services to simplify user interaction and streamline device operations, providing a conversational interface that interprets user intent and performs actions without requiring users to manually specify details.

Benefits of technology

The assistant unifies and simplifies the user experience across multiple applications and services, reducing the burden of learning device functionalities and enhancing usability for diverse user groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026090448000001_ABST
    Figure 2026090448000001_ABST
Patent Text Reader

Abstract

To facilitate user interaction with devices and to make local and / or remote services easier to use. [Solution] The method includes the steps of: interpreting speech input utterances to generate a set of candidate speech interpretations; generating a user intent expression 716 for the candidate speech interpretations; generating paraphrases of the user intent expression 716 and presenting them to the user; performing task and dialogue analysis; presenting the task and domain interpretations to the user using an intent paraphrasing algorithm; displaying intermediate results in the form of real-time progress; and formatting the response for an appropriate output modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent system, and more particularly to an intelligent system for classes of applications for intelligent automatic assistants.

Background Art

[0002] Today's electronic devices can access an ever-increasing amount of diverse functions, services, and information via the Internet and from other sources. As many consumer devices, such as smartphones and tablet computers, can execute software applications to perform various tasks and provide various information, the functionality of such devices has been rapidly increasing. Each application, function, website, or feature often has its own user interface and operating paradigm, and many of them can be too much to learn or burdensome for users. Furthermore, for many users, it can even be difficult to find the functionality and / or information available on electronic devices or various websites. Thus, such users may be frustrated, overwhelmed by their quantity, or simply unable to effectively utilize the available resources.

[0003] Particularly, novice users, i.e., those with some impairment, and / or the elderly, busy people, distracted people, and / or those operating a vehicle may have difficulty effectively interfacing with electronic devices and / or effectively using online services. Such users are particularly likely to have difficulty with the large number of diverse and inconsistent functions, applications, and websites that may be available for use.

[0004] Thus, in many cases, existing systems present users with interfaces that are too inconsistent and too numerous to be difficult to use and navigate, often preventing users from effectively using the technology. [Overview of the Initiative]

[0005] According to various embodiments of the present invention, an intelligent automated assistant is implemented in an electronic device to facilitate user interaction with the device and to make local and / or remote services more effectively usable. In various embodiments, the intelligent automated assistant engages with the user through integrated dialogue using natural language dialogs and invokes external services when appropriate to obtain information or perform various actions.

[0006] According to various embodiments of the present invention, the intelligent automated assistant integrates various functions provided by different software components (e.g., for supporting natural language recognition and dialogue, multimodal input, personal information management, task flow management, and orchestration of distributed services). Furthermore, in order to provide the user with an intelligent interface and useful functionality, the intelligent automated assistant of the present invention coordinates these components and services in at least some embodiments. The ability to perform conversational interfaces and information acquisition and subsequent task execution is realized in at least some embodiments by coordinating various components such as language components, dialogue components, task management components, information management components, and / or multiple external services.

[0007] According to various embodiments of the present invention, an intelligent automatic assistant system is configured, designed, and / or operable to provide a variety of different operations, functionalities, and / or features and / or to combine multiple features, operations, and applications of the electronic device on which it is installed. In some embodiments, the intelligent automatic assistant system of the present invention can perform any or all of the following: actively derive input from a user, interpret the user's intent, eliminate ambiguity between conflicting interpretations, request and receive clarification information as needed, and perform (or initiate) an action based on the identified intent. The action is performed, for example, by invoking and / or interface with any application or service available on the electronic device, as well as services available via an electronic network such as the Internet. In various embodiments, the invocation of such external services is performed via an API or by any other suitable mechanism. Thus, the intelligent automatic assistant systems of various embodiments of the present invention can unify, simplify, and improve the user experience with respect to many different applications and functions of the electronic device, as well as services available via the Internet. Accordingly, the user is relieved of the burden of learning the functionalities available on the device and web-connected services, how to interface with such services to obtain what the user requests, and how to interpret the output received from such services. The assistant of this invention can mediate between the user and such various services.

[0008] Furthermore, in various embodiments, the assistant of the present invention provides a conversational interface that the user can find more intuitively and with less burden than conventional graphical user interfaces. The user can interact with the assistant in the form of a conversational dialog using any of the many available input / output mechanisms, such as voice, graphical user interfaces (buttons and links), and text input. The system is implemented using any of the many different platforms, such as device APIs, the web, and email, or any combination thereof. Requests for additional input are presented to the user in the context of such a conversation. Short-term and long-term memory is used to interpret user input in appropriate context, assuming previous events and communications within a given session, as well as history and profile information about the user.

[0009] Furthermore, in various embodiments, contextual information derived from the device's features, operation, or user interaction with an application is used to streamline the operation of that device or other features, operation, or application of another device. For example, an intelligent automated assistant can use the context of a call to streamline the initiation of a text message (for example, to determine that a text message is being sent to the same person without the user having to explicitly specify the recipient). Thus, the intelligent automated assistant of the present invention can interpret commands such as "send him a text message," where "him" is interpreted according to contextual information derived from the current call and / or any features, operation, or application of the device. In various embodiments, the intelligent automated assistant takes into account various available contextual data to determine such information, so that the user does not have to manually re-specify the phonebook contacts to be used, the contact data to be used, and the telephone number to be used to make contact.

[0010] In various embodiments, the assistant further takes external events into consideration and responds accordingly, for example, by initiating an action, initiating communication with the user, providing an alert, and / or modifying a previously initiated action in consideration of the external events. If input is requested from the user, the conversational interface is used again.

[0011] In one embodiment, the system employs additional functionality powered by external services with which the system can interact, based on a set of interrelated domains and tasks. In various embodiments, these external services include web-enabled services and functionality associated with the hardware device itself. For example, in one embodiment where an intelligent automated assistant is implemented on a smartphone, personal digital assistant, tablet computer, or other device, the assistant controls many of the device's operations and functions, such as dialing a phone number, sending a text message, setting a reminder, or adding an event to a calendar.

[0012] In various embodiments, the system of the present invention is implemented to support any one of many different domains. Examples include the following:

[0013] • Local services (including location-specific and time-specific services such as restaurants, movies, ATMs, events, and meeting places) • Personal and social memory services (including activity items, notes, calendar events, and shared links) • Electronic commerce (including online purchases of items such as books, DVDs, and music) • Travel services (including flights, hotels, and tourist attractions) Those skilled in the art will understand that the above list of domains is merely illustrative. Furthermore, the system of the present invention can be implemented in any combination of domains.

[0014] In various embodiments, the intelligent automated assistant systems disclosed herein are configured or designed to include functionality for automating the application of data and services available via the Internet for discovering, finding, and selecting products and services, particularly for purchasing, reserving, or ordering them. In addition to automating the process of using these data and services, at least one embodiment of the intelligent automated assistant system disclosed herein allows for the combined use of several data sources and services at once. For example, an embodiment combines product information from several review sites, checks availability and pricing from multiple vendors, checks location and time constraints, and facilitates finding personalized solutions to the user's problems. Furthermore, at least one embodiment of an intelligent automated assistant system disclosed herein is configured or designed to include functionality for automating the use of data and services available via the Internet to discover, research, and select, in particular to book or learn about, any other sources of social interaction or entertainment that may be found on the Internet, including but not limited to movies, events, performances, exhibitions, shows and tourist attractions, places to go (travel destinations, hotels and other places to stay, landmarks and other places of interest), places to eat and drink (restaurants and bars, etc.), times and places to meet other people, and any other sources of social interaction or entertainment that may be found on the Internet.Furthermore, at least one embodiment of an intelligent automated assistant system disclosed herein is configured or designed to include functionality that enables the operation of applications and services via natural language dialogue provided by a dedicated application having a graphical user interface, which includes: searching (including location-based searching); navigation (maps and directions); database lookup (such as finding businesses or people by name or other characteristics); obtaining weather conditions and forecasts; checking the prices of market items or the status of financial transactions; monitoring traffic or flight conditions; accessing and updating calendars and schedules; managing reminders, alerts, tasks and projects; communication via email or other messaging platforms; and local or remote operation of the device (e.g., making phone calls, controlling lights and temperature, controlling home security devices, playing music or videos). Also, at least one embodiment of an intelligent automated assistant system disclosed herein is configured or designed to include functionality that identifies, generates and / or provides personalized recommendations for activities, products, services, sources of entertainment, time management or other types of recommendation services, which benefit from natural language dialogue and automated access to data and services.

[0015] In various embodiments, the intelligent automatic assistant of the present invention controls many features and operations of an electronic device. For example, the intelligent automatic assistant can invoke services that interface with functionality and applications in the device via APIs or other means to perform functions and operations that would otherwise be initiated using a conventional user interface in the device. Such functions and operations include, for example, setting alarms, making calls, sending text or email messages, and adding calendar events. Such functions and operations are performed as add-on functions in the context of a conversational dialog between the user and the assistant. Such functions and operations are specified by the user in the context of such a dialog or are performed automatically based on the context of the dialog. Thus, it will be understood by those skilled in the art that the assistant is used as a control mechanism to initiate and control various operations of an electronic device, and that this may be used as an alternative mechanism to conventional mechanisms such as buttons or graphical user interfaces. [Brief explanation of the drawing]

[0016] The accompanying drawings illustrate several embodiments of the present invention, and the principles of the present invention as they are described. It will be understood by those skilled in the art that the specific embodiments shown in the drawings are merely illustrative and are not intended to limit the scope of the present invention. [Figure 1] Figure 1 is a block diagram showing an example of one embodiment of an intelligent automatic assistant system. [Figure 2] Figure 2 shows an example of interaction between an intelligent automated assistant and a user according to at least one embodiment. [Figure 3] Figure 3 is a block diagram showing a computing device suitable for realizing at least a part of an intelligent automatic assistant according to at least one embodiment. [Figure 4]FIG. 4 is a block diagram showing an architecture that implements at least a portion of an intelligent automated assistant in a stand-alone computing system, according to at least one embodiment. [Figure 5] FIG. 5 is a block diagram showing an architecture that implements at least a portion of an intelligent automated assistant in a distributed computing network, according to at least one embodiment. [Figure 6] FIG. 6 is a block diagram showing a system architecture that shows several different types of clients and operating modes. [Figure 7] FIG. 7 is a block diagram showing a client and a server that communicate with each other to implement the present invention according to one embodiment. [Figure 8] FIG. 8 is a block diagram showing a part of an active ontology according to one embodiment. [Figure 9] FIG. 9 is a block diagram showing an example of another embodiment of an intelligent automated assistant. [Figure 10] FIG. 10 is a flowchart showing an operation method of an active input derivation component according to one embodiment. [Figure 11] FIG. 11 is a flowchart showing an active type input derivation method according to one embodiment. [Figure 12] FIG. 12 is a screenshot showing some parts of some procedures of an active type input derivation according to one embodiment. [Figure 13] FIG. 13 is a screenshot showing some parts of some procedures of an active type input derivation according to one embodiment. [Figure 14] FIG. 14 is a screenshot showing some parts of some procedures of an active type input derivation according to one embodiment. [Figure 15] FIG. 15 is a screenshot showing some parts of some procedures of an active type input derivation according to one embodiment. [Figure 16]FIG. 16 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 17] FIG. 17 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 18] FIG. 18 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 19] FIG. 19 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 20] FIG. 20 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 21] FIG. 21 is a screenshot showing some parts of some procedures of active type input derivation according to an embodiment. [Figure 22] FIG. 22 is a flowchart showing an active input derivation method for voice input according to an embodiment. [Figure 23] FIG. 23 is a flowchart showing an active input derivation method for GUI input according to an embodiment. [Figure 24] FIG. 24 is a flowchart showing an active input derivation method at the level of a dialog flow according to an embodiment. [Figure 25] FIG. 25 is a flowchart showing a method for actively monitoring related events according to an embodiment. [Figure 26] FIG. 26 is a flowchart showing a multimodal active input derivation method according to an embodiment. [Figure 27] FIG. 27 is a set of screenshots showing an example of various functions, operations, behaviors, and / or other features provided by domain model components and service orchestration according to an embodiment. [Figure 28] FIG. 28 is a flowchart showing an example of a method for natural language processing according to an embodiment. [Figure 29] Figure 29 is a screenshot showing natural language processing according to one embodiment. [Figure 30] Figure 30 is a screenshot showing an example of various functions, operations, actions, and / or other features provided by a dialog flow processor component according to one embodiment. [Figure 31] Figure 31 is a screenshot showing an example of various functions, operations, actions, and / or other features provided by a dialog flow processor component according to one embodiment. [Figure 32] Figure 32 is a flowchart showing the operation method of a dialog flow processor component according to one embodiment. [Figure 33] Figure 33 is a flowchart showing an automated calling and response procedure according to one embodiment. [Figure 34] Figure 34 is a flowchart showing an example of a task flow for a conditional selection task according to one embodiment. [Figure 35] Figure 35 is a screenshot showing an example of the operation of a conditional selection task according to one embodiment. [Figure 36] Figure 36 is a screenshot showing an example of the operation of a conditional selection task according to one embodiment. [Figure 37] Figure 37 is a flowchart showing an example of a procedure for executing a service orchestration procedure according to one embodiment. [Figure 38] Figure 38 is a flowchart showing an example of a service call procedure according to one embodiment. [Figure 39] Figure 39 is a flowchart showing an example of a multiphase output procedure according to one embodiment. [Figure 40] Figure 40 is a screenshot showing an example of output processing according to one embodiment. [Figure 41] Figure 41 is a screenshot showing an example of output processing according to one embodiment. [Figure 42]Figure 42 is a flowchart showing an example of multimodal output processing according to one embodiment. [Figure 43A] Figure 43A is a screenshot illustrating an example of using short-term personal memory components to maintain the dialog context while changing locations, according to one embodiment. [Figure 43B] Figure 43B is a screenshot illustrating an example of using short-term personal memory components to maintain the dialog context while changing locations, according to one embodiment. [Figure 44A] Figure 44A is a screenshot showing an example of using a long-term personal memory component according to one embodiment. [Figure 44B] Figure 44B is a screenshot showing an example of using a long-term personal memory component according to one embodiment. [Figure 44C] Figure 44C is a screenshot showing an example of using long-term personal memory components according to one embodiment. [Figure 45] Figure 45 shows an example of an abstract model of a conditional selection task. [Figure 46] Figure 46 shows an example of a dialogue flow model that facilitates user guidance during the search process. [Figure 47] Figure 47 is a flowchart showing a conditional selection method according to one embodiment. [Modes for carrying out the invention]

[0017] Various technologies will be described in detail with reference to several examples of embodiments, such as those shown in the accompanying drawings. In the following description, many specific details will be described so that one or more embodiments and / or features described and referenced herein may be fully understood. However, it will be apparent to those skilled in the art that one or more embodiments and / or features described and referenced herein may be implemented without some or all of the specific details. In other examples, well-known processing steps and / or structures are not described in detail so as not to obscure some of the embodiments and / or features described or referenced herein.

[0018] One or more distinct inventions are described in this application. Furthermore, many embodiments of one or more inventions described herein are described and presented for illustrative purposes only. The embodiments described are not intended to be limiting in any sense. One or more inventions are broadly applicable to many embodiments, as will be readily apparent from the disclosure. These embodiments are described in sufficient detail so that a person skilled in the art can carry out one or more inventions. It will be understood that other embodiments may be used, and that structural, logical, software, electrical, and other modifications may be made without departing from the scope of one or more inventions. Thus, it will be understood that one or more inventions may be carried out with various modifications and changes. Certain features of one or more inventions are described with reference to one or more specific embodiments or drawings that form part of the present invention. Hereinafter, specific embodiments of one or more inventions are given as examples. However, it should be understood that such features are not limited to those used in the one or more specific embodiments or drawings referenced in the description. The present invention is neither a list of features of one or more inventions that must be present in all embodiments, nor a literal description of all embodiments of one or more inventions.

[0019] The section titles and invention names provided in this patent application are for convenience only and are not intended to limit the present invention.

[0020] Devices communicating with each other do not need to communicate continuously unless otherwise specified. Furthermore, devices communicating with each other may communicate directly or indirectly through one or more intermediaries.

[0021] A description of one embodiment that includes several components communicating with each other does not indicate that all such components are required. In contrast, various optional components are described to illustrate various possible embodiments of one or more inventions.

[0022] Further processing steps, method steps, or algorithms will be described in order, but such processing, methods, and algorithms may be configured to operate in a different order. In other words, any order or sequence of steps described in this patent application does not essentially imply that the steps must be performed in that order. The steps of the processing described may, in practice, be performed in any order. Furthermore, some steps may be performed simultaneously, even if they are not described or shown as occurring simultaneously (for example, one step is described after another). Furthermore, the processing shown in the drawings does not imply that the illustrated processing excludes other changes and variations, that any of the illustrated processing or its steps are necessary for one or more inventions, or that the illustrated processing is preferred.

[0023] When a single device or article is described, it will be readily apparent that two or more devices / articles (whether they work together or not) may be used in place of that single device / article. Similarly, when two or more devices or articles are described (whether they work together or not), it will be readily apparent that a single device / article may be used in place of two or more devices or articles.

[0024] The functionality and / or features of a device may be embodied by one or more other devices not explicitly described as having such functionality / features. Therefore, other embodiments of one or more inventions do not necessarily have to include the device itself.

[0025] The technologies and mechanisms described or referenced herein may be described in the singular for ease of understanding. However, unless otherwise specified, particular embodiments include multiple iterations of the technology or multiple specific examples of the mechanism.

[0026] While this description is made within the scope of intelligent automated assistant technology, it should be understood that various aspects and techniques described herein may be developed and / or applied in other technical fields, including human and / or computerized interaction with software.

[0027] Other aspects of intelligent automatic assistant technology (e.g., those utilized, provided, and / or implemented in embodiments of one or more intelligent automatic assistant systems described herein) are disclosed in one or more of the following references.

[0028] The contents of U.S. Patent Provisional Application No. 61 / 295,774, “Intelligent Automated Assistant,” filed on January 18, 2010, attorney docket No. SIRIP003P, are incorporated herein by reference. • The contents of U.S. Patent Application No. 11 / 518,292, “Method And Apparatus for Building an Intelligent Automated Assistant,” filed on September 8, 2006, are incorporated herein by reference. • The contents of U.S. Provisional Patent Application No. 61 / 186,414, “System and Method for Semantic Auto-Completion,” filed on June 12, 2009, are incorporated herein by reference.

[0029] Hardware architecture Generally, the intelligent automatic assistant technologies disclosed herein are implemented in hardware or in combination of software and hardware. For example, these technologies are implemented in operating system kernels, independent user processes, library packages linked to network applications, or specially built machines or network interface cards. In a particular embodiment, the technologies disclosed herein are implemented in software such as an operating system or in applications running on an operating system.

[0030] At least some of the software / hardware hybrid implementations of the embodiments of the intelligent automated assistants disclosed herein are implemented in programmable machines that are selectively enabled or reconfigured by computer programs stored in memory. Such network devices have multiple network interfaces configured or designed to utilize various network communication protocols. General architectures for some of these machines will become apparent from the descriptions disclosed herein. According to certain embodiments, at least some of the features and / or functionalities of the embodiments of the intelligent automated assistants disclosed herein are implemented in one or more general-purpose network host machines such as end-user computer systems, computers, network servers or server systems, mobile computing devices (e.g., personal digital assistants, mobile phones, smartphones, laptops or tablet computers, etc.), consumer electronic devices, music players or any other suitable electronic devices, routers or switches, or any combination thereof. In at least some embodiments, at least some of the features and / or functionalities of the embodiments of the intelligent automated assistants disclosed herein are implemented in one or more virtual computing environments (e.g., network computing clouds, etc.).

[0031] Referring next to Figure 3, a block diagram is shown illustrating a computing device 60 suitable for realizing at least a part of the features and / or functionality of the intelligent automatic assistant disclosed herein. The computing device 60 is, for example, an end-user computer system, a network server or server system, a mobile computing device (e.g., a personal digital assistant, a mobile phone, a smartphone, a laptop or tablet computer, etc.), a consumer electronic device, a music player or any other suitable electronic device, or any combination thereof or part thereof. The computing device 60 is configured to communicate with other computing devices such as clients and / or servers via a communication network such as the Internet using known protocols for communication, whether wireless or wired.

[0032] In one embodiment, the computing device 60 includes a central processing unit (CPU) 62, an interface 68, and a bus 67 (such as a PCI (Peripheral Component Interconnect) bus). When operating under the control of appropriate software or firmware, the CPU 62 has a role in realizing specific functions associated with the functions of a specially configured computing device or machine. For example, in at least one embodiment, a user's personal digital assistant (PDA) may be configured or designed to function as an intelligent automated assistant system using the CPU 62, memory 61, 65, and interface 68. In at least one embodiment, the CPU 62 is made to perform one or more different intelligent automated assistant functions and / or operations under the control of software modules / components, including, for example, an operating system and any appropriate application software and drivers.

[0033] The CPU 62 includes one or more processors 63, such as Motorola or Intel microprocessors or MIPS microprocessors. In some embodiments, the processors 63 include specially designed hardware for controlling the operation of the computing device 60 (e.g., application-specific integrated circuits (ASICs), electrically erasable programmable read-only memory (EEPROM), and field-programmable gate arrays (FPGAs). In a particular embodiment, memory 61 (non-volatile random access memory (RAM) and / or read-only memory (ROM)) further forms part of the CPU 62. However, there are many different ways in which memory is integrated into the system. The memory block 61 is used for various purposes, such as caching and / or storing data, and programming instructions.

[0034] As used herein, the term “processor” is not limited to an integrated circuit called a processor in the prior art, but broadly includes microcontrollers, microcomputers, programmable logic controllers, application-specific integrated circuits, and any other programmable circuit.

[0035] In one embodiment, interface 68 is provided as an interface card (sometimes called a “line card”). Generally, these interfaces control the transmission and reception of data packets over a computing network and, in some cases, support other peripheral devices used with the computing device 60. The interfaces provided include Ethernet interfaces, Frame Relay interfaces, cable interfaces, DSL interfaces, and Token Ring interfaces, etc. Furthermore, various interfaces are provided, such as USB (Universal Serial Bus), Serial, Ethernet, Firewire, PCI, Parallel, Radio Frequency (RF), Bluetooth™, Near Field Communication (e.g., using Near Field Magnetism), 802.11 (WiFi), Frame Relay, TCP / IP, ISDN, High Speed ​​Ethernet interfaces, Gigabit Ethernet interfaces, Asynchronous Transfer Mode (ATM) interfaces, High Speed ​​Serial Interface (HSSI), Point of Sale (POS) interfaces, and Fiber Optic Distributed Data Interface (FDDI). Generally, such interfaces 68 include ports suitable for communicating with the appropriate medium. In some examples, these interfaces further include a separate processor, and in some examples, include volatile and / or non-volatile memory (e.g., RAM).

[0036] The system shown in Figure 3 illustrates one particular architecture of a computing device 60 for realizing the technology of the invention described herein, but it is not the only device architecture on which at least a portion of the features and technology described herein are realized. For example, an architecture having one or more processors 63 may be used, such processors 63 residing in a single device or distributed across multiple devices. In one embodiment, a single processor 63 handles communication and routing calculations. In various embodiments, various intelligent automatic assistant features and / or functionalities are realized in an intelligent automatic assistant system including a client device (such as a smartphone or personal digital assistant running client software) and a server system (such as a server system described in more detail below).

[0037] Regardless of the network device configuration, the system of the present invention employs one or more memories or memory modules (e.g., memory block 65) configured to store data, program instructions for general network operation and / or other information relating to the functionality of the intelligent automated assistant technology described herein. The program instructions control, for example, the operation of the operating system and / or one or more applications. One or more memories are further configured to store data structures, keyword classification information, advertising information, user click and impression information and / or other specific non-program information described herein.

[0038] Since such information and program instructions are used to implement the systems / methods described herein, at least some embodiments of network devices include non-temporary machine-readable storage media configured or designed to store program instructions and state information, etc., for performing various operations described herein. Examples of such non-temporary machine-readable storage media include, but are not limited to, magnetic media such as hard disks, floppy disks and magnetic tapes, optical media such as CD-ROM disks, magneto-optical media such as floppy disks, and hardware devices such as read-only memory elements (ROM), flash memory, memristor memory and random access memory (RAM) that are configured specifically to store and execute program instructions. Examples of program instructions include machine code, such as that generated by a compiler, and files containing high-level code that is executed by a computer using an interpreter.

[0039] In one embodiment, the system of the present invention is realized in a standalone computing system. Referring now to Figure 4, a block diagram is shown illustrating an architecture for realizing at least a portion of an intelligent automatic assistant in a standalone computing system according to at least one embodiment. The computing device 60 includes a processor 63 that runs software for realizing the intelligent automatic assistant 1002. The input device 1206 is any type of input device suitable for receiving user input, including, for example, a keyboard, a touchscreen, a microphone (e.g., for voice input), a mouse, a touchpad, a trackball, a directional switch, a joystick, and / or any combination thereof. The output device 1207 is a screen, a speaker, a printer, and / or any combination thereof. The memory 1210 is a random access memory having a structure and architecture known in the prior art that the processor 63 uses when running the software. The storage device 1208 is any magnetic storage device, optical storage device, and / or electrical storage device that stores data in digital form, such as flash memory, magnetic hard drives, and / or CD-ROMs.

[0040] In another embodiment, the system of the present invention is implemented in a distributed computing network and has, for example, one or more clients and / or servers. Referring now to Figure 5, a block diagram is shown illustrating an architecture that implements at least a portion of the intelligent automated assistant in a distributed computing network according to at least one embodiment.

[0041] In the configuration shown in Figure 5, one or more clients 1304 are provided. Each client 1304 runs software that implements the client-side portion of the present invention. Furthermore, one or more servers 1340 may be provided to handle requests received from the clients 1304. The clients 1304 and servers 1340 communicate with each other via an electronic network 1361, such as the Internet. The network 1361 is implemented using any known network protocol, including, for example, wired and / or wireless protocols.

[0042] In one embodiment, the server 1340 calls an external service 1360 when it needs to retrieve additional information or refer to stored data relating to a previous interaction with a particular user. Communication with the external service 1360 takes place, for example, via a network 1361. In various embodiments, the external service 1360 includes web-enabled services and / or functionalities associated with or installed on the hardware device itself. For example, in one embodiment where the assistant 1002 is implemented on a smartphone or other electronic device, the assistant 1002 can retrieve information, contacts, and / or other sources stored in a calendar application ("app").

[0043] In various embodiments, the assistant 1002 can control many features and operations of the electronic device on which it is installed. For example, the assistant 1002 calls an external service 1360 that interfaces with the device's functionality and applications via an API or other means to perform functions and operations that would otherwise be initiated using a conventional user interface on the device. Such functions and operations include, for example, setting alarms, making calls, sending text or email messages, and adding calendar events. Such functions and operations are performed as add-on functions in the context of a conversational dialog between the user and the assistant 1002. Such functions and operations are specified by the user in the context of such a dialog or are performed automatically based on the context of the dialog. Thus, the assistant 1002 is used as a control mechanism to initiate and control various operations of an electronic device, and it is used as an alternative mechanism to conventional mechanisms such as buttons or graphical user interfaces.

[0044] For example, the user provides input to the assistant 1002 such as "I need to wake tomorrow at 8 a.m." Once the assistant 1002 determines the user's intent, it uses the techniques described herein to call an external service 1340 to interface with the device's alarm clock function or application. The assistant 1002 sets the alarm on behalf of the user. Thus, the user uses the assistant 1002 as an alternative to conventional mechanisms for setting alarms or performing other functions of the device. If the user's request is ambiguous or requires further explanation, the assistant 1002 may use various techniques described herein, including active derivation, paraphrasing, and suggestions, to obtain the information necessary to ensure that the appropriate service 1340 is called and the intended action is performed. In one embodiment, the assistant 1002 prompts the user for confirmation before calling service 1340 to perform a function. In one embodiment, the user can selectively disable the assistant 1002's ability to call a particular service 1340, or can disable all such service calls as requested.

[0045] The system of the present invention is realized by many different clients 1304 and operating modes. Referring now to Figure 6, a block diagram is shown illustrating a system architecture illustrating several different clients 1304 and operating modes. The various clients 1304 and operating modes shown in Figure 6 are merely illustrative, and it will be understood by those skilled in the art that the system of the present invention can be realized using clients 1304 and / or operating modes other than those shown. Furthermore, the system may include any or all of such clients 1304 and / or operating modes individually or in any combination. Examples shown include:

[0046] A computer device 1402 including input / output devices and / or sensors. Client components are deployed in any such computer device 1402. At least one embodiment is implemented using a web browser 1304A or other software application that enables communication with a server 1340 via a network 1361. The input / output channels may be any type of channel, including, for example, visual and / or auditory channels. For example, in one embodiment, the system of the present invention is implemented using a voice communication method, enabling an embodiment of an assistant for the visually impaired in which a web browser equivalent is voice-driven and uses voice for output.

[0047] • A mobile device 1406 that includes I / O and sensors, which are implemented as an application by the client on mobile device 1304B. This includes, but is not limited to, mobile phones, smartphones, personal digital assistants, tablet devices, and networked game consoles.

[0048] • A consumer device 1410, including I / O and sensors, which is implemented as an embedded application in device 1304C.

[0049] • Automobiles and other vehicles 1414, including dashboard interfaces and sensors, which are implemented by the client as embedded system applications 1304D. This includes, but is not limited to, automobile navigation systems, voice control systems and in-car entertainment systems, etc.

[0050] The client is a networked computing device 1418 such as a router, which is implemented as a device-resident application 1304E, or any other device that resides on or interfaces with the network.

[0051] One embodiment of the assistant is an email client 1424 connected via an email modality server 1426. The email modality server 1426 acts as a communication bridge, for example, by receiving user input as an email message sent to the assistant and sending output from the assistant to the user as a response.

[0052] One embodiment of the assistant is an instant messaging client 1428 connected via a messaging modality server 1430. The messaging modality server 1430 acts as a communication bridge, receiving user input as messages sent to the assistant and sending output from the assistant to the user as messages in response.

[0053] One embodiment of the assistant is a voice telephone 1432 connected via a VoIP (Voice over IP (Internet Protocol)) modality server 1430. The VoIP modality server 1430 acts as a communication bridge, taking in user input as voice spoken to the assistant and sending output from the assistant to the user, for example, as synthesized speech, when responding.

[0054] For messaging platforms including, but not limited to, email, instant messaging, discussion forums, group chat sessions, live help, or customer support sessions, Assistant 1002 acts as a participant in the conversation. Assistant 1002 monitors the conversation and responds to individuals or groups using one or more techniques and methods described herein for one-on-one interactions.

[0055] In various embodiments, the functionality for realizing the technology of the present invention is distributed among one or more client and / or server components. For example, various software modules are implemented to perform various functions related to the present invention, and such modules are implemented in various forms for execution on server and / or client components. Referring now to Figure 7, an example of a client 1304 and a server 1340 communicating with each other to realize the present invention according to one embodiment is shown. Figure 7 shows one possible configuration in which the software modules are distributed among the client 1304 and the server 1340. The illustrated configuration is merely illustrative, and it will be understood by those skilled in the art that such modules can be distributed in many different ways. Furthermore, one or more clients 1304 and / or servers 1340 may be provided, and the modules are distributed among those clients 1304 and / or servers 1340 in any of many different ways.

[0056] In the example in Figure 7, the input derivation functionality and output processing functionality are distributed between client 1304 and server 1340. The client portion 1094a of input derivation and the client portion 1092a of output processing are located on client 1304, while the server portion 1094b of input derivation and the server portion 1092b of output processing are located on server 1340. The following components are located on server 1340.

[0057] • Complete Glossary 1058b • Complete library of language pattern recognition software 1060b • Short-term personal memory master version 1052b • Long-term personal memory master version 1054b

[0058] In one embodiment, client 1304 maintains a subset and / or part of its components locally to improve responsiveness and reduce dependence on network communication. Such subsets and / or parts are maintained and updated according to well-known cache management techniques. Such subsets and / or parts include, for example, the following:

[0059] • Glossary subset 1058a • A subset of the language pattern recognition library, 1060a • Short-term personal memory cache 1052a • Long-term personal memory cache 1054a

[0060] The additional components are implemented as part of server 1340 and include, for example, the following: • Language interpreter 1070 • Dialogflow Processor 1080 • Output processor 1090 • Domain entity database 1072 • Task flow model 1086 • Service Orchestration 1082 • Service function model 1088

[0061] Each of these components is described in more detail below. Server 1340 obtains additional information by interfaceing with external service 1360 as needed.

[0062] Conceptual architecture Referring next to Figure 1, a schematic block diagram of a particular embodiment of the intelligent automatic assistant 1002 is shown. As will be described in more detail herein, various embodiments of the intelligent automatic assistant system are generally configured, designed and / or operable to provide various operations, functionalities and / or features related to intelligent automatic assistant technology. Also, as will be described in more detail herein, many of the various operations, functionalities and / or features of the intelligent automatic assistant system disclosed herein enable or provide various advantages and / or benefits to various entities that interface with the intelligent automatic assistant system. The embodiment shown in Figure 1 is implemented using one of the hardware architectures described above or using a different type of hardware architecture.

[0063] For example, according to various embodiments, at least some intelligent automatic assistant systems are configured, designed and / or operable to provide a variety of operations, functions and / or features, such as one or more of the following operations, functions and / or features (or combinations thereof).

[0064] • Automate the application of data and services available via the internet to discover, find, and select products and services, particularly for purchase, reservation, or order. In addition to automating the process of using this data and services, the Intelligent Auto Assistant 1002 enables the combined use of several data sources and services at once. For example, the Intelligent Auto Assistant combines product information from several review sites, checks availability and pricing from multiple sellers, checks location and time constraints, and makes it easier to find personalized solutions to the user's problems.

[0065] Automate the use of data and services available via the Internet to discover, research, and select, in particular to book or learn about, any other sources of social interaction or entertainment that may be found on the Internet, including, but not limited to, movies, events, performances, exhibitions, shows, and tourist attractions; places to go (including, but not limited to, travel destinations, hotels and other places to stay, landmarks and other places of interest); places to eat and drink (restaurants and bars, etc.); times and places to meet other people; and any other sources of social interaction or entertainment that may be found on the Internet.

[0066] The application and service can be operated via natural language dialogs provided by a dedicated application having a graphical user interface that includes: searching (including location-based searching), navigation (maps and directions), database lookup (such as finding businesses or people by name or other characteristics), retrieving weather conditions and forecasts, checking market item prices or financial transaction status, monitoring traffic or flight status, accessing and updating calendars and schedules, managing reminders, alerts, tasks and projects, communication via email or other messaging platforms, and local or remote operation of the device (e.g., making phone calls, controlling lights and temperature, controlling home security devices, playing music or videos). In one embodiment, the assistant 1002 is used to start, operate and control many of the functions and applications available on the device.

[0067] • Provides personalized recommendations for activities, products, services, entertainment sources, time management, or any other type of recommendation service, leveraging natural language dialogue and automated access to data and services.

[0068] According to various embodiments, at least a portion of the various functions, operations, behaviors and / or other features provided by the intelligent automatic assistant 1002 are realized in one or more client systems, one or more server systems and / or a combination thereof.

[0069] According to various embodiments, at least a portion of the various functions, operations, actions and / or other features provided by the assistant 1002 are realized by at least one embodiment of an automated call and response procedure, such as the one illustrated and described with respect to Figure 33.

[0070] Furthermore, the various embodiments of the Assistant 1002 described herein include or provide a number of different advantages and / or benefits compared to existing intelligent automatic assistant technologies, such as one or more (or a combination thereof) of the following:

[0071] The integration of speech-text and natural language understanding technologies is constrained by a set of explicit models of domains, tasks, services, and dialogues. Unlike assistant technologies that aim to create general-purpose artificial intelligence systems, the embodiments described herein apply multiple sources of constraints to reduce the number of solutions to a more manageable size. As a result, ambiguous interpretations of language are reduced, the number of relevant domains or tasks is reduced, and the ways in which intent in a service can be made operational are reduced. By focusing on specific domains, tasks, and dialogues, scope across domains and tasks can be achieved through human-managed glossaries and mappings from intent to service parameters.

[0072] • The ability to resolve user problems by calling services over the internet using APIs for the user's problems. Unlike search engines that only return links and content, some embodiments of the automated assistant 1002 described herein automate exploration and problem-solving activities. The ability to call multiple services for a given request provides users with broader functionality than that achieved by visiting a single site, for example, to generate a product or service or to find out what to do.

[0073] • Application of personal information and personal interaction history in interpreting and performing user requests. Unlike conventional search engines or question-answering services, the embodiments described herein use information such as personal interaction history (e.g., history of dialogues and subsequent selections from results), personal physical context (e.g., user's location and time), and personal information collected in the context of the interaction (e.g., name, email address, physical address, telephone number, account number, and preferences). By using these sources of information, it becomes possible, for example, to:

[0074] • More appropriate interpretation of user input (e.g., using personal history and physical context when interpreting language) • More personal results (e.g., biased towards preferences or recent choices) • Improved efficiency for users (for example, by automating steps including service sign-up or form completion) • Use of dialog history when interpreting natural language user input. Embodiments maintain a personal history and interpret new input using the dialog context, such as the current location, time, domain, task step, and task parameters, to apply natural language understanding to user input. Conventional search engines and command processors interpret at least one query regardless of the dialog history. The ability to use dialog history allows for a more natural interpretation, which is similar to ordinary human conversation.

[0075] • Active input derivation, in which Assistant 1002 actively guides and constrains user input based on the same model and information used to interpret user input. For example, Assistant 1002 may apply a dialogue model to suggest the next step in a dialogue with a user who is refining a request, or it may complete partially typed input based on domain and context-specific possibilities, or it may use semantic interpretation to choose between an ambiguous interpretation of speech as text or an ambiguous interpretation of text as intent.

[0076] Explicit modeling and dynamic management of services with dynamic and robust service orchestration. The architecture of the embodiment described allows Assistant 1002 to interface with many external services, dynamically determine which services provide information for a particular user request, map user request parameters to various service APIs, call multiple services at once, integrate results from multiple services, appropriately failover to failed services, and / or efficiently maintain service implementation as APIs and functions evolve.

[0077] Active ontology is used as a method and apparatus for constructing Assistant 1002. This simplifies the software engineering and data maintenance of the automated assistant system. Active ontology is the integration of data modeling and execution environments for the assistant. They provide a framework for linking various model and data sources (domain concepts, task flows, glossaries, language pattern recognizers, dialogue contexts, user personal information, and mappings from domain and task requests to external services). The active ontology and other architectural inventions described herein enable the construction of deep functionality within a domain to unify multiple information sources and services, and to do so across a set of domains.

[0078] In at least one embodiment, the intelligent automatic assistant 1002 is operable to utilize and / or generate various data and / or other types of information when performing a particular task and / or operation. This includes, for example, input data / information and / or output data / information. For example, in at least one embodiment, the intelligent automatic assistant 1002 is operable to access, process, and / or utilize information from one or more different sources, such as one or more local and / or remote memories, devices and / or systems. Furthermore, in at least one embodiment, the intelligent automatic assistant 1002 is operable to generate one or more different types of output data / information, for example, stored in the memory of one or more local and / or remote devices and / or systems.

[0079] Examples of the various input data / information that the intelligent automatic assistant 1002 accesses and / or uses include, but are not limited to, one or more (or a combination thereof) of the following:

[0080] • Voice input from mobile devices such as mobile phones and tablets, computers with microphones, Bluetooth headsets, automotive voice control systems, telephone systems, recording in answering services, audio voicemail in integrated messaging services, clock radio and other consumer applications with voice input, telephone exchanges, home entertainment control systems and game consoles.

[0081] - Keyboards of computers or mobile devices, keypads of remote controls or other consumer electronic devices, email messages sent to the assistant, instant messages or similar short messages sent to the assistant, messages received from players in a multi-user gaming environment, and text input from streamed text in message delivery.

[0082] • Location information input from sensors or location-based systems. Examples include Global Positioning System (GPS) and Assisted GPS (A-GPS) in mobile phones. In one embodiment, location information is combined with explicit user input. In one embodiment, the system of the present invention can detect when the user is at home based on known address information and current location determination. Thus, certain inferences are made regarding the types of information the user might be interested in when at home compared to when outside, and the types of services and actions that should be called upon for the user depending on whether the user is at home or not.

[0083] • Time information from the client device's clock. This includes, for example, the time on the phone or other client device indicating local time and time zone. Furthermore, the time is used in the context of the user request to interpret phrases such as "in an hour" and "tonight."

[0084] • Compass, accelerometer, gyroscope, and / or motion speed data, as well as other sensor data from embedded systems such as mobile devices, handheld devices, or automotive control systems. This includes device positioning data from remote controls to equipment and game consoles.

[0085] Click events, menu selection events, and other events from a graphical user interface (GUI) on any device having a GUI. Further examples include touches on touchscreens.

[0086] • Alarm clocks, calendar alerts, price change triggers, location triggers, and sensor-based event and other data-driven triggers such as press notifications from servers to devices.

[0087] The input to the embodiments described herein further includes the context of a user interaction history, including dialogs and request history.

[0088] Examples of the various output data / information generated by the intelligent automatic assistant 1002 include, but are not limited to, one or more (or a combination thereof) of the following:

[0089] • Text output sent directly to the output device and / or the device's user interface. • Text and graphics sent to the user via email • Text and graphics sent to the user via messaging services • Audio output. This includes one or more (or a combination thereof) of the following:

[0090] • Synthesized voice • Sampled audio • Recorded message • Graphic layout of information including photos, rich text, videos, audio, and hyperlinks. For example, content rendered in a web browser.

[0091] Actuator outputs for controlling physical actions on the device, such as turning the power on or off, making sounds, changing colors, vibrating, or controlling lighting.

[0092] • Calling mapping applications, making phone voice dials, sending emails or instant messages, playing media, entering data into calendar, task manager and memo applications, and calling other applications on the device, such as other applications.

[0093] Actuator outputs for controlling physical actions on devices connected to or controlled by other devices, such as remote camera operation, wheelchair control, music playback on remote speakers, and video playback on remote displays.

[0094] It is understood that the intelligent automatic assistant 1002 in Figure 1 is an example of a broader range of intelligent automatic assistant systems that can be implemented. Other embodiments of the intelligent automatic assistant system (not shown) may include additional, fewer, and / or different components / features compared to, for example, the components / features illustrated in the example of the intelligent automatic assistant system embodiment in Figure 1.

[0095] User interaction Referring to Figure 2, an example of interaction between at least one embodiment of the intelligent automatic assistant 1002 and a user is shown. The example in Figure 2 assumes that the user is speaking to the intelligent automatic assistant 1002 using an input device 1206, which is a voice input mechanism, and that the output is a graphic layout to an output device 1207, which is a scrollable screen. Conversation screen 101A features a conversational user interface that shows what the user said 101B ("I'd like a romantic place for Italian food near my office"), a summary of the result 101C ("OK, I found these Italian restaurants which reviews say are romantic close to your work"), and a set of results 101D (the first three restaurants in the list are displayed). In this example, the user clicks the first result in the list, and the result automatically opens to show further information about the restaurant, which is displayed on information screen 101E. The information screen 101E and the conversation screen 101A are displayed on the same output device, such as a touchscreen or another display device. In other words, the example shown in Figure 2 represents two different output states for the same output device.

[0096] In one embodiment, the information screen 101E displays information collected and combined from various services, including, for example, any or all of the following:

[0097] • Store address and geolocation (location information?) • Distance from the user's current location • Reviews from multiple sources

[0098] In one embodiment, the information screen 101E further includes some examples of services that the assistant 1002 provides for the user, including the following: • Dial the number to make a phone call to the store ("outgoing call") • Remember this restaurant for future reference ("save"). • Share directions and information about this restaurant with someone via email ("share"). • Display the location and directions to this restaurant on a map ("map it" (Show on map)) • Save personal notes about this restaurant ("my notes")

[0099] As shown in the example in Figure 2, in one embodiment, the assistant 1002 includes intelligence beyond simple database applications, such as the following. • It processes not just keywords, but also expressions of intent in natural language. • Infer semantic intent from the language input. For example, interpret "place for Italian food" as "Italian restaurants." • Turn semantic intent into a strategy for using online services and make that strategy work for the user (for example, turning the desire for a romantic place into a strategy for checking online review sites for reviews that describe the place as "romantic"). Intelligent Auto Assistant Components

[0100] According to various embodiments, the intelligent automatic assistant 1002 includes a plurality of various components, devices, modules, processes, and systems, etc., which are implemented and / or instantiated using, for example, hardware and / or a combination of hardware and software. For example, as shown in the embodiment of Figure 1, the assistant 1002 includes one or more of the following types of systems, components, devices, and processes, etc. (or combinations thereof):

[0101] • One or more active ontology 1050s • Active input derivation component 1094 (including client portion 1094a and server portion 1094b) • Short-term personal memory component 1052 (including master version 1052b and cache 1052a) • Long-term personal memory component 1054 (including master version 1054b and cache 1054a) • Domain model component 1056 • Glossary component 1058 (including the complete glossary 1058b and subset 1058a) • Language pattern recognition component 1060 (including full library 1060b and subset 1560a) • Language interpreter component 1070 • Domain entity database 1072 • Dialogflow processor component 1080 • Service orchestration component 1082 • Service component 1084 • Task flow model component 1086 • Dialog flow model component 1087 • Service model component 1088 Output processor component 1090

[0102] As explained in relation to Figure 7, in an embodiment using a specific client / server, some or all of these components are distributed between client 1304 and server 1340.

[0103] Next, for illustrative purposes, at least some of the various components of a particular embodiment of the intelligent automatic assistant 1002 will be described in more detail with reference to the example of the embodiment of the intelligent automatic assistant 1002 shown in Figure 1.

[0104] Active Ontology 1050 The active ontology 1050 acts as a unified infrastructure for integrating data from models, components, and / or other parts of the embodiment of the intelligent automated assistant 1002. In the field of computer and information science, ontologs provide structures for data and knowledge representations such as classes / kinds, relationships, attributes / characteristics, and their instantiation in instances. For example, ontologs are used to build models of data and knowledge. In some embodiments of the intelligent automated assistant 1002, the ontology is part of a modeling framework for building models such as domain models.

[0105] In the context of the present invention, the "active ontology" 1050 operates as an execution environment in which separate processing elements are configured like an ontology (for example, having separate attributes and relationships with other processing elements). These processing elements perform at least a portion of the tasks of the intelligent automated assistant 1002. Any number of active ontologeries 1050 can be provided.

[0106] In at least one embodiment, the active ontology 1050 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or combinations thereof) of the following:

[0107] It operates as a modeling and development environment, integrating models and data from various model and data components. This includes, but is not limited to, the following: • Domain Model 1056 ·Glossary 1058 • Domain entity database 1072 • Task flow model 1086 • Dialogue flow model 1087 • Service function model 1088

[0108] • The ontology-based editing tool operates as a data modeling environment that enables the development of new models, data structures, database schemas, and representations.

[0109] It operates as a live execution environment, instantiating elements of domain 1056, task 1086 and / or dialogue model 1087, language pattern recognizers, and / or values ​​for glossary 1058, as well as user-specific information that can be found in short-term personal memory 1052, long-term personal memory 1054 and / or service orchestration 1182. For example, some nodes of the active ontology correspond to domain concepts such as restaurants and their characteristics, such as restaurant names. During live execution, these active ontology nodes are instantiated by matches of a particular restaurant entity and its name, and by how that name corresponds to a word in natural language input utterances. Thus, in this embodiment, the active ontology specifies the concept that a restaurant is an entity that includes a name match, and operates as a modeling environment that stores natural language parsing and data from the entity database and their dynamic connections to the modeling nodes.

[0110] For example, it enables communication and cooperation between intelligent automated assistant components and processing elements, such as one or more (or a combination thereof) of the following: • Active input derivation component 1094 • Language interpreter component 1070 • Dialogflow processor component 1080 • Service orchestration component 1082 • Service component 1084

[0111] In one embodiment, at least a portion of the functions, operations, behaviors and / or other features of the active ontology 1050 described herein are at least partially realized using various methods and apparatus described in U.S. Patent Application No. 11 / 518,292, filed on September 8, 2006, “Method and Apparratus for Building an Intelligent Automated Assistant”.

[0112] In at least one embodiment, a given instance of the active ontology 1050 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the active ontology 1050 include, but are not limited to, one or more (or a combination thereof) of the following:

[0113] Static data available from one or more components of the intelligent automatic assistant 1002.

[0114] Data that is dynamically instantiated for each user session and maintains, for example, the user-specific input and output states exchanged between the components of the intelligent automated assistant 1002, the contents of short-term personal memory, and inferences made from the state of the user session, etc., without limitation.

[0115] Thus, the active ontology 1050 is used to unify the elements of various components in the intelligent automated assistant 1002. The active ontology 1050 allows document creators, designers, or system builders to integrate components so that the elements of one component are identified by the elements of other components. Therefore, document creators, designers, or system builders can more easily combine and integrate components.

[0116] Next, referring to Figure 8, an example of a part of the active ontology 1050 according to one embodiment is shown. This example is intended to facilitate illustration of some of the various functions, operations, behaviors and / or other features provided by the active ontology 1050.

[0117] The active ontology 1050 in Figure 8 includes representations of a restaurant and a dining event. In this example, the restaurant is a concept 1610 that includes characteristics such as name 1612, dishes served 1615, and location 1613, which is modeled as a structured node with characteristics for address 1614. The concept of a dining event is modeled as a node 1616 that includes a dining party 1617 (having size 1619) and duration 1618.

[0118] The active ontology includes and / or references domain model 1056. For example, Figure 8 shows a foodservice domain model 1622 linked to the restaurant concept 1610 and the meal event concept 1616. In this instance, the active ontology 1050 includes the foodservice domain model 1622, and in particular, at least two nodes of the active ontology 1050, namely restaurant 1610 and meal event 1616, are included in and / or referenced in the foodservice domain model 1622. This domain model specifically represents the idea that foodservice includes meal events that occur in restaurants. The restaurant 1610 and meal event 1616 nodes of the active ontology are further included in and / or referenced in other components of the intelligent automated assistant, as indicated by the dotted lines in Figure 8.

[0119] The active ontology includes and / or references the task flow model 1086. For example, Figure 8 shows an event planning task flow model 1630 that models the planning of domain-independent events, i.e., meal events 1616, which are domain-specific event types. Here, the active ontology 1050 includes the general event planning task flow model 1630, which includes nodes representing events and other concepts involved in planning those events. The active ontology 1050 further includes nodes for meal events 1616, which are a specific type of event. In this example, meal events 1616 are included in or referenced in the domain model 1622 and the task flow model 1630, both of which are included in and / or referenced in the active ontology 1050. Here again, meal events 1616 unify elements of various components that are included in and / or referenced in other components of the intelligent automated assistant, as shown by the dotted lines in Figure 8.

[0120] The active ontology includes and / or references the dialog flow model 1087. For example, Figure 8 shows a dialog flow model 1642 for obtaining the constraint values ​​required for a transaction instantiated with a constraint party size as represented by concept 1619. Here again, the active ontology 1050 provides a framework for relating and unifying various components such as the dialog flow model 1087. In this case, the dialog flow model 1642 has a general concept of the constraints instantiated in this particular example for the node of party size 1619 in the active ontology. This particular dialog flow model 1642 operates with an abstract concept of constraints, independent of the domain. The active ontology 1050 represents the party size characteristic 1619 of the party node 1617 related to the meal event node 1616. In one such embodiment, the intelligent automated assistant 1002 uses the active ontology 1050 to unify the characteristics of a party size 1619, which is part of a cluster of nodes representing a meal event concept 1616, which is part of a domain model 1622 for dining out, with the concept of constraints in a dialogue flow model 1642.

[0121] The active ontology includes and / or references service model 1088. For example, Figure 8 shows a model of a restaurant reservation service associated with a dialog flow step for obtaining values ​​required for the restaurant reservation service 1672 to execute a transaction. In this instance, service model 1672 for the restaurant reservation service specifies that the reservation requires a value for party size 1619 (the number of people to sit at the table to be reserved). The concept of party size 1619, which is part of the active ontology 1050, is linked to or associated with a general dialog flow model 1642 that asks the user about constraints for the transaction. In this instance, party size is the required constraint for dialog flow model 1642.

[0122] The active ontology includes and / or references the domain entity database 1072. For example, Figure 8 shows the domain entity database for restaurant 1652 associated with restaurant node 1610 in the active ontology 1050. The active ontology 1050 represents a general concept of restaurant 1610, as used by various components of the intelligent automated assistant 1002, which is instantiated by data about the specific restaurant in the restaurant database 1652.

[0123] The active ontology includes and / or references the glossary database 1058. For example, Figure 8 shows the glossary database of cuisines 1662 such as Italian and French, and words associated with each cuisine such as "French," "continental," and "provincial." The active ontology 1050 includes a restaurant node 1610 related to a cuisine node 1615, and the cuisine node 1615 is associated with a representation in the cuisine database 1662. A specific entry in the database 1662 for a cuisine such as "French" is associated via the active ontology 1050 as an instance of the concept of cuisine 1615.

[0124] The Active Ontology includes and / or references any databases that map to concepts or other representations in Ontology 1050. The Domain Entity Database 1072 and the Glossary Database 1058 are just two examples of how the Active Ontology 1050 integrates databases with each other and with other components of the Automated Assistant 1002. The Active Ontology allows document creators, designers, or system builders to specify non-trivial mappings between representations in databases and representations in Ontology 1050. For example, the database schema for the Restaurant Database 1652 represents a restaurant as a table of strings and numbers, or as a picture from a larger database of businesses, or as any other representation suitable for database 1652. In this example of Active Ontology 1050, Restaurant 1610 is a conceptual node with characteristics and relationships organized differently from a database table. In this example, nodes in Ontology 1050 are associated with elements in the database schema. The integration of the database and ontology 1050 provides a unified representation for interpreting and acting upon specific data entries in the database in relation to a larger set of models and data in the active ontology 1050. For example, the word "French" may be an entry in the cuisine database 1662. In this example, because database 1662 is integrated into the active ontology 1050, the same word "French" also has the interpretation of a possible dish served in a restaurant, which is relevant when planning a dining event, and this dish serves as a constraint used when using restaurant reservation services, etc. The active ontology models the database and integrates it into the execution environment so that it can interact with other components of the automated assistant 1002.

[0125] As described above, the active ontology 1050 allows document creators, designers, or system builders to integrate components. Therefore, in the example in Figure 8, the elements of components such as constraints in the dialogue flow model 1642 are identified by the elements of other components such as the required parameters of the restaurant reservation service 1672.

[0126] The active ontology 1050 is embodied, for example, as a model, database, and component configuration such that the relationships between the model, database, and components are one of the following: • Container-related and / or contained • Relationship with links and / or pointers • Interfaces via APIs within and between programs

[0127] Referring, for example, to Figure 9, an example of another embodiment of the intelligent automatic assistant system 1002 is shown. Here, the components of the domain model 1056, glossary 1058, language pattern recognizer 1060, short-term personal memory 1052, and long-term personal memory 1054 are organized under a common container associated with the active ontology 1050, while other components such as the active input derivation component 1094, language interpreter 1070, and dialogue flow processor 1080 are associated with the active ontology 1050 via API relationships.

[0128] Active input derivation component 1094 In at least one embodiment, the active input derivation component 1094 (which is implemented in a standalone configuration or a configuration including server and client components, as described above) is operable to perform and / or realize various functions, operations, actions and / or other features such as one or more (or combinations thereof) of the following:

[0129] • Derives, facilitates, and / or processes information about user or user environment input, and / or requests or demands. For example, if a user is searching for a restaurant, the input deriving module obtains information about the user's constraints or preferences regarding location, time, cuisine, and price, etc.

[0130] • For example, it facilitates various inputs from various sources such as one or more (or a combination thereof) of the following: • Input from a keyboard or any other input device that generates text. • Keyboard input in a user interface that provides proposed dynamic completion for partial input • Input from a voice input system • Input from a graphical user interface (GUI) where the user clicks, selects, or directly interacts with graphic objects to indicate options. • Input from other applications that generate text, including email, text messaging, or other text communication platforms, and send it to the automated assistant.

[0131] By performing active input derivation, the assistant 1002 can eliminate ambiguity of intent in the early stages of input processing. For example, in one embodiment where the input is provided by voice, the waveform is sent to the server 1340, where terms are extracted and semantically interpreted. The results of such semantic interpretation are used to drive active input derivation, providing the user with alternative candidate words that can be selected based on semantic relevance and phonetic matching.

[0132] In at least one embodiment, the active input derivation component 1094 actively, automatically, and dynamically guides the user to an input acted upon by one or more services provided by an embodiment of the assistant 1002. Referring now to Figure 10, a flowchart illustrating how the active input derivation component 1094 operates according to one embodiment is shown.

[0133] The procedure begins (20). In step 21, the assistant 1002 provides an interface for one or more input channels. For example, the user interface provides a speaking user option, a typing user option, or a tapping user option at any stage of the conversational interaction. In step 22, the user selects an input channel by initiating input in one modality, such as by pressing a button to start recording voice or by displaying an interface for typing.

[0134] In at least one embodiment, the assistant 1002 provides default suggestions for the selected modality (23). That is, the assistant 1002 provides relevant options 24 in the current context before the user enters any input in that modality. For example, in a text input modality, the assistant 1002 provides a list of common words to initiate a text request or command, such as one or more imperative verbs (e.g., find, buy, reserve, get, call, check, and schedule), nouns (e.g., restaurants, movies, events, and businesses), or menu-like options that name domains of conversation (e.g., weather, sports, and news).

[0135] If the user selects one of the default options in step 25 and sets a preference for automatic submission (30), the procedure immediately returns. This is similar to the behavior of traditional menu selection.

[0136] However, the initial options are taken in as partial inputs, or the user has begun to input partial inputs (26). In at least one embodiment, the user chooses at some point in the input to indicate that the partial input is complete (22), thereby causing the procedure to return.

[0137] In step 28, the most recent input is added to the cumulative input, regardless of whether it was selected or entered.

[0138] In 29, the system proposes the next possible related input, assuming other sources of constraints on what constitutes a related and / or meaningful input, and the current input.

[0139] In at least one embodiment, the source of constraints on user input (for example, used in steps 23 and 29) is one or more of the various models and data sources included in Assistant 1002, which include, but are not limited to, one or more (or combinations thereof) of the following:

[0140] • Glossary 1058. For example, a word or phrase matching the current input is suggested. In at least one embodiment, the glossary is associated with one or more nodes from among the active ontology, domain model, task model, dialogue model, and / or service model.

[0141] • A domain model 1056 that instantiates a domain model or constrains inputs consistent with the domain model. For example, in at least one embodiment, the domain model 1056 is used to propose concepts, relationships, properties and / or instances consistent with the current input.

[0142] A language pattern recognizer 1060 is used to recognize phrases, grammatical structures, or other patterns in the current input, and to suggest completions to make the pattern complete.

[0143] • Domain entity database 1072 used to suggest possible domain entities that match the input (e.g., business name, movie title, and event name).

[0144] Short-term memory 1052 used to match previous input or a portion of previous input, and / or any other characteristics or facts relating to the history of user interaction. For example, a partial input is matched with cities the user has traveled to during the session, whether virtual (e.g., indicated by a query) and / or physical (e.g., determined from a location sensor).

[0145] In at least one embodiment, a semantic paraphrase of a recent input, request, or result is matched with the current input. For example, if the user previously requested "live music" and obtained a concert list, and then typed "music" in the active input derivation environment, the suggestions would include "live music" and / or "concerts".

[0146] Long-term personal memory 1054 is used to suggest matching items from long-term memory. Such matching items include, for example, one or more or a combination thereof of stored domain entities (e.g., “favorite” restaurants, movies, theaters and places), to-do items, list items, calendar entries, names of people in contacts / phone books, and street or city names shown in contacts / phone books.

[0147] Task flow model 1086 is used to suggest inputs based on the following possible steps in the task flow.

[0148] Dialogflow Model 1087 is used to suggest inputs based on the following possible steps in the dialog flow.

[0149] • A service function model 1088 is used to suggest possible services to adopt based on their name, category, function, or any other characteristics in the model. For example, the user types part of a preferred review site name, and the assistant 1002 suggests a complete command to query review sites for reviews.

[0150] In at least one embodiment, the active input derivation component 1094 presents the user with a conversational interface, such as an interface in which the user and assistant communicate by uttering words to each other in a conversational manner. The active input derivation component 1094 is operable to execute and / or realize various conversational interfaces.

[0151] In at least one embodiment, the active input derivation component 1094 is operable to execute and / or implement various conversational interfaces that use multiple conversational exchanges to instruct the assistant 1002 to input information from the user according to a dialog model. The dialog model represents a procedure for executing a dialog, such as a series of steps required to derive information necessary to perform a service.

[0152] In at least one embodiment, the active input derivation component 1094 provides the user with constraints and guidance in real time while the user is typing, speaking, or creating input. For example, active derivation guides the user to type text input that is recognizable by one embodiment of the assistant 1002 and / or serviced by one or more services provided by an embodiment of the assistant 1002. This is advantageous over passively waiting for unconstrained input from the user because it allows the user's effort to focus on potentially useful input and / or allows the embodiment of the assistant 1002 to apply interpretation of the input in real time while the user is typing.

[0153] At least a portion of the functions, operations, behaviors and / or other features of the active input derivation described herein are at least partially realized using the various methods and apparatus described in U.S. Patent Application No. 11 / 518,292, filed September 8, 2006, “Method and Apparatus for Building an Intelligent Automated Assistant.”

[0154] According to a particular embodiment, multiple instances or threads of the active input derivation component 1094 are realized and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software.

[0155] According to various embodiments, one or more different threads or instances of the active input derivation component 1094 are started in response to the detection of one or more conditions or events that satisfy one or more different minimum threshold criteria that trigger the start of at least one instance of the active input derivation component 1094. Various examples of conditions or events that trigger the start and / or realization of one or more different threads or instances of the active input derivation component 1094 include, but are not limited to, one or more (or combinations thereof) of the following:

[0156] • Initiating a user session. For example, if a user session starts an application which is an embodiment of Assistant 1002, the interface provides the user with an opportunity to begin input, for example, by pressing a button to start a voice input system or by clicking a text field to start a text input session. Detected user input. - When Assistant 1002 explicitly instructs the user to input information, such as when Assistant 1002 requests a response to a question or provides a menu of next steps to be selected. Assistant 1002 facilitates the user performing a transaction, such as filling out a form and collecting data for that transaction.

[0157] In at least one embodiment, a given instance of the active input derivation component 1094 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the active input derivation component 1094 include, but are not limited to, one or more (or combinations thereof) of the following:

[0158] • A database of possible words that can be used in text input. • Grammar of possible phrases used in text-based conversations • Database of possible interpretations of voice input • Database of previous inputs from the user or other users • Data from any of the various models and data sources that are part of the embodiment of Assistant 1002. This includes, but is not limited to, one or more (or a combination thereof) of the following:

[0159] • Domain Model 1056 ·Glossary 1058 • Language pattern recognition unit 1060 • Domain entity database 1072 • Short-term memory 1052 • Long-term personal memory 1054 • Task flow model 1086 • Dialogue flow model 1087 • Service function model 1088

[0160] According to various embodiments, the active input derivation component 1094 applies an active derivation procedure to one or more (or a combination thereof) of the following, for example: • Typed input • Voice input • Input from a graphical user interface (GUI), including gestures • Input from suggestions provided in the dialog • Events from the calculation and / or detection environment

[0161] Active type input derivation Referring to Figure 11, a flowchart illustrating a method for deriving an active type input according to one embodiment is shown.

[0162] The process is initiated (110). The assistant 1002 receives partial text input, for example, via the input device 1206 (111). The partial text input includes, for example, characters previously typed into the text input field. The user may indicate at any point that typing is complete, for example, by pressing the Enter key (112). If not, the suggestion generator generates candidate suggestions 116 (114). These suggestions are syntactic suggestions, semantic suggestions, and / or suggestions of other kinds, based on any of the information sources or constraints described herein. If a suggestion is selected (118), the input is transformed to include the selected suggestion (117).

[0163] In at least one embodiment, the proposal includes an extension to the current input. For example, the proposal for "rest" is "restaurants".

[0164] In at least one embodiment, the suggestion involves replacing a portion of the current input. For example, the suggestion for "rest" is "places to eat".

[0165] In at least one embodiment, the suggestion includes replacing and paraphrasing parts of the current input. For example, if the current input is "find restaurants of style", the suggestion is "italian", and if that suggestion is selected, the entire input is rewritten to "find Italian restaurants".

[0166] In at least one embodiment, the resulting returned input is annotated such that information about the selection made in 118 is stored along with the text input (119). This allows, for example, the underlying semantic concepts or entities of the string to be associated with the returned string, thereby improving the accuracy of subsequent language interpretation.

[0167] Referring to Figures 12 to 21, screenshots are shown illustrating some parts of several steps of an active type input derivation according to one embodiment. The screenshots illustrate an example of one embodiment of Assistant 1002, as implemented in a smartphone such as the iPhone, which is commercially available from Apple, Inc. in Cupertino, California. Input is provided to such a device via a touchscreen, including on-screen keyboard functionality. The screenshots illustrate only one embodiment, and those skilled in the art will understand that the technology of the present invention can be implemented in other devices using other layouts and configurations.

[0168] In Figure 12, screen 1201 contains a set of top-level suggestions 1202 that are shown when no input is provided for field 1203. This corresponds to step 23 in Figure 10 when there is no input, which is applied to step 114 in Figure 11 when there is no input.

[0169] In Figure 13, screen 1301 shows an example of using the glossary to provide suggested completions 1303 for partial user input 1305 entered into field 1203 using the on-screen keyboard 1304. These suggested completions 1303 are part of the functionality of active input derivation 1094. The user enters partial user input 1305 containing the string "comm". The glossary component 1058 provides a mapping of this string to three different types of instances listed as suggested completions 1303. The phrase "community&local events" is a category of the events domain, "chambers of commerce" is a category of the local business exploration domain, and "Jewish Community Center" is an instance name of a local business. The glossary component 1058 provides data references and management of namespaces such as these. The user can tap the Go button 1306 to indicate that they have finished entering input. This causes the assistant 1002 to proceed to the text column completed as a unit of user input.

[0170] In Figure 14, screen 1401 shows an example containing the entire phrase with typed parameters for the proposed semantic completion 1303 for the substring "wh" 1305. These types of completions are made possible by using one or more of the various models and sources of input constraints described herein. For example, in one embodiment shown in Figure 14, "what is happening in city" is the active derivation of the location parameter of the local event domain. "where is business name" is the active derivation of the business name of the local business search domain. "what is showing at the venue name" is the active derivation of the constraint for the location name of the local event domain. "what is playing at the movie theater" is the active derivation of the constraint for the movie theater name of the local event domain. These examples demonstrate that the proposed completions are generated by the model rather than simply obtained from a database of previously entered queries.

[0171] In Figure 15, screen 1501 shows a continuation of the same example after the user has entered additional text 1305 into field 1203. The suggested completion 1303 is updated to match the additional text 1305. In this example, data from the domain entity database 1072 was used; that is, locations whose names begin with "f" were used. Note that this is not all words that begin with "f", but rather a very small set of semantically relevant suggestions. Here again, the suggestions are generated by applying a model that is a domain model representing local events occurring at locations in this example, and these are businesses that have names. The suggestions actively derive input that creates potentially meaningful entries when using the local events service.

[0172] In Figure 16, screen 1601 shows a continuation of the same example after the user has selected one of the suggested completions 1303. Active derivation continues by prompting the user to input further to specify the type of information desired, here by presenting several modifiers 1602 that the user can select. In this example, these modifiers are generated by the domain, task flow, and dialog flow models. The domain is a local event, which includes a category of events that take place in that place on that day and have an event name and main actors. In this embodiment, the fact that these five options are offered to the user is generated from the dialog flow model indicating that the user should request constraints that have not yet been entered and from the service model indicating that these five constraints are parameters to the assistant's available local event services. A selection of phrases suitable for use as modifiers, such as "by category" and "featured," is also generated from the domain word list database.

[0173] In Figure 17, screen 1701 shows a continuation of the same example after the user has selected one of the modifiers 1602.

[0174] In Figure 18, screen 1801 shows a continuation of the same example where the selected modifier 1602 is added to field 1203 and additional modifiers 1602 are presented. The user can select one of the modifiers 1602 and / or provide additional text input via keyboard 1304.

[0175] In Figure 19, screen 1901 shows a continuation of the same example where the selected modifier 1602 is added to field 1203 and further modifiers 1602 are presented. In this example, previously entered constraints are not actively derived again.

[0176] In Figure 20, screen 2001 shows a continuation of the same example when the user taps the Go button 1306. The user's input is shown in box 2002, a message is shown in box 203, and feedback is provided to the user regarding the query being executed in response to the user's input.

[0177] In Figure 21, screen 2101 shows a continuation of the same example when results are found. A message is shown in box 2102. Result 2103 includes input elements that allow the user to view further details, save the identified event, purchase a ticket, or add a note.

[0178] Within one screen 2101, other displayed screens are scrollable, allowing the user to scroll upwards to view screen 2001 or other previously presented screens and, if desired, make changes to the query.

[0179] Active voice input derivation Referring to Figure 22, a flowchart is shown illustrating a method for actively deriving input from voice input according to one embodiment.

[0180] The process is initiated (221). Assistant 1002 receives the audio input in the form of an audible signal (121). The voice-to-text service 122 or processor generates a set of candidate text interpretations 124 of the audible signal. In one embodiment, the voice-to-text service 122 is implemented using, for example, the Nuance Recognizer, commercially available from Nuance Communications, Inc. in Burlington, MA.

[0181] In one embodiment, the assistant 1002 employs a statistical language model to generate candidate text interpretations 124 of the voice input 121.

[0182] In one further embodiment, the statistical language model is adapted to look up words, names, and phrases that occur in the various models of the assistant 1002 shown in Figure 8. For example, in at least one embodiment, the statistical language model is given words, names, and phrases from some or all of the words, names, or phrases associated with any node of the domain model 1056 (e.g., words and phrases related to restaurant and dining events), the task flow model 1086 (e.g., words and phrases related to planning events), the dialogue flow model 1087 (e.g., words and phrases related to constraints needed to collect input for restaurant reservations), the domain entity database 1072 (e.g., restaurant names), the glossary database 1058 (e.g., dish names), the service model 1088 (e.g., service provider names such as OpenTable), and / or the active ontology 1050.

[0183] In one embodiment, the statistical language model is adapted to retrieve words, names, and phrases from long-term personal memory 1054. For example, the statistical language model is given text from to-do items, list items, personal notes, calendar entries, names of people in contacts / phone books, email addresses, and street or city names indicated in contacts / phone books.

[0184] The ranking components analyze the candidate interpretations 124 and rank them according to their degree of fit to the syntactic and / or semantic models of the intelligent automated assistant 1002 (126). Any source of constraints on user input may be used. For example, in one embodiment, the assistant 1002 ranks the speech-text interpreter output according to how well those interpretations syntactically and / or semantically parse domain models, task flow models and / or dialogue models, etc. That is, the assistant evaluates how well various combinations of words in the text interpretations 124 fit to the concepts, relationships, entities and characteristics of the active ontology 1050 and its associated models. For example, if the voice-text service 122 generates two candidate interpretations, "italian food for lunch" and "italian shoes for lunch," the semantic relevance ranking 126 will rank "italian food for lunch" higher if it matches more nodes in the assistant 1002's active ontology 1050 (for example, if the words "italian," "food," and "lunch" all match nodes in ontology 1050 and are connected by relationships in ontology 1050, while the word "shoes" does not match in ontology 1050 and matches a node that is not part of the food service domain network).

[0185] In various embodiments, an algorithm or procedure used by the assistant 1002 for interpreting text input, including any embodiment of the natural language processing procedure shown in Figure 28, is used to rank and score candidate text interpretations 124 generated by the speech-to-text service 122.

[0186] In one embodiment, if the ranking component 126 determines that the highest-ranked voice interpretation among the interpretations 124 is ranked above a specified threshold (128), the highest-ranked interpretation is automatically selected (130). If no interpretations are ranked above the specified threshold, the user is presented with possible candidate voice interpretations 134 (132). The user can then select from the displayed options (136).

[0187] In various embodiments, user selection 136 from displayed options is achieved by any input mode, including any mode of multimodal input as described in relation to Figure 16, for example. Such input modes include, but are not limited to, actively derived type input 2610, actively derived voice input 2620, and / or actively presented GUI 2640 for input. In one embodiment, the user can select from candidate interpretations 134, for example, by tapping or speaking. When speaking, the possible interpretations of a new voice input are largely constrained by a small set of options 134 provided. For example, if "Did you mean Italian food or Italian shoes?" is provided, the user simply says "food," and the assistant matches this to "Italian food," not confusing it with other global interpretations of the input.

[0188] Regardless of whether the input is automatically selected (130) or selected by the user (136), the resulting input 138 is returned. In at least one embodiment, the returned input is annotated (138) so that information about the selection made in step 136 is stored along with the text input. This allows, for example, the underlying semantic concepts or entities of the string to be associated with the returned string, improving the accuracy of subsequent language interpretation. For example, if "Italian food" is provided as one of the candidate interpretations 134 based on the semantic interpretation of Cuisine=ItalianFood, the machine-readable semantic interpretation is sent out as an annotated text input 138 along with the user's selection of the string "Italian food".

[0189] In at least one embodiment, the candidate text interpretation 124 is generated based on the speech interpretation received as the output of the speech-to-text service 122.

[0190] In at least one embodiment, candidate text interpretations 124 are generated by paraphrasing the phonetic interpretation in terms of meaning. In some embodiments, there may be multiple paraphrases of the same phonetic interpretation, providing various meanings of words or examples of homophones. For example, if the voice-text service 122 indicates "place for meet," the candidate interpretations presented to the user may be paraphrased as "place to meet (local businesses)" and "place for meat (restaurants)."

[0191] In at least one embodiment, the candidate text interpretation 124 includes a suggestion for modifying a substring.

[0192] In at least one embodiment, the candidate text interpretation 124 includes a proposal to modify a substring of the candidate interpretation using syntactic and semantic analysis as described herein.

[0193] In at least one embodiment, when the user selects a candidate interpretation, it is returned.

[0194] In at least one embodiment, the user is provided with an interface for editing the interpretation before it is returned.

[0195] In at least one embodiment, the user is provided with an interface to continue with further voice input before the input is returned. This allows for the incremental construction of the input utterance, and syntactic and semantic corrections, suggestions, and guidance to be obtained all at once.

[0196] In at least one embodiment, the user is provided with an interface to proceed directly from 136 to step 111 of the method for active type input derivation (described above in relation to Figure 11). This allows for interleaving of type input or voice input and obtaining syntactic and semantic modifications, suggestions, and guidance in a single step.

[0197] In at least one embodiment, the user is provided with an interface to proceed directly from step 111 of an embodiment of active type input derivation to an embodiment of active speech input derivation. This allows for interleaving of type input and speech input, and enables syntactic and semantic corrections, suggestions, and guidance to be obtained in a single step.

[0198] Active GUI input derivation Next, referring to Figure 23, a flowchart is shown illustrating a method for actively deriving input for GUI input according to one embodiment.

[0199] The process is initiated (140). The assistant 1002 presents a graphical user interface (GUI) to the output device 1207 (141), which includes, for example, links and buttons. The user interacts with at least one GUI element (142). The data 144 is received and converted to a unified format (146). The converted data is then returned.

[0200] In at least one embodiment, some elements of the GUI are dynamically generated from an active ontology model rather than being written into a computer program. For example, Assistant 1002 may provide a set of constraints to guide a restaurant reservation service, which is an area on the screen to be tapped. Each area represents a constraint name and / or value. For example, the screen has rows of dynamically generated GUI layouts, each with areas for constraints on cuisine, location, and price range. When the active ontology model is changed, the GUI screen is automatically updated without reprogramming.

[0201] Active Dialogue Proposal Input Derivation Figure 24 is a flowchart illustrating a method for active input derivation at the dialog flow level according to one embodiment. Assistant 1002 proposes possible responses 152 (151). The user selects a proposed response (154). The received input is converted to a unified format (154). The converted data is then returned.

[0202] In at least one embodiment, the suggestion provided in step 151 is provided as a subsequent step in the dialog and / or task flow.

[0203] In at least one embodiment, the proposal provides an option to improve queries using parameters from, for example, domain and / or task models. For example, it is offered to change the location or time of the assumed request.

[0204] In at least one embodiment, the proposal provides an option to select from an ambiguous alternative interpretation given by a language interpretation procedure or component.

[0205] In at least one embodiment, the proposal provides an option to select from an ambiguous alternative interpretation given by a language interpretation procedure or component.

[0206] In at least one embodiment, the proposal provides an option to select from the following steps in the network flow-related dialog flow model 1087. For example, the dialog flow model 1087 suggests that after collecting constraints for one domain (e.g., dining at a restaurant), the assistant 1002 should suggest other related domains (e.g., a nearby movie theater).

[0207] Active monitoring of related events In at least one embodiment, asynchronous events are treated as inputs, just like other modalities of active derivation inputs. Thus, such events are provided as inputs to the assistant 1002. Once interpreted, such events are treated like any other input.

[0208] For example, a change in flight status triggers an alert notification sent to the user. If it indicates a flight delay, Assistant 1002 continues the dialogue by suggesting an alternative flight or making other suggestions based on the detected event.

[0209] Such events can be of any kind. For example, Assistant 1002 may detect that the user has just arrived home, or is lost (has deviated from a designated route), or that a stock price has reached a threshold, or that a TV program of the user's interest is about to begin, or that a musician of the user's interest is coming to the area on tour. In any of these situations, Assistant 1002 proceeds to a dialogue in much the same way that the user would initiate the question themselves. In one embodiment, the event may be based on data provided by another device, for example, to inform the user when a colleague will be back from lunch (the colleague's device can signal such an event to the user's device, at which point Assistant 1002 installed on the user's device will respond accordingly).

[0210] In one embodiment, the event is a notification or alert from a calendar, clock, reminder, or to-do application. For example, an alert from a calendar application regarding a dinner date would initiate a dialogue with Assistant 1002 about the meal event. The dialogue would proceed as if the user had just spoken or typed information about the upcoming meal event, such as "dinner for 2 in San Francisco."

[0211] In one embodiment, the context of a possible event trigger 162 includes information about a person, place, time, and other data. This data is used as part of the input to the assistant 1002, which is used in various steps of the process.

[0212] In one embodiment, data from the context of the event trigger 162 is used to remove ambiguity from voice or text input from the user. For example, if a calendar event alert includes the names of people invited to the event, that information makes it easier to remove ambiguity from input that may match several people with the same or similar names.

[0213] Referring next to Figure 25, a flowchart is shown illustrating a method for actively monitoring relevant events according to one embodiment. In this example, the event trigger events are a set of inputs 162. The assistant 1002 monitors such events (161). The detected events are filtered and sorted for semantic relevance using models, data, and information available from other components of the intelligent automated assistant 1002 (164). For example, if a short-term or long-term memory record for a user indicates that the user is on that flight and / or has questioned the assistant 1002 about that flight, events reporting a change in the flight status are given higher relevance. This storage and filtering presents only significant events for review by the user, who selects one or more events and acts upon them.

[0214] The event data is converted to a unified input format (166) and returned.

[0215] In at least one embodiment, the assistant 1002 proactively provides services associated with proposed events to gain the user's attention. For example, if a flight status alert indicates that the user may miss a flight, the assistant 1002 suggests a task flow to the user to replan the itinerary or book a hotel.

[0216] Examples of input derivation components The following examples are intended to illustrate some of the various functions, operations, actions, and / or other features provided by the active input derivation component 1094. Example: Command completion (What can a user say to Assistant 1002?)

[0217] The user is presented with a text input box containing common commands for entering "what do you want to do?". Depending on the context and user input, one of several system responses is provided. An example is shown below.

[0218] Example: Null input TIFF2026090448000002.tif117157

[0219] Example: Inputting the first word TIFF2026090448000003.tif68157

[0220] Example: Keyword Input TIFF2026090448000004.tif76157

[0221] Example: Instructions for inputting arguments TIFF2026090448000005.tif68157

[0222] Case study: Proposal of standards TIFF2026090448000006.tif84157

[0223] Example: Addition of a reference TIFF2026090448000007.tif108157

[0224] Example: Addition of a location or other constraint TIFF2026090448000008.tif108157

[0225] Example: Start from a constraint, unknown task, or domain TIFF2026090448000009.tif99157

[0226] Example: Name completion Here, the user types some text without accepting any commands or extends the command with an entity name. The system tries to complete the name depending on the context. Furthermore, the system removes domain ambiguity. Example: Words without context TIFF2026090448000010.tif92156

[0227] Example: Names including context TIFF2026090448000011.tif124156

[0228] Example: Selection of a value from a set Here, the user is responding to a system requirement to input a value for a specific parameter such as location, time, cuisine, or genre. The user selects a value from the list or enters a value. As the user types, the matching items in the list are shown as options. Examples are shown below. Example: Selection of a value class TIFF2026090448000012.tif91157

[0229] Example: Reuse of a previous command The queries above are options for auto-completion by the auto-completion interface. These queries are matched as strings (if the input field is empty and there are no known constraints), or they are suggested as relevant in specific situations. Example: Completion in the previous query TIFF2026090448000013.tif83157

[0230] Example: Searching for items in personal memory Assistant 1002 stores specific events and / or entities in personal memory associated with the user. Autocompletion is performed based on such stored items. An example is shown below. Example: Complementary information on events and entities in personal memory TIFF2026090448000014.tif116157

[0231] Multimodal active input derivation In at least one embodiment, the active input derivation component 1094 processes input from multiple input modalities. At least one modality is implemented by an active input derivation procedure that utilizes a method of selecting from a particular type of input and proposed options. As described herein, these are embodiments of the active input derivation procedure for text input, voice input, GUI input, input in the context of a dialog and / or input obtained as a result of an event trigger.

[0232] In at least one embodiment, a single instance of the intelligent automatic assistant 1002 has support for one or more (or any combination thereof) of type input, voice input, GUI input, dialog input, and / or event input.

[0233] Referring next to Figure 26, a flowchart illustrating a method for multimodal active input derivation according to one embodiment is shown. The method is initiated (100). Inputs are received simultaneously from one or more or any combination of input modalities in any order. Thus, the method includes actively deriving type input 2610, voice input 2620, GUI input 2640, input 2650 in the context of a dialog and / or input 2660 obtained as a result of an event trigger. Any or all of these input sources are returned unified in a unified input format 2690. The unified input format 2690 allows other components of the intelligent automatic assistant 1002 to be designed and operate regardless of the specific modality of the input.

[0234] By providing active guidance for multiple modalities and levels, constraints and guidance beyond those available to individual modalities are possible for the input. For example, since the types of suggestions provided for selecting from speech, text, and dialogue steps are independent, their combination shows a significant improvement compared to adding active derivation techniques to individual modalities or levels.

[0235] By combining multiple sources of constraints as described herein (such as syntax / language, glossaries, entity databases, domain models, task models, and service models) with multiple locations where these constraints are actively applied (speech, text, GUI, dialogue, and asynchronous events), a new level of functionality is provided for human-machine interaction.

[0236] Domain Model Component 1056 The components of the domain model 1056 include representations of domain concepts, entities, relationships, characteristics, and instances. For example, the food service domain model 1622 includes the concept of a restaurant, which is a business with a name, address, and telephone number, and the concept of a dining event, which has a party size, date, and time associated with the restaurant.

[0237] In at least one embodiment, the domain model component 1056 of the assistant 1002 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or a combination thereof) of the following: The domain model component 1056 is used by the automated assistant 1002 for several processes, including input derivation 100, natural language interpretation 200, dispatch to services 400, and output generation 600. The domain model component 1056 provides a list of words that match a domain concept or entity such as a restaurant name, which is used for the active derivation of the input 100 and natural language processing 200. • Domain model component 1056 classifies candidate words in the process to determine, for example, that a word is a restaurant name. Domain model component 1056 shows the relationships between parts of information for interpreting natural language, for example, when food is associated with a business entity (for example, "local Mexican food" is interpreted as "find restaurants with style=Mexican", and this inference is possible because of the information in domain model 1056). • Domain model component 1056 organizes information about services used in service orchestration 1082, for example, when a specific web service provides restaurant reviews. The domain model component 1056 provides information for generating natural language paraphrases and other output formats, for example, by providing standard ways of describing concepts, relationships, characteristics, and instances.

[0238] According to certain embodiments, multiple instances or threads of the domain model component 1056 are simultaneously realized and / or started using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the domain model component 1056 are executed, realized and / or started by one or more (or a combination thereof) of the following types of systems, components, systems, devices, procedures and processes: The domain model components 1056 are implemented as data structures representing concepts, relationships, characteristics, and instances. These data structures are stored in memory, files, or databases. Access to the domain model component 1056 is provided via the direct API, network API, and / or database query interface. The creation and maintenance of domain model components 1056 can be achieved, for example, by direct file editing, database transactions, and / or by using domain model editing tools. The domain model component 1056 is implemented as part of or in connection with the active ontology 1050, combining the instantiation of the model for the server and the user with the model itself.

[0239] According to various embodiments, one or more different threads or instances of the domain model component 1056 are started in response to the detection of one or more conditions or events that satisfy various minimum threshold criteria that trigger the start of at least one instance of the domain model component 1056. For example, the start and / or realization of one or more different threads or instances of the domain model component 1056 are triggered when domain model information is needed, including during input derivation, input interpretation, task and domain identification, natural language processing, service orchestration, and / or formatting of output to the user.

[0240] In at least one embodiment, a given instance of the domain model component 1056 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. For example, data from the domain model component 1056 is associated with modeling components of other models including, for example, the glossary 1058, the language pattern recognizer 1060, the dialog flow model 1087, the task flow model 1086, the service function model 1088, and the domain entity database 1072. For example, a business in the domain entity database 1072 classified as a restaurant is recognized by a type identifier maintained in the dine-out domain model component.

[0241] Examples of domain model components Referring now to FIG. 27, a set of screenshots is shown that illustrate an example of various functions, operations, actions, and / or other features provided by the domain model component 1056 according to one embodiment.

[0242] In at least one embodiment, the domain model component 1056 is a unified data representation that enables the representation of the information shown in the restaurant-related screens 103A and 103B, which combines data from several distinct data sources and services, for example, name, address, business category, phone number, identifier for saving to long-term personal memory, identifier for sharing via email, reviews from multiple sources, map coordinates, and personal notes, etc.

[0243] Language interpreter component 1070 In at least one embodiment, the language interpreter component 1070 of the assistant 1002 is operable to perform and / or implement various functions, operations, actions, and / or other features such as, for example, one or more of the following (or combinations thereof).

[0244] • Analyzes user input and identifies the set of syntactic analysis results. User input includes any information from the user's device context that helps to understand the user and their intent, including, for example, a sequence of words, matching gestures or GUI elements involved in deriving the input, the current context of the dialog, the current device application and its current data object, and / or one or more (or a combination thereof) of any other dynamic personal data obtained about the user, such as location and time. For example, in one embodiment, user input is in the form of a unified annotated input format 2690 obtained as a result of active input derivation 1094. The parsing result is an association between user input data and data structures in other representations of user input, including concepts, relationships, characteristics, instances, and / or other nodes, as well as / or models, databases, and / or user intent and / or context. The association in the parsing result is a complex mapping from a set and sequence of words, signals, and other elements of user input to one or more associated concepts, relationships, characteristics, instances, other nodes, and / or data structures as described herein.

[0245] The system analyzes user input and identifies a set of syntactic parsing results, which are parsing results that associate user input data with structures representing syntactic parts of speech, clauses, and phrases, including the names of multiple words, sentence structure, and / or other grammatical graph structures. The syntactic parsing results are described in element 212 of the natural language processing procedure, which is explained in relation to Figure 28.

[0246] - The system parses user input and identifies a set of semantic parsing results, which are parsing results that relate the user input data to structures representing concepts, relationships, characteristics, entities, quantities, propositions, and / or other expressions of meaning and user intent. In one embodiment, these expressions of meaning and intent are represented by a set of models or databases and / or elements and / or instances, and / or nodes of an ontology, as described in element 220 of the natural language processing procedure described in relation to Figure 28.

[0247] • Remove ambiguity between different syntactic or semantic parsing results, as described in element 230 of the natural language processing procedure, which is explained in relation to Figure 28.

[0248] • Determine whether partially typed input is syntactically and / or semantically meaningful in an autocompletion procedure, such as the one described in relation to Figure 11.

[0249] • To facilitate the generation of the proposed completion 114 in an autocompletion procedure, as explained in relation to Figure 11.

[0250] • In a voice input procedure as described in relation to Figure 22, it is determined whether the interpretation of the voice input is syntactically and / or semantically meaningful.

[0251] According to a particular embodiment, multiple instances or threads of the language interpreter component 1070 are realized and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software.

[0252] According to various embodiments, one or more different threads or instances of the language interpreter component 1070 are started in response to the detection of one or more conditions or events that satisfy one or more different minimum threshold criteria that trigger the start of at least one instance of the language interpreter component 1070. Various examples of conditions or events that trigger the start and / or realization of one or more different threads or instances of the language interpreter component 1070 include, but are not limited to, one or more (or combinations thereof) of the following:

[0253] While deriving inputs that include, but are not limited to, the following: • We propose auto-completion that can be done via typing (114) (Figure 11). • Rank the interpretation of the speech (126) (Figure 22). • When providing ambiguity as proposed response 152 in a dialogue (Figure 24).

[0254] When the result of deriving the input is available, including when the input is derived by any mode of the active multimodal input derivation 100.

[0255] In at least one embodiment, a given instance of the language interpreter component 1070 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of such database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the language interpreter component include, but are not limited to, one or more (or combinations thereof) of the following:

[0256] • Domain Model 1056 ·Glossary 1058 • Domain entity database 1072 • Short-term memory 1052 • Long-term personal memory 1054 • Task flow model 1086 • Dialogue flow model 1087 • Service function model 1088

[0257] Referring further to Figure 29, a screenshot illustrating natural language processing according to one embodiment is shown. The user inputs language input 2902 consisting of the phrase "who is playing this weekend at the filmore" (by voice or text). This phrase is echoed back to the user on screen 2901. The language interpreter component 1070 processes the input 2902 and generates a parsing result. The parsing result associates the input with a request indicating a local event scheduled on one of the days of the next weekend at any event location whose name matches "filmore". A paraphrase of the parsing result is shown on screen 2901 as 2903.

[0258] Next, referring to Figure 28, a flowchart is shown illustrating an example of a natural language processing method according to one embodiment.

[0259] The process is initiated (200). Language input 202 is received, such as the string "who is playing this weekend at the filmore" in the example in Figure 29. In one embodiment, the input is supplemented with current context information, such as the user's current location and local time. In word / phrase matching 210, the language interpreter component 1070 finds associations between user input and concepts. In this example, associations are found between the string "playing" and the concept of a list of locations for events, between the string "this weekend" (along with the user's current local time) and specific examples of approximate time periods representing the next weekend, and between the string "filmore" and the name of the location. Word / phrase matching 210 uses data from, for example, the language pattern recognition unit 1060, the glossary database 1058, the active ontology 1050, the short-term personal memory 1052, and the long-term personal memory 1054.

[0260] The language interpreter component 1070 generates a syntactic candidate parsing 212 that includes the selected parsing result and further includes other parsing results. For example, the other parsing results include a result in which "playing" is associated with other domains such as games or categories of events such as sports events.

[0261] Short-term memory 1052 and / or long-term memory 1054 are used by the language interpreter component 1070 when generating syntactic candidate parsings 212. Therefore, known information about inputs and / or users previously applied in the same session is used to improve performance, reduce ambiguity, and enhance the conversational nature of the dialogue. Data from the active ontology 1050, domain model 1056, and task flow model 1086 are used to enable evidence-based reasoning when determining a valid syntactic candidate parsing 212.

[0262] In semantic matching 220, the language interpreter component 1070 considers possible combinations of parsing results according to their degree of fit to semantic models such as the domain model and database. In this case, parsing involves associating (1) "playing" (a word in user input) which is "Local Event At Venue" (part of the domain model 1056 represented by a cluster of nodes in the active ontology 1050) with (2) "filmore" (another word in input) which matches the entity name in the domain entity database 1072 for the location of the local event, represented by elements of the domain model and nodes (Venue Name) in the active ontology.

[0263] Semantic matching 220 uses data from, for example, the active ontology 1050, short-term personal memory 1052, and long-term personal memory 1054. For example, semantic matching 220 uses data from previous references to locations or local events in a dialog (from short-term personal memory 1052) or to a person's favorite locations (from long-term personal memory 1054).

[0264] A set of semantic candidate syntactic analysis results or semantic potential syntactic analysis results 222 is generated.

[0265] In the deambiguation step 230, the language interpreter component 1070 considers the basis for the semantic candidate syntactic analysis result 222. In this example, the combination of parsing "playing," which is "Local Event At Venue," and matching it with "filmore," which is Venue Name, matches the domain model at a higher level than, for example, another combination in which "playing" is associated with the sports domain model and "filmore" has no association in the sports domain.

[0266] The deambiguation 230 uses data from, for example, the structure of the active ontology 1050. In at least one embodiment, the connections between nodes in the active ontology provide evidence-based support for deambiguating the semantic candidate parsing results 222. For example, in one embodiment, if three nodes in the active ontology semantically match and are all connected in the active ontology, this provides a stronger basis for semantic parsing compared to the case where those matching nodes are not connected or are connected by a longer connection path in the active ontology 1050. For example, in one embodiment of semantic matching 220, parsing matching Local Event At Venue and Venue Name is given stronger evidence-based support because the combined representation of those aspects of the user's intent is connected by relationships and / or links in the active ontology 1050. In this instance, the Local Event node is connected to the Venue node, the Venue node is connected to the Venue Name node, and the Venue Name node is connected to the entity name in the database of place names.

[0267] In at least one embodiment, the connections between nodes of an active ontology that provide evidence-based support to eliminate ambiguity between semantic candidate parsing results 222 are directed arcs that form an inference grid, where matching nodes provide evidence for the nodes connected by the directed arcs.

[0268] In 232, the language interpreter component 1070 sorts and selects the top-level semantic syntactic analysis as the expression of the user's intent 290 (232).

[0269] Domain Entity Database 1072 In at least one embodiment, the domain entity database 1072 is operable to perform and / or realize various functions, operations, behaviors and / or other features, such as one or more (or a combination thereof) of the following:

[0270] • Stores data about domain entities. Domain entities are things in the world or computing environment that are modeled in a domain model. Examples include, but are not limited to, one or more (or combinations thereof) of the following:

[0271] · All kinds of business • Movies, videos, songs and / or other music products, and / or any other designated entertainment products · All kinds of products ·event · Calendar Entry • Cities, states, countries, regions, and / or other geographical, geopolitical, and / or geospatial points or territories • Landmarks and designated locations such as airports • Provide database services to those databases. This includes, but is not limited to, simple queries, complex queries, transactions, and triggered events.

[0272] According to certain embodiments, multiple instances or threads of the domain entity database 1072 are simultaneously realized and / or started using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the domain entity database 1072 are executed, realized and / or started by database software and / or hardware residing in the client 1304 and / or server 1340.

[0273] An example of a domain entity database 1072 used in connection with the present invention according to one embodiment is a database of one or more businesses that store, for example, names and locations. The database is used, for example, to refer to words contained in an input request for matching businesses and / or to refer to the location of a business whose name is known. It will be understood by those skilled in the art that many other configurations and implementations are possible.

[0274] Glossary component 1058 In at least one embodiment, the glossary component 1058 is operable to perform and / or realize various functions, operations, actions and / or other features such as one or more (or combinations thereof) of the following:

[0275] • Provides a database that associates concepts, characteristics, relationships, or instances of domain models or task models with words and strings.

[0276] The glossary, derived from the glossary components, is used by the automated assistant 1002 for several processes, including, for example, input derivation, natural language interpretation, and output generation.

[0277] According to certain embodiments, multiple instances or threads of the glossary component 1058 are implemented and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the glossary component 1058 are implemented as data structures associating names and strings of concepts, relationships, characteristics and instances. These data structures are stored in memory, files or databases. Access to the glossary component 1058 is implemented via direct APIs, network APIs and / or database query interfaces. Creation and maintenance of the glossary component 1058 are achieved by direct editing of files, database transactions or the use of domain model editing tools. The glossary component 1058 is implemented as part of or in association with the active ontology 1050. It will be understood by those skilled in the art that many other configurations and implementations are possible.

[0278] According to various embodiments, one or more different threads or instances of the glossary component 1058 are started in response to the detection of one or more conditions or events that satisfy one or more different minimum threshold criteria, which trigger the start of at least one instance of the glossary component 1058. In one embodiment, the glossary component 1058 is accessed whenever glossary information is requested, including, for example, during input derivation, input interpretation, and output formatting for the user. It will be understood by those skilled in the art that other conditions or events may trigger the start and / or realization of one or more different threads or instances of the glossary component 1058.

[0279] In at least one embodiment, a given instance of the glossary component 1058 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. In one embodiment, the glossary component 1058 accesses data from an external database, such as a data warehouse or dictionary.

[0280] Language pattern recognition component 1060 In at least one embodiment, the language pattern recognition component 1060 is operable to perform and / or realize various functions, operations, actions and / or other features, such as searching for patterns in language or speech input that represent grammatical compound words, idiomatic compound words and / or other compound words of input tokens. These patterns correspond to one or more (or combinations thereof) of words, names, phrases, data, parameters, commands and / or signals of speech acts.

[0281] According to certain embodiments, multiple instances or threads of the pattern recognition component 1060 are simultaneously implemented and / or started using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the language pattern recognition component 1060 are executed, implemented and / or started by one or more files, databases and / or programs containing representations in a pattern matching language. In at least one embodiment, the language pattern recognition component 1060 is represented declaratively rather than in program code. This allows the component to be created and maintained by tools other than editors and programming tools. Examples of declarative representations include, but are not limited to, one or more (or combinations thereof) of regular expressions, pattern matching rules, natural language grammars, state transition machine-based parsers and / or other parsing models.

[0282] Those skilled in the art will understand that other types of systems, components, devices, procedures, and processes (or combinations thereof) may be used to realize the language pattern recognition component 1060.

[0283] According to various embodiments, one or more different threads or instances of the language pattern recognition component 1060 are started in response to the detection of one or more conditions or events that satisfy one or more different minimum threshold criteria that trigger the start of at least one instance of the language pattern recognition component 1060. Various examples of conditions or events that trigger the start and / or realization of one or more different threads or instances of the language pattern recognition component 1060 include, but are not limited to, one or more (or combinations thereof) of the following: The structure of the language pattern recognition system constrains and guides user input during the active derivation of input. • The language pattern recognition system is performing natural language processing to make the input easier to interpret as language. • During task and dialogue identification, the language pattern recognition system facilitates the identification of tasks, dialogues, and / or steps.

[0284] In at least one embodiment, a given instance of the language pattern recognition component 1060 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various data accessed by the language pattern recognition component 1060 include, but are not limited to, data from any of the various models and data sources that are part of embodiments of the assistant 1002. This includes, but is not limited to, one or more (or combinations thereof) of the following:

[0285] • Domain Model 1056 ·Glossary 1058 • Domain entity database 1072 • Short-term memory 1052 • Long-term personal memory 1054 • Task flow model 1086 • Dialogue flow model 1087 • Service function model 1088

[0286] In one embodiment, data access from other parts of the Assistant 1002 embodiment is coordinated by the Active Ontology 1050.

[0287] Referring to Figure 14, some examples of the various functions, operations, behaviors, and / or other features provided by the language pattern recognition component 1060 are shown. Figure 14 shows language patterns recognized by the language pattern recognition component 1060. For example, the phrase "what is happening" (in a city) is associated with the task of planning events and the domain of local events.

[0288] Dialogflow Processor Component 1080 In at least one embodiment, the dialog flow processor component 1080 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or combinations thereof) of the following:

[0289] Based on the language interpretation 200, assume an expression of the user's intent 290 to identify the task the user wants to perform and / or the problem the user wants to solve. For example, the task is to find a restaurant.

[0290] • For a given problem or task, an expression of the user's intent 290 is assumed, and parameters for the task or problem are identified. For example, the user is looking for a recommended restaurant that serves Italian food near the user's home. The constraints that the restaurant is recommended, serves Italian food, and is near the home are parameters for the task of finding a restaurant.

[0291] - Assuming the interpretation of the task and the current dialogue with the user as represented in personal short-term memory 1052, select an appropriate dialogue flow model and determine the step of the flow model corresponding to the current state.

[0292] According to a particular embodiment, multiple instances or threads of the dialog flow processor component 1080 are realized and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software.

[0293] In at least one embodiment, a given instance of the Dialogflow processor component 1080 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the Dialogflow processor component 1080 include, but are not limited to, one or more (or combinations thereof) of the following: • Task flow model 1086 • Domain Model 1056 • Dialogue flow model 1087

[0294] Next, referring to Figures 30 and 31, screenshots are shown illustrating examples of various functions, operations, behaviors and / or other features provided by the Dialogflow Processor component according to one embodiment.

[0295] As shown in screen 3001, the user requests a dinner reservation by providing the voice or text input "book me a table for dinner". Assistant 1002 generates a prompt 3003 requesting the user to specify the time and party size.

[0296] Once these parameters are provided, screen 3101 is displayed. Assistant 1002 outputs a dialog box 3102 indicating that the results are presented, and a prompt 3103 asking the user to click on a time. List 3104 is then displayed.

[0297] In one embodiment, such a dialog is implemented as follows: The dialog flow processor component 1080 is given a representation of the user's intent from the language interpreter component 1070 and determines that the appropriate response is to request information from the user that is necessary to perform the next step in the task flow. In this case, the domain is a restaurant, the task is to make a reservation, and the dialog step is to request information from the user that is necessary to accomplish the next step in the task flow. This dialog step is exemplified by prompt 3003 on screen 3001.

[0298] Referring to Figure 32, a flowchart is shown illustrating the operation method of the dialog flow processor component 1080 according to one embodiment. The flowchart in Figure 32 will be explained in relation to the examples shown in Figures 30 and 31.

[0299] The method is initiated (200). A user intent expression 290 is received. As described in relation to Figure 28, in one embodiment, the user intent expression 290 is a set of semantic syntactic parses. In the example shown in Figures 30 and 31, the domain is restaurant, the verb is "book" associated with making a restaurant reservation, and the time parameter is the night of the current date.

[0300] In step 310, the dialog flow processor component 1080 determines whether the interpretation of this user intent is strongly supported and / or better supported than another ambiguous parsing. In this example, the interpretation is strongly supported and there are no conflicting ambiguous parsings. On the other hand, if there is a conflicting ambiguous parsing or a very uncertain parsing, step 322 is performed to configure the dialog flow step so that the execution stage prompts the dialog for further information from the user.

[0301] In step 312, the dialog flow processor component 1080 determines a suitable interpretation of semantic parsing using other information to determine the task to be executed and its parameters. The information is obtained, for example, from the domain model 1056, the task flow model 1086 and / or the dialog flow model 1087, or any combination thereof. In this example, the task is identified as making a reservation, which involves finding a place that is available and reservable and performing a transaction to reserve the table. The task parameters are time constraints in addition to the other parameters inferred in step 312.

[0302] In 320, the task flow model is examined to determine the appropriate next step. Information is obtained, for example, from the domain model 1056, the task flow model 1086 and / or the dialogue flow model 1087, or any combination thereof. In this example, it is determined that the next step in this task flow is to derive the missing parameters for the restaurant availability search, resulting in the prompt 3003 shown in Figure 30, which requests the party size and reservation time.

[0303] As described above, Figure 31 shows a screen 3101 that includes a dialog element 3102 presented after the user responds to a request for party size and reservation time. In one embodiment, screen 3101 is presented as a result of another single execution via an automated call and response procedure, as described in conjunction with Figure 33, which triggers another call to the dialog and flow procedure shown in Figure 32. At the start of this dialog and flow procedure, after receiving the user's preferences, the dialog flow processor component 1080 determines a different task flow step in step 320 to perform an availability search. Once the request 390 is constructed, it includes enough task parameters for the dialog flow processor component 1080 and the service orchestration component 1082 to dispatch to the restaurant reservation service.

[0304] Dialog Flow Model Component 1087 In at least one embodiment, the dialog flow model component 1087 is operable to provide a dialog flow model that represents steps taken in a particular type of conversation between the user and the intelligent automated assistant 1002. For example, a dialog flow for a generic task that performs a transaction includes the steps of obtaining the data required for the transaction and verifying the transaction parameters before committing it.

[0305] Task flow model component 1086 In at least one embodiment, the task flow model component 1086 is operable to provide a task flow model, which represents steps taken to solve a problem or address a request. For example, a task flow for making a dinner reservation includes finding a desired restaurant, checking its availability, and making a transaction to make a reservation at that restaurant for a specific time.

[0306] According to certain embodiments, multiple instances or threads of the task flow model component 1086 are realized and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the task flow model component 1086 are realized by a program, a state transition machine or other method that identifies the appropriate step in the flow graph.

[0307] In at least one embodiment, the task flow model component 1086 uses a task modeling framework called a generic task. A generic task is an abstract concept that models the steps of a task and their requested inputs and generated outputs, without being domain-specific. For example, a generic task for a transaction includes the steps of collecting the data required for the transaction, executing the transaction, and outputting the transaction results, all of which do not refer to any particular transaction domain or service to accomplish. A generic task is instantiated for a domain such as shopping, but is independent of the shopping domain and is equally appropriate for domains such as reservations and scheduling.

[0308] At least a portion of the functions, operations, actions and / or other features associated with the task flow model component 1086 and / or procedure described herein are at least partially realized using the concepts, features, components, processes and / or other embodiments described herein in relation to the generic task modeling framework.

[0309] Furthermore, at least a portion of the functions, operations, actions and / or other features associated with the task flow model components 1086 and / or procedures described herein are at least partially realized using concepts, features, components, processes and / or other embodiments related to conditional selection tasks as described herein. For example, one embodiment of a generic task is realized using a conditional selection task model.

[0310] In at least one embodiment, a given instance of the task flow model component 1086 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the task flow model component 1086 include, but are not limited to, one or more (or a combination thereof) of the following:

[0311] • Domain Model 1056 ·Glossary 1058 • Domain entity database 1072 • Short-term memory 1052 • Long-term personal memory 1054 • Dialogue flow model 1087 • Service function model 1088

[0312] Referring to Figure 34, a flowchart is shown illustrating an example of the task flow of a conditional selection task 351 according to one embodiment.

[0313] Conditional selection is a type of generic task whose objective is to select an item from a set of items worldwide based on a set of constraints. For example, conditional selection task 351 is instantiated for the domain of restaurants. Conditional selection task 351 is initiated by requesting criteria and constraints from user 352. For example, the user is interested in Asian food and wants to find a place to eat near their office.

[0314] In step 353, the assistant 1002 presents items that meet the explicitly stated criteria and constraints for the user to view. In this example, this may be a list of restaurants and their characteristics used for selection.

[0315] In step 354, the user is given the opportunity to improve the criteria and constraints. For example, the user improves the requirement by saying "near my office." The system then presents the new set of results in step 353.

[0316] Referring to Figure 35, an example of a screen 3501 is shown, which includes a list 3502 of items presented by a conditional selection task 351 according to one embodiment.

[0317] In step 355, the user selects from a set of matching items. One of several subsequent tasks 359 is made available, for example, reservation 356, save 357, or share 358. In various embodiments, the subsequent tasks 359 include interacting with web-enabled services and / or device-local functionality (such as setting calendar appointments, making calls, sending emails or text messages, and setting alarms).

[0318] In the example in Figure 35, the user selects an item in List 3502 to view further details and take additional action. Referring to Figure 36, an example of screen 3601 after the user has selected an item from List 3502 is shown. Additional information and options corresponding to the subsequent task 359 regarding the selected item are displayed.

[0319] In various embodiments, the flow step is provided to the user in any of several input modalities, including but not limited to any combination of explicit dialog prompts and GUI links.

[0320] Service component 1084 The service component 1084 represents a set of services that the intelligent automatic assistant 1002 invokes on behalf of the user. Any of the invoked services are provided within the service component 1084.

[0321] In at least one embodiment, the service component 1084 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or combinations thereof) of the following:

[0322] Typically, functionality is provided via APIs delivered through a web-based user interface for the service. For example, a review website might provide a service API that automatically returns reviews of a given entity when called programmatically. The API provides the intelligent automated assistant 1002 with services obtained by the user interacting with the website's user interface.

[0323] Typically, functionality is provided through APIs offered by the user interface to the application. For example, a calendar application provides a service API that automatically returns calendar entries when called programmatically. The API provides the intelligent automatic assistant 1002 with services obtained by the user operating the application's user interface. In one embodiment, the assistant 1002 can initiate and control any of the many different functions available to the device. For example, if the assistant 1002 is installed on a smartphone, personal digital assistant, tablet computer, or other device, the assistant 1002 can perform functions such as initiating applications, making calls, sending emails and / or text messages, adding calendar events, and setting alarms. In one embodiment, such functions are enabled using a service component 1084.

[0324] • Provides services that are not currently implemented in the user interface but are available to the assistant via API for larger tasks. For example, in one embodiment, an API that obtains an address and returns machine-readable geographic coordinates is used by the assistant 1002 as a service component 1084 even if it does not have a direct user interface on the web or device.

[0325] According to certain embodiments, multiple instances or threads of the service component 1084 are simultaneously realized and / or started using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the service component 1084 are executed, realized and / or started by one or more (or a combination thereof) of the following types of systems, components, systems, devices, procedures and processes:

[0326] • Implementation of APIs as defined by the service, whether local, remote, or any combination. • The database must be included within the database service for which the automated assistant 1002 or assistant 1002 is available.

[0327] For example, a website that provides a user with an interface for viewing movies is used by one embodiment of the intelligent automated assistant 1002 as a copy of the database used by the website. The service component 1084 provides the data with internal APIs as if they were provided via network APIs, even though the data is stored locally.

[0328] As another example, a service component 1084 of an intelligent automated assistant 1002 that assists in restaurant selection and meal planning includes any or all of the following sets of services available from third parties via a network: A collection of restaurant listing services that display a list of restaurants matching by name, location, or other criteria. A collection of restaurant rating services that return a ranking for a specified restaurant. A collection of restaurant review services that provide written reviews for designated restaurants. • Geocoding service to locate restaurants on a map. • A reservation service that enables the booking of restaurant tables based on a program.

[0329] Service orchestration component 1082 The service orchestration component 1082 of the intelligent automated assistant 1002 executes a service orchestration procedure.

[0330] In at least one embodiment, the service orchestration component 1082 is operable to perform and / or realize various functions, operations, actions and / or other features such as one or more (or combinations thereof) of the following: • Dynamically and automatically determines services that meet user requests and / or specified domains and tasks. • Dynamically and automatically invoke multiple services in any combination of concurrent and ordered operations. • Dynamically and automatically convert task parameters and constraints to meet the input requirements of the service API. • Dynamically and automatically monitor and collect results from multiple services. • Dynamically and automatically merge service result data from various services into a unified result model. • Orchestrate multiple services to meet the requirements constraints. • Orchestrate multiple services to annotate an existing set of results with supplementary information.

[0331] • Output the results of calling multiple services in a service-independent, unified representation that unifies the results from various services (for example, as a result of calling several restaurant services that return a list of restaurants, merge data on at least one restaurant from several services, and remove redundancy).

[0332] For example, in a given situation, there are several ways to accomplish a particular task. For instance, user input such as "remind me to leave for my meeting across town at 2pm" specifies an action that can be accomplished in at least three ways: setting an alarm clock, creating a calendar event, or calling a to-do manager. In one embodiment, the service orchestration component 1082 makes a determination regarding the best way to satisfy the request.

[0333] The service orchestration component 1082 also determines the best combination of services to call to perform a given task. For example, to find and reserve a table for dinner, the service orchestration component 1082 determines the services to call to perform the functions of reviewing, retrieving availability, and making a reservation. The determination of which services to use depends on one of many different factors. For example, in at least one embodiment, information such as reliability, the service's ability to handle a particular type of request, and user feedback are used as factors in determining which services are appropriate to call.

[0334] According to a particular embodiment, multiple instances or threads of the service orchestration component 1082 are realized and / or started simultaneously using one or more processors, and / or hardware and / or other combinations of hardware and software.

[0335] In at least one embodiment, a given instance of the service orchestration component 1082 uses an explicit service function model 1088 to represent the functions and other characteristics of an external service and infers those functions and characteristics while achieving the features of the service orchestration component 1082. This provides advantages over manually programming a set of services, including, for example, one or more (or a combination thereof) of the following:

[0336] • Ease of development • Robustness and reliability during execution • Ability to dynamically add and remove services without interfering with the code. • The ability to implement general distributed query optimization algorithms driven by characteristics and functionality, rather than hardcoding for specific services or APIs.

[0337] In at least one embodiment, a given instance of the service orchestration component 1082 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Examples of the various types of data accessed by the service orchestration component 1082 include, but are not limited to, one or more (or a combination thereof) of the following:

[0338] • Instantiation of the domain model Syntactic and semantic parsing of natural language input • Instantiation of the task model (based on values ​​for parameters) • Dialogue and task flow models, and / or selected steps within those models • Service function model 1088 • Any other information available in Active Ontology 1050

[0339] Referring to Figure 37, an example of the procedure for performing a service orchestration procedure according to one embodiment is shown.

[0340] In this particular example, we assume a single user is interested in finding a suitable place for dinner at a restaurant and is using an intelligent automated assistant 1002 in conversation to facilitate this service.

[0341] Consider the task of finding a restaurant that is upscale, well-regarded, close to a specific location, can be booked at a specific time, and serves a specific type of food. These domain and task parameters are given as input 390.

[0342] The process is initiated (400). In 402, it is determined whether a given request requires any service. In some situations, for example, if Assistant 1002 can perform the desired task itself, service delegation is not requested. For example, in one embodiment, Assistant 1002 can respond to the actual question without calling service delegation. Therefore, if the request does not require a service, a standalone flow step is performed in 403, and as a result, 490 is returned. For example, if the task request is to request information about the automated Assistant 1002 itself, the dialog response is handled without calling any external services.

[0343] If it is determined in step 402 that service delegation is required, the service orchestration component 1082 proceeds to step 404. In 404, the service orchestration component 1082 matches the declarative description of the service's functions and characteristics in the service function model 1088 with the task requirements. At least one service provider supporting instantiated behavior provides declarative and qualitative metadata details, such as one or more (or a combination thereof) of the following:

[0344] • Data fields from which results are returned • Classes of parameters that the service provider is known to statically support. • Policy functions for parameters that the service provider can support after dynamic inspection of parameter values • Performance evaluation that defines how the service is executed (e.g., relational database, web service, triple store, full-text index, or any combination thereof) • Characteristic quality evaluation that statically defines the expected quality of the characteristic values ​​returned with the result object. • Assessment of the overall quality of the results that the service is expected to return.

[0345] For example, by inferring the classes of parameters that a service supports, the service model reveals that services 1, 2, 3, and 4 provide restaurants near a specific location (parameter), services 2 and 3 filter or rank restaurants by quality (another parameter), services 3, 4, and 5 return reviews for restaurants (returned data field), service 6 lists the types of food offered by restaurants (returned data field), and service 7 checks the availability of restaurants for a specific time period (parameter). Services 8-99 provide functionality not required for this particular domain and task.

[0346] Using this declarative and qualitative metadata, tasks, task parameters, and other information available from the Assistant's execution environment, the service orchestration component 1082 determines the best set of service providers to invoke (404). The best set of service providers supports one or more task parameters (returning results that satisfy one or more parameters) and takes into account the performance evaluation of at least one service provider and the overall quality evaluation of at least one service provider.

[0347] The result of step 404 is a list of dynamically generated services to call for this specific user and request.

[0348] In at least one embodiment, the service orchestration component 1082 takes into account the reliability of the service and its ability to respond to specific information requests.

[0349] In at least one embodiment, the service orchestration component 1082 avoids inconsistency by calling overlapping or redundant services.

[0350] In at least one embodiment, the service orchestration component 1082 takes into account personal information about the user (short-term personal memory component) to select a service. For example, the user may prefer a certain evaluation service over others.

[0351] In step 450, the service orchestration component 1082 dynamically and automatically invokes multiple services for the user. In at least one embodiment, these services are invoked dynamically in response to user requests. According to a particular embodiment, multiple instances or threads of a service are invoked simultaneously. In at least one embodiment, these are invoked over a network using an API, over a network using a web service API, over the internet using a web service API, or any combination thereof.

[0352] In at least one embodiment, the frequency with which the service is invoked is limited and / or controlled on a programmatic basis.

[0353] Referring to Figure 38, an example of a service call procedure 450 according to one embodiment is shown. A service call is used, for example, to obtain additional information or to perform a task using an external service. In one embodiment, request parameters are translated into a format appropriate for the service's API. Once results are received from the service, these results are translated into a result representation for presentation to the user within the assistant 1002.

[0354] In at least one embodiment, the service invoked by the service invocation procedure 450 is a web service, an application or operating system function running on the device, etc.

[0355] The request expression 390 is provided, for example, including task parameters. For at least one service available from the service function model 1088, the service call procedure 450 performs a conversion step 452, an invocation step 454, and an output mapping step 456.

[0356] In conversion step 452, the current task parameters of the request expression 390 are converted to a format used by at least one service. Parameters for services provided as an API or database are different from, and at least different from, the data representation used in the task request. Thus, the objective of step 452 is to map at least one task parameter to one or more corresponding formats and values ​​in at least one service being called.

[0357] For example, the names of businesses such as restaurants differ among the services that cater to such businesses. Therefore, step 452 includes converting every name into a format that is best suited to at least one service.

[0358] As another example, locations are known with varying degrees of precision using different units and rules across services. Service 1 requires a postal code, Service 2 requires GPS coordinates, and Service 3 requires a postal address.

[0359] The service is invoked via an API (454) and its data is collected. In at least one embodiment, the results are cached. In at least one embodiment, services that do not return within a specified level of performance (e.g., as specified in a service quality assurance scheme, i.e., an SLA) are dropped.

[0360] In the output mapping step 456, the data returned by the service is mapped again to the unified result representation 490. This step includes processing various formats and units.

[0361] In step 412, results from multiple services are inspected and merged. In one embodiment, when inspected results are collected, a domain-defined equivalence policy function is called for each pair of results to determine which results represent identical concepts in the real world. If pairs of equivalent results are found, a domain-defined set of characteristic policy functions is used to merge the characteristic values ​​into the merged result. The characteristic policy function determines the optimal merge strategy using an evaluation of characteristic quality from the service function model, task parameters, domain context and / or long-term personal memory 1054.

[0362] For example, lists of restaurants from various restaurant providers are merged, and duplicates are removed. In at least one embodiment, the criteria for identifying duplicates include fuzzy name matching, fuzzy location matching, fuzzy matching to multiple characteristics of a domain entity such as name, location, telephone number and / or website address, and / or combinations thereof.

[0363] In step 414, the results are sorted and removed to return a result list of the desired length.

[0364] In at least one embodiment, a request relaxation loop is further applied. If the service orchestration component 1082 determines in step 416 that the current result list is insufficient (for example, the current result list has fewer items than the desired number of matching items), the task parameters are relaxed to allow for more results (420). For example, if there are too few restaurants of the desired type that can be found within N miles from the target location, the request is relaxed to look at an area wider than N miles and / or relax other parameters of the search.

[0365] In at least one embodiment, the service orchestration method is applied during a second pass to annotate the results with auxiliary data useful for the task.

[0366] In step 418, the service orchestration component 1082 determines whether an annotation is needed. For example, an annotation is needed if a task requires plotting results on a map, but the main service does not return the geographic coordinates necessary for mapping.

[0367] In 422, the service function model 1088 is re-examined to find a service that returns the desired additional information. In one embodiment, annotation processing determines whether to annotate the merged result with additional or more appropriate data. Annotation processing makes this determination by delegating to a domain-defined characteristic policy function for at least one characteristic of at least one merged result. The characteristic policy function uses an evaluation of the merged characteristic value and characteristic quality, an evaluation of the characteristic quality of one or more other service providers, the domain context and / or user profile to determine whether more appropriate data is obtained. If it is determined that one or more service providers annotate one or more characteristics of the merged result, a cost function is called to determine the set of service providers best suited to perform the annotation.

[0368] At least one service provider from the set of optimal annotation service providers is invoked using the list of merge results to obtain result 424 (450). Changes made to at least one merge result by at least one service provider are tracked during this process, and the changes are merged using the same characteristic policy function processing used in step 412. Those results are merged into the existing set of results (426).

[0369] The resulting data is sorted (428) and standardized to unified representation 490.

[0370] One advantage of the methods and systems described above with respect to the service orchestration component 1082 is that they can be advantageously applied and / or utilized in various technical fields other than those particularly related to intelligent automated assistants. Examples of other technical fields of the aspects and / or features of the service orchestration procedure include, for example, one or more of the following:

[0371] • Dynamic "mashups" in websites, as well as web-based applications and services. • Optimization of distributed database queries • Dynamic service-oriented architecture configuration

[0372] Service Function Model Component 1088 In at least one embodiment, the service function model component 1088 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or a combination thereof) of the following: • Provides machine-readable information about the service's functionality in order to perform calculations of a specific class. • Provide machine-readable information about the service's functionality to answer queries for a specific class. • Provides machine-readable information about the classes of transactions offered by various services. • Provides machine-readable information regarding parameters for APIs provided by various services. • Provides machine-readable information about parameters used in database queries against databases provided by various services.

[0373] Output processor component 1090 In at least one embodiment, the output processor component 1090 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or combinations thereof) of the following:

[0374] - Format the output data, represented by a unified internal data structure, into a format and layout that allows it to be appropriately rendered in various modalities. The output data includes, for example, natural language communication between an intelligent automated assistant and the user, data on domain entities such as restaurants, movies, and product characteristics, domain-specific data results from information services such as weather forecasts, flight status checks, and prices, and / or interactive links and buttons that allow the user to respond by directly interacting with the output representation.

[0375] • Render output data for modalities, including, for example, any combination of graphical user interfaces, text messages, email messages, sound, animation, and / or audio output.

[0376] • Dynamically renders data to various graphical user interface display engines based on requests. For example, it uses various output processing layouts and formats that depend on the web browser and / or device being used.

[0377] • Dynamically renders output data using various speech sounds.

[0378] • Dynamically renders the image using the modality specified based on the user's preferences.

[0379] • Dynamically render output using user-specific "skins" that customize the appearance and feel.

[0380] • Sends a stream of output packages to the modality and displays intermediate states, feedback, or results throughout the interaction phase with Assistant 1002.

[0381] According to certain embodiments, multiple instances or threads of the output processor component 1090 are simultaneously realized and / or started using one or more processors 63, and / or other combinations of hardware and / or hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the output processor component 1090 are executed, realized and / or started by one or more (or combinations thereof) of the following types of systems, components, systems, devices, procedures and processes: • A software module in a client or server of one embodiment of an intelligent automated assistant. • A service that can be called remotely. • Use a combination of procedure codes and templates.

[0382] Referring to Figure 39, a flowchart is shown illustrating an example of a multiphase output procedure according to one embodiment. The multiphase output procedure includes processing step 702 of the automatic assistant 1002 and the multiphase output step 704.

[0383] In step 710, a voice input utterance is acquired, and the voice-text components (such as those described in relation to Figure 22) interpret the speech to generate a set of candidate speech interpretations 712. In one embodiment, the voice-text components are implemented using a Nuance Recognizer, for example, a Nuance Recognizer commercially available from Nuance Communications, Inc. in Burlington, Massachusetts, MA. The candidate speech interpretations 712 are presented to the user in 730, for example, in a paraphrased form. For example, the interface may show an alternative interpretation, "did you say?", listing several possible alternative text interpretations of the same speech sample.

[0384] In at least one embodiment, a user interface is provided to allow the user to interrupt and select from candidate speech interpretations.

[0385] In step 714, the candidate speech interpretations 712 are sent to the language interpreter 1070, which generates a user intent expression 716 for at least one candidate speech interpretation 712. In step 732, paraphrases of the user intent expression 716 are generated and presented to the user. (See the related step 132 of step 120 in Figure 22.) In at least one embodiment, the user interface allows the user to interrupt and select from natural language interpretation paraphrases 732.

[0386] In step 718, task and dialogue analysis is performed. In step 734, the task and domain interpretation are presented to the user using an intent paraphrasing algorithm.

[0387] Referring to Figure 40, a screenshot is shown illustrating an example of output processing according to one embodiment. Screen 4001 includes an echo 4002 of the user's voice input generated by step 730. Screen 4001 further includes a paraphrase 4003 of the user's intent generated by step 734. In one embodiment, as shown in the example of Figure 40, special formatting / highlighting is used for keywords such as "events" and is used to facilitate user training for interaction with the intelligent automated assistant 1002. For example, by visually observing the formatting of the displayed text, the user can easily identify and interpret keywords such as "events," "next Wednesday," and "San Francisco" that the intelligent automated assistant recognizes.

[0388] Referring to Figure 39, once a request is dispatched to a service (720) and results are dynamically collected, intermediate results are displayed in the form of real-time progress (736). For example, a list of restaurants is returned, and their reviews are dynamically loaded as results arrive from the review service. The services include services that access information stored locally on the device and / or from any other source and / or web-enabled services.

[0389] A unified representation of the response 722 is generated and formatted for the appropriate output modality (724). After the final output format is complete, different types of paraphrases are provided in 738. At this stage, the entire set of results is parsed and compared with the initial request. A summary of the results or an answer to the question is provided.

[0390] Referring to Figure 41, another example of output processing according to one embodiment is shown. Screen 4101 shows the paraphrased text interpretation 4102 generated by step 732, the real-time progress 4103 generated by step 736, and the paraphrased summary 4104 generated by step 738. Detailed results 4105 are also included.

[0391] In one embodiment, the assistant 1002 can generate output in multiple modes. Referring to Figure 42, a flowchart is shown illustrating an example of multimodal output processing according to one embodiment.

[0392] The process is initiated (600). The output processor 1090 takes in the unified representation of the response 490 and formats the response according to the appropriate and applicable modality and device (612). Step 612 includes information from the device and modality model 610 and / or the domain data model 614.

[0393] When response 490 is formatted (612), any combination of several different output mechanisms is used. Examples shown in Figure 42 include the following: • Generate a text message output (620), which is sent to the text message channel (630). • Generate an email output (622), which is sent out as an email message (632). • Generate GUI output (624), which is sent to a device or web browser for rendering (634). • Generates an audio output (626), which is sent to the audio generation module (636).

[0394] Those skilled in the art will understand that many other output mechanisms are available.

[0395] In one embodiment, the content of the output message generated by the polyphasic output procedure 700 is adapted to the mode of the multimodal output processing 600. For example, if the output modality is speech (626), the language used to paraphrase the user input 730, text interpretation 732, task and domain interpretation 734, progress 736, and / or result summary 738 is somewhat verbose or uses sentences that are easier to understand in audible form than in written form. In one embodiment, the language is adapted in a step of the polyphasic output procedure 700. In other embodiments, the polyphasic output procedure 700 generates an intermediate result, which is further refined into a specific language by the multimodal output processing 600.

[0396] Short-term personal memory component 1052 In at least one embodiment, the short-term personal memory component 1052 is operable to perform and / or realize various functions, operations, behaviors and / or other features, such as one or more (or combinations thereof) of the following: • Maintains a history of recent dialogs between the assistant's embodiments and the user, including a history of user input and its interpretation. • Maintains a history of recent user selections in the GUI, such as opened or explored items, dialed phone numbers, mapped items, and played movie trailers. The history of dialogs and user interactions is stored in the client's database, on the server in the client session state such as user-specific sessions or web browser cookies, or in RAM used by the client. • Stores a list of recent user requests. • Stores a sequence of results from recent user requests. • Stores a clickstream history of UI events, including button presses, taps, gestures, voice-activated triggers, and / or any other user input. • Stores device sensor data (location, time, position / orientation, motion, light level, and sound level, etc.) that correlates with the interaction with the assistant.

[0397] According to a particular embodiment, multiple instances or threads of the short-term personal memory component 1052 are realized and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software.

[0398] According to various embodiments, one or more different threads or instances of the short-term personal memory component 1052 are started in response to the detection of one or more conditions or events that satisfy one or more different minimum threshold criteria that trigger the start of at least one instance of the short-term personal memory component 1052. For example, the short-term personal memory component 1052 is invoked when the assistant 1002 has a user session with the embodiment, in the event of at least one input form or action by the user or a response by the system.

[0399] In at least one embodiment, a given instance of the short-term personal memory component 1052 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. For example, the short-term personal memory component 1052 accesses data from the long-term personal memory component 1054 (e.g., to obtain user identity and personal preferences) and / or accesses data from a local device regarding time and location contained in the short-term memory entries.

[0400] Referring to Figures 43A and 43B, screenshots are shown illustrating an example of using the short-term personal memory component 1052 to maintain the dialogue context while changing location according to one embodiment. In this example, the user asks about the local weather and simply says "in New York." Screen 4301 shows the initial response, including the local weather. When the user says "in New York," the assistant 1002 uses the short-term personal memory component 1052 to access the dialogue context and determines that the current domain is weather forecasts. This allows the assistant 1002 to interpret the new utterance "in New York" to mean "what is the weather forecast in New York this coming Tuesday?" Screen 4302 shows the appropriate response, including the weather forecast for New York.

[0401] In the examples in Figures 43A and 43B, the short-term memory stored not only the input word "is it going to rain the day after tomorrow?", but also the system's semantic interpretation of the input as a weather domain and the time parameter set to "the day after tomorrow".

[0402] Long-term personal memory component 1054 In at least one embodiment, the long-term personal memory component 1054 is operable to perform and / or realize various functions, operations, behaviors and / or other features, such as one or more (or a combination thereof) of the following: - Fixedly storing personal information and user data, including user preferences, identity verification, authentication certificates, accounts, and addresses. • To store information collected by the user using embodiments of Assistant 1002, such as bookmarks, favorites, and clipping equivalents. - To permanently store a list of saved business entities, including restaurants, hotels, stores, theaters, and other locations. In one embodiment, the long-term personal memory component 1054 also stores enough information to show a full list of entities, including not only names or URLs, but also telephone numbers, locations on maps, and photographs. • To permanently store saved movies, videos, music, shows, and other entertainment items. - To permanently store the user's personal calendar, to-do list, reminders and alerts, contact database, and social network list, etc. • To permanently store shopping lists and request lists for products and services, acquired coupons and discount codes, etc. - Fixedly store transaction history and receipts, including reservations, purchases, and tickets for events.

[0403] According to certain embodiments, multiple instances or threads of the long-term personal memory component 1054 are simultaneously realized and / or started using one or more processors 63, and / or hardware and / or other combinations of hardware and software. For example, in at least some embodiments, various aspects, features and / or functionalities of the long-term personal memory component 1054 are executed, realized and / or started using one or more databases and / or files that reside in (or are associated with) the client 1304 and / or server 1340 and / or reside in storage.

[0404] According to various embodiments, one or more different threads or instances of the long-term personal memory component 1054 are started in response to the detection of one or more conditions or events that satisfy various minimum threshold criteria that trigger the start of at least one instance of the long-term personal memory component 1054. Various examples of the various conditions or events that trigger the start and / or realization of one or more different threads or instances of the long-term personal memory component 1054 include, but are not limited to, one or more (or combinations thereof) of the following:

[0405] Long-term personal memory entries are acquired as a side effect when a user interacts with one embodiment of the assistant 1002. Any type of interaction with the assistant generates additions to long-term personal memory, including browsing, exploring, discovering, shopping, scheduling, purchasing, making reservations, and communicating with other users through the assistant. Long-term personal memory is accumulated as a result of the user signing up for an account or service, allowing Assistant 1002 to access accounts for other services, and using Assistant 1002's services on a client device that has access to other personal information databases such as calendars, to-do lists, and contact lists.

[0406] In at least one embodiment, a given instance of the long-term personal memory component 1054 accesses and / or utilizes information from one or more associated databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements located, for example, on client 1304 and / or server 1340. Examples of various data accessed by the long-term personal memory component 1054 include, but are not limited to, data from other personal information databases such as contact or friend lists, calendars, to-do lists, other list managers, personal accounts, and wallet managers provided by external services such as 1360.

[0407] Referring to Figures 44A to 44C, screenshots are shown illustrating an example of using the long-term personal memory component 1054 according to one embodiment. In this example, features are provided that include access to saved entities such as restaurants, movies, and businesses that can be found through an interaction session with one embodiment of the assistant 1002 (referred to as "My Stuff"). In screen 4401 of Figure 44A, the user has found a restaurant. The user taps Save to My Stuff 4402 to save information about the restaurant to the long-term personal memory component 1054.

[0408] Screen 4403 in Figure 44B shows user access to My Stuff. In one embodiment, the user can select from categories to navigate to the desired item.

[0409] Screen 4404 in Figure 44C shows the My Restaurant category, which includes items previously stored in My Stuff.

[0410] Automatic calling and response procedures Referring to Figure 33, a flowchart illustrating an automated calling and response procedure according to one embodiment is shown. The procedure in Figure 33 is implemented in connection with one or more embodiments of the intelligent automated assistant 1002. It should be understood that the intelligent automated assistant 1002 as shown in Figure 1 is merely one example of a wide range of embodiments of the intelligent automated assistant system that may be implemented. Other embodiments of the intelligent automated assistant system (not shown) may include additional components / features, fewer components / features, and / or different components / features compared to the components / features shown in the example of the intelligent automated assistant 1002 shown in Figure 1.

[0411] In at least one embodiment, the automated calling and answering procedure of Figure 33 is operable to perform and / or realize various functions, operations, actions and / or other features, such as one or more (or a combination thereof) of the following:

[0412] The automatic calling and answering procedure in Figure 33 provides an interface control flow loop for the conversation interface between the user and the intelligent automatic assistant 1002. At least one execution of the automatic calling and answering procedure constitutes one conversational exchange. The conversation interface is an interface in which the user and the assistant 1002 communicate by uttering utterances to each other in a conversational manner. The automatic calling and response procedure in Figure 33 provides an execution control flow for the intelligent automatic assistant 1002. Specifically, the procedure controls the collection of input, the processing of input, the generation of output, and the presentation of output to the user. The automatic calling and response procedure in Figure 33 coordinates communication between the components of the intelligent automatic assistant 1002. Specifically, the procedure directs where the output of one component is supplied to another component, and where all inputs from the environment and actions in the environment are performed.

[0413] In at least some embodiments, portions of the automated calling and answering procedures are implemented in other devices and / or systems of the computer network.

[0414] According to certain embodiments, multiple instances or threads of an automated call and response procedure are implemented and / or started simultaneously using one or more processors 63, and / or hardware and / or other combinations of hardware and software. In at least some embodiments, one or more parts or selected parts of an automated call and response procedure are implemented in one or more clients 1304, one or more servers 1340, and / or a combination thereof.

[0415] For example, in at least some embodiments, various aspects, features, and / or functionalities of the automated calling and response procedure are executed, implemented, and / or initiated by software components, network services, and / or databases, or combinations thereof.

[0416] According to various embodiments, one or more different threads or instances of an automated call and response procedure are started in response to the detection of one or more conditions or events that satisfy one or more different criteria (e.g., minimum threshold criteria) that trigger the start of at least one instance of the automated call and response procedure. Examples of the various conditions or events that trigger the start and / or realization of one or more different threads or instances of an automated call and response procedure include, but are not limited to, one or more (or combinations thereof) of the following:

[0417] • User sessions with instances of the intelligent automated assistant 1002, including, but not limited to, one or more of the following: For example, a mobile device application that launches a mobile device application that implements one embodiment of the intelligent automatic assistant 1002. • A computer application that launches an application that implements, for example, one embodiment of the intelligent automatic assistant 1002. • Dedicated buttons on mobile devices that are pressed, such as the "speech input button". Buttons on peripheral devices connected to a computer or mobile device, such as headsets, telephone headsets or base stations, GPS navigation systems, consumer electronics, remote controls or any other devices with buttons associated with call assistance. • A web session initiated from a web browser to a website that enables the Intelligent Automated Assistant 1002 For example, an interaction initiated from within an existing web browser session to a website that implements the Intelligent Automated Assistant 1002 service, where the Intelligent Automated Assistant 1002 service is requested. • Email message sent to modality server 1426 mediating communication with one embodiment of the intelligent automatic assistant 1002 A text message is sent to a modality server 1426 that mediates communication with one embodiment of the intelligent automatic assistant 1002. A phone call is made to the modality server 1434, which is mediating communication with one embodiment of the intelligent automatic assistant 1002. An event such as an alert or notification is sent to an application that provides one embodiment of the intelligent automatic assistant 1002. If the device providing the intelligent automatic assistant 1002 is turned ON and / or started.

[0418] According to various embodiments, one or more different threads or instances of an automated call and response procedure are started and / or implemented manually, automatically, statically, dynamically, simultaneously, and / or in combination thereof. Furthermore, various instances and / or embodiments of the automated call and response procedure are started at one or more different time intervals (e.g., during a specific time interval, at regular intervals, at irregular intervals, and as needed, etc.).

[0419] In at least one embodiment, a given instance of an automated call and response procedure utilizes and / or generates various data and / or other types of information when performing a particular task and / or operation. This includes, for example, input data / information and / or output data / information. For example, in at least one embodiment, at least one instance of an automated call and response procedure accesses, processes, and / or utilizes information from one or more different sources, such as one or more databases. In at least one embodiment, at least a portion of the database information is accessed via communication with one or more local and / or remote memory elements. Furthermore, at least one instance of an automated call and response procedure generates one or more different types of output data / information, which are stored, for example, in local memory and / or remote memory elements.

[0420] In at least one embodiment, the initial configuration of a given instance of an automated call and response procedure is performed using one or more different initialization parameters. In at least one embodiment, at least a portion of the initialization parameters are accessed via communication with one or more local and / or remote memory elements. In at least one embodiment, at least a portion of the initialization parameters provided to an instance of the automated call and response procedure may correspond to and / or be derived from input data / information.

[0421] In the specific example shown in Figure 33, it is assumed that a single user is accessing an instance of the intelligent automated assistant 1002 over a network from a client application with voice input capabilities. The user is interested in finding a suitable restaurant for dinner and is using the intelligent automated assistant 1002 in conversation to facilitate this service.

[0422] In step 100, the user is prompted to enter a request. The client's user interface provides several input modes, as described in relation to Figure 26. These input modes include, for example, the following: This is an interface for type input, and it invokes the active type input derivation procedure shown in Figure 11. This is an interface for voice input, which invokes an active voice input derivation procedure as shown in Figure 22. This is an interface for selecting input from a menu, and it invokes an active GUI input derivation as shown in Figure 23.

[0423] Those skilled in the art will understand that other input modes may be provided.

[0424] In one embodiment, step 100 includes presenting the remaining options from the previous conversation with the assistant 1002 using the techniques described in the active dialogue suggestion input derivation procedure, for example, as described in relation to Figure 24.

[0425] For example, one of the methods of active input derivation in step 100, the user says to the assistant 1002, "Where may I get some good Italian around here?" For example, the user spoke this to the voice input component. One embodiment of the active input derivation component 1094 calls a voice-text service, requests confirmation from the user, and presents the confirmed user input as a unified annotated input format 2609.

[0426] One embodiment of the language interpreter component 1070 is invoked in step 200, as described in relation to Figure 29. The language interpreter component 1070 parses the text input and generates a list of possible interpretations of the user intent 290. In one parsing, the word "italian" is associated with restaurants of style Italian, "good" is associated with recommendation property of restaurants, and "around here" is associated with a location parameter describing the distance from a global sensor reading (e.g., the user's location as given by GPS on a mobile device).

[0427] In step 300, the user intent expression 290 is passed to the dialog flow processor 1080, which implements an embodiment of the dialog and flow analysis procedure as described in relation to Figure 32. The dialog flow processor 1080 determines the most likely interpretation of the intent, maps this interpretation to instances of the domain model and parameters of the task model, and determines the next flow step of the dialog flow. In this example, the restaurant domain model is instantiated by a conditional selection task to find restaurants based on constraints (constraints on type of cuisine, recommendation level, and proximity). The dialog flow model indicates that the next step is to retrieve several examples of restaurants that satisfy those constraints and present them to the user.

[0428] In step 400, one embodiment of the flow and service orchestration procedure 400 is invoked via a service orchestration component 1082. This invokes a set of services 1084 for the user's request to find a restaurant. In one embodiment, these services 1084 provide some data in a common result. This data is merged, and the resulting list of restaurants is represented in a service-independent, unified format.

[0429] In step 500, the output processor 1092 generates a result dialog summary such as "I found some recommended Italian restaurants near here." The output processor 1092 combines this summary with the output result data and sends the combined data to a module that formats the output for the user's specific mobile device in step 600.

[0430] In step 700, this device-specific output package is sent to the mobile device, and the device's client software renders it on the mobile device's screen (or other output device).

[0431] The user views this presentation and decides to explore the various options. If the user exits (790), the method terminates. If the user does not exit (790), another iteration of the loop is started by returning to step 100.

[0432] The automated calling and response procedure is applied, for example, to a user query "how about mexican food?". Such input is derived in step 100. In step 200, the input is interpreted as "restaurants of style Mexican" and combined with other states (held in short-term personal memory 1052) to support the interpretation of the intent, except for a change in one parameter of the type of restaurant. In step 300, this updated intent improves the request, which is then given to the service orchestration component 1082 in step 400.

[0433] In step 400, the updated request is dispatched to multiple services 1084, resulting in a new set of restaurants that is summarized in a dialog in step 500, formatted for the device in step 600, and sent over the network in step 700 to display the new information on the user's mobile device.

[0434] In this scenario, the user finds their preferred restaurant, places it on a map, and sends directions to a friend.

[0435] Those skilled in the art will understand that various embodiments (not shown) of the automatic calling and answering procedure may include additional features and / or operations compared to the features and / or operations shown in the particular embodiment of Figure 33, and / or may omit at least some of the features and / or operations of the automatic calling and answering procedure shown in the particular embodiment of Figure 33.

[0436] Conditional selection In one embodiment, the intelligent automatic assistant 1002 uses conditional selection in its interaction with the user to more effectively identify and present items that the user may be interested in.

[0437] Conditional selection is a type of generic task. A generic task is an abstraction that characterizes common types of domain objects, inputs, outputs, and control flows across a class of tasks. A conditional selection task is performed by selecting an item from a set of choices of a domain object (e.g., a restaurant) based on selection constraints (e.g., a desired dish or place). In one embodiment, the assistant 1002 facilitates the user's exploration of a space of possible choices, derives the user's constraints and preferences, presents choices, and provides actions to perform on those choices, such as reserving, purchasing, remembering, or sharing. The task is completed when the user selects one or more items for which to perform an action.

[0438] Conditional selection is useful in many situations. For example, when choosing a movie to watch, a restaurant for dinner, a hotel to stay at, or a place to buy a book. Generally, conditional selection is useful when you know a category and need to select an instance of that category that possesses several desired characteristics.

[0439] One traditional approach to conditional selection is a directory service. The user selects a category, and the system provides a list of options. In a local directory, the user restricts the directory to a location, such as a city. For example, in a "yellow pages" service, the user selects a phone book for a city, looks up a category, and the phone book displays one or more items within that category. The main problem with directory services is the large number of possible related options (e.g., restaurants in a given city).

[0440] Another traditional approach is the database application. This provides a way to generate a set of choices by deriving queries from the user, searching for matching items, and presenting the items in a way that highlights their key features. The user then browses the rows and columns of the resulting set, likely sorting the results or modifying the query until they find some suitable candidates. The problem with database services is that the user must be able to make human requests function as canonical queries and use abstract machines of sorting, filtering, and browsing to explore the resulting data. These are difficult for most users to do, even with a graphical user interface.

[0441] A third conventional method is open search, such as "local search." While the search is easy to perform, there are several problems with the search service that make it difficult for users to complete conditional selection tasks. In particular, the following: Similar to directory searching, users need to narrow down the list rather than simply entering a category to see one or more possible options. • When users can limit their choices through constraints, the constraints used are not always clear (e.g., can they find places that are within walking distance or places that are open late?). • The method for specifying constraints is unclear (for example, can we request the type of cuisine or restaurant? What is the possible price?) • Multiple preferences conflict. Typically, there is no objectively "optimal" answer to a given situation (for example, wanting to find a place that is nearby, offers great service, serves cheap and delicious food, and is open late into the night). Preferences are relative and depend on what is available. For example, if a user wants to reserve a table at a high-end restaurant, they will choose that restaurant even if it is expensive. However, generally speaking, users prefer cheaper options.

[0442] In various embodiments, the Assistant 1002 of the present invention facilitates the streamlining of conditional selection tasks. In various embodiments, the Assistant 1002 employs databases and search services, as well as other functionalities, to clarify what is being sought, consider what is available, and reduce the effort required from the user in determining a satisfactory solution.

[0443] In various embodiments, the assistant 1002 makes conditional selection more understandable to humans by one of many different methods.

[0444] For example, in one embodiment, Assistant 1002 makes properties operable as constraints. The user explicitly states their requirements regarding the properties of the desired result. Assistant 1002 makes this input operable as a normal constraint. For example, instead of saying, "find one or more restaurants less than 2 miles from the center of Palo Alto whose cuisine includes Italian food," the user may simply say, "Italian restaurants in Palo Alto." Assistant 1002 makes properties requested by the user that are not parameters to the database operable. For example, if the user requests romantic restaurants, the system may make this operable as a text search or tag matching constraints. In this way, Assistant 1002 makes it easier for users to overcome some of the problems they may have with conditional selection. For users, it is easier to imagine and describe a satisfactory solution than to describe the conditions that distinguish between an inappropriate solution and a suitable solution.

[0445] In one embodiment, the assistant 1002 suggests useful selection criteria, and the user only needs to state the criteria that are important at that time. For example, the assistant 1002 might ask, "Which of these matters: price (cheaper is better), location (closer is better), rating (higher rated is better)?" The assistant 1002 may also suggest criteria that require a specific value. For example, it might suggest, "You can say what kind of cuisine you would like or a food item you would like."

[0446] In one embodiment, the assistant 1002 facilitates decision-making among different options based on many competing criteria (e.g., price, quality, availability, and convenience).

[0447] By providing such guidance, Assistant 1002 facilitates the user in performing multi-parameter determination in one of the following ways:

[0448] One approach is to reduce the spatial dimensionality and combine raw data, such as ratings from multiple sources, into a synthetic "recommendation" score. The synthetic score takes into account domain knowledge about the data sources (for example, Zagat ratings can predict quality better than Yelp).

[0449] Another approach is to focus on a subset of the criteria and change the question from "what are all the possible criteria to consider and how to combine them?" to "which is the most important criterion in a given situation?" (e.g., "which is more important, price or proximity?").

[0450] Another way to make a simple decision is to assume a default value and an order of preferences (for example, all things are equal, and higher ratings, closer proximity, and lower prices are more desirable). The system may remember the user's previous responses indicating default values ​​and preferences.

[0451] Fourthly, the system provides key characteristics of items in a set of options not shown in the original request. For example, a user requests local Italian food. The system provides a set of restaurant options and uses them to provide a list of popular tags or taglines from guidebooks used by reviewers (e.g., "a nice spot for a date," "great pasta"). This allows the user to select a specific item and complete the task. The study shows that most users make decisions by evaluating specific instances rather than deciding according to criteria and rationally accepting what appears at the top. It also shows that users learn about features from detailed examples. For example, when choosing a car, a buyer might not care about the navigation system until they see that some cars have one (and the navigation system becomes an important criterion). Assistant 1002 presents key characteristics of listed items, suggesting aspects that are aligned with the user's selection or optimization.

[0452] Conceptual Data Model In one embodiment, the assistant 1002 provides assistance with conditional selection tasks by simplifying the conceptual data model, which is an abstract concept presented to the user in the interface of the assistant 1002. To overcome the aforementioned psychological problems, in one embodiment, the assistant 1002 provides a model that allows the user to describe their requests not in terms of constraint expressions, but in terms of several easily understandable and reproducible characteristics of appropriate choices. Thus, the characteristics are made easily comprehensible in natural language requests (e.g., adjectives that change keyword markers) and prompts ("you may also favor recommended restaurants..."). In one embodiment, a data model is used that allows the assistant 1002 to determine a general method for guidance instantiated by a domain of interest (e.g., restaurants vs. hotels) and domain-specific characteristics.

[0453] In one embodiment, the conceptual data model used by Assistant 1002 includes a selection class, which is a representation of the space of things to choose. For example, in a restaurant search application, the selection class is the class of restaurants. The selection class is abstract and has subclasses such as "things to do while in a destination." In one embodiment, the conceptual data model assumes that, in a given problem-solving situation, the user is interested in choosing from a single selection class. This assumption simplifies the interaction and allows Assistant 1002 to declare authority boundaries (e.g., "I know about restaurants, hotels, and movies" as opposed to "I know about life in the city").

[0454] Assuming a selection class, in one embodiment, the data model presented to the user for a conditional selection task includes, for example, items, item characteristics, selection criteria, and constraints.

[0455] An item is an instance of the selection class.

[0456] Item features are characteristics, attributes, or calculated values ​​that are presented and / or associated with at least one item. For example, the name and phone number of a restaurant are item features. Features may be unique (restaurant name or dish) or relational (e.g., distance from the current place of interest). Features may be static (e.g., restaurant name) or dynamic (e.g., rating). Features may be composite values ​​calculated from other data (e.g., a score for "monetary value"). Item features are abstractions for the user created by the domain modeler; that is, item features do not need to correspond to underlying data from backend services.

[0457] A selection criterion is an item characteristic used to compare the relevance or value of items. In other words, a selection criterion is a way of saying that an item is preferable. Selection criteria are modeled as characteristics of the item itself, whether they are intrinsic properties or calculated. For example, proximity (defined as the distance from the place of interest) is a selection criterion. Location in space-time is a property, not a selection criterion, and is used with the place of interest to calculate the distance from it.

[0458] Selection criteria have an inherent preference order; that is, the value of any particular criterion is used to secure items in the best-priority order. For example, the proximity criterion has an inherent preference that closer is preferable. Location, on the other hand, has no inherent preference value. This constraint allows the system to make default assumptions and guide selections when the user only states the criteria. For example, a user interface might be provided to "sort by rating," assuming that higher ratings are preferable.

[0459] One or more selection criteria are characteristics of the items. These are characteristics that relate to choosing from among the available items. However, item characteristics are not necessarily related to preferences (for example, the name and phone number of a restaurant are not generally relevant to choosing from them).

[0460] In at least one embodiment, a constraint is a restriction on a desired value of a selection criterion. Formally, constraints can be expressed as a set belonging relation (e.g., cuisine type includes Italian), a pattern match (e.g., restaurant review text includes "romantic"), a fuzzy inequality (e.g., distance less than a few miles), a quality threshold (e.g., highly rated), or a more complex function (e.g., worth the price). To make things simple enough for the average person, this data model relaxes at least one constraint to symbolic values ​​that are matched as words. Time and distance are excluded from this relaxation. In one embodiment, the operators and thresholds used to implement the constraints are not disclosed to the user. For example, a constraint on a selection criterion called "cuisine" is expressed as a symbolic value such as "Italian" or "Chinese". A constraint on rating is "recommended" (a binary choice). In one embodiment, for time and distance, Assistant 1002 uses its own expressions to deal with the constraint values ​​and input ranges. For example, distance is "walking distance" and time is "tonight". In one embodiment, the assistant 1002 uses special processing to match such inputs to highly accurate data.

[0461] In at least one embodiment, some constraints are required. This simply means that the task cannot be completed without this data. For example, even if the user knows the name, it is difficult to select a restaurant without a concept of the desired location.

[0462] In summary, a domain is modeled as a selection class containing the characteristics of items important to the user. Several characteristics are used to select and order the items provided to the user. These characteristics are called selection criteria. Constraints are symbolic restrictions on the selection criteria that narrow the set of items to matching items.

[0463] Multiple criteria often conflict, and constraints frequently partially coincide. Data models reduce the problem of selection through optimization (finding the optimal solution) to a matching problem (finding items that are appropriate for a given set of criteria and match a set of symbolic constraints). Algorithms for selecting criteria and constraints and determining their ordering are described in the next section.

[0464] Methodologies for Conditional Selection In one embodiment, the assistant 1002 performs conditional selection by taking as input a list of ordered criteria with implicit or explicit constraints on at least one of them, and selecting a set of candidate items that contain the main features. Computationally, the selection task is characterized as a nested search. First, it identifies a selection class, identifies key selection criteria, specifies constraints (the range of acceptable solutions), and searches instances in the order that best fits to find an acceptable item.

[0465] Referring to Figure 45, an example of an abstract model 4500 of a conditional selection task, which is a nested search, is shown. In this example, the assistant 1002 identifies a selection class from all location search types 4501 (4505). The identified class is restaurants. Within the set of all restaurants 4502, the assistant 1002 selects a criterion (4506). In this example, the criterion is identified as distance. Within the set of restaurants in PA 4503, the assistant 1002 specifies a constraint for the search (4507). In this example, the identified constraint is "Italian cuisine". Within the set of Italian restaurants in PA 4504, the assistant selects an item to present to the user (4508).

[0466] In one embodiment, such nested exploration is performed by the assistant 1002 when it has relevant input data, rather than being a flow that derives data and presents results. In one embodiment, such a control flow is managed through a dialog between the user and the assistant 1002, which operates through other procedures such as dialogs and task flow models. Conditional selection provides a framework for constructing dialogs and task flow models at an abstract level (i.e., appropriate for conditional selection tasks regardless of domain).

[0467] Referring to Figure 46, an example of a dialog 4600 that facilitates user guidance during the search process so that relevant input data is obtained is shown.

[0468] In the example of dialog 4600, the first step is to specify the type of thing the user is looking for, which is a selection class. For example, the user does this by saying "dining in Palo Alto". This allows Assistant 1002 to infer the task and domain (4601).

[0469] Once Assistant 1002 understands the combination of the task and domain (selection class = restaurant), the next step is to understand the selection criteria important to the user, for example, by requesting criteria and / or constraints (4603). In the example above, "in palo alto" indicates the place of interest. In the context of restaurants, the system interprets the place as a proximity constraint (technically, a constraint on proximity criteria). Assistant 1002 describes what is needed and receives input. If there is enough information to constrain the set of choices to a reasonable size, Assistant 1002 rephrases the input and presents one or more restaurants that satisfy the proximity constraints and are sorted in some useful order (4605). The user selects from this list (4607) or refines the criteria and constraints (4606). Assistant 1002 infers about the constraints already expressed, suggests other criteria that may be helpful using domain-specific knowledge, and requests constraints on those criteria. For example, when Assistant 1002 recommends a restaurant within walking distance of the hotel, useful criteria to consider are the availability of the food and the tables.

[0470] The conditional selection task is completed when the user selects an instance of the selection class (4607). In one embodiment, an additional subsequent task 4602 is enabled by the assistant 1002, which then provides a service indicating the selection while providing some other value. Example 4608 is sharing the selection with other users by making a restaurant reservation, setting a reminder on a calendar, and / or sending an invitation. Making a restaurant reservation, for example, ensures that the selection is made. Other options include adding the restaurant to a calendar or sending an invitation to a friend that includes directions.

[0471] Referring to Figure 47, a flowchart illustrating a conditional selection method according to one embodiment is shown. In one embodiment, the assistant 1002 operates opportunistically and bidirectionally, allowing the user to jump into an internal loop by explicitly specifying, for example, a task, domain, criteria, and constraints once or more times in the input.

[0472] The method is initiated (4701). Input is received from the user according to one of the modes described herein (4702). Based on the input, if the task is unknown, the assistant 1002 requests clarification from the user (4705).

[0473] In step 4717, the assistant 1002 determines whether the user provides additional input. If additional input is provided, the assistant 1002 returns to step 4702. If no additional input is provided, the method terminates (4799).

[0474] If the task is known in step 4703, Assistant 1002 determines whether the task is a conditional selection (4704). If it is not a conditional selection, Assistant 1002 proceeds to the specified task flow (4706).

[0475] If the task is a conditional selection in step 4704, the assistant 1002 determines whether the selection class can be determined (4707). If it is not a conditional selection, the assistant 1002 provides options for known selection classes (4708) and returns to step 4717.

[0476] If a selection class is determined in step 4707, the assistant 1002 determines whether all requested constraints are determined (4709). If none of the requested constraints are determined, the assistant 1002 requests input of the requested information (4710) and returns to step 4717.

[0477] If all requested constraints are determined in step 4709, Assistant 1002 determines whether any result items can be found under the assumption of the constraints (4711). If no items satisfy the constraints, Assistant 1002 provides a way to relax the constraints (4712). For example, Assistant 1002 may use a filter / sort algorithm to relax the constraints from the lowest priority to the highest priority. In one embodiment, if some items satisfy the constraints, Assistant 1002 rephrases the situation (for example, outputting "I could not find Recommended Greek restaurants that deliver on Sundays in San Carlos. However, I found 3 Greek restaurants and 7 Recommend restaurants in San Carlos."). In one embodiment, if no items match any of the constraints, Assistant 1002 rephrases the situation and requests input for different constraints (for example, "Sorry, I could not find any restaurants in Anytown, Texas. You may pick a different location."). Assistant 1002 returns to step 4717.

[0478] If a result item is found in step 4711, the assistant 1002 provides a list of items (4713). In one embodiment, the assistant 1002 rephrases the currently specified criteria and constraints (for example, "Here are some recommended Italian restaurants in San Jose." (recommended=yes,cuisine=Italian,proximity=<in San Jose> ). In one embodiment, the assistant 1002 presents a sorted and paginated list of items that satisfy known constraints. If an item exhibits only a portion of the constraints, such conditions are shown as part of the item display. In one embodiment, the assistant 1002 provides the user with a way to select an item by initiating another task for that item, such as reserving, remembering, scheduling, or sharing. In one embodiment, for any given item, the assistant 1002 presents the main item features for selecting an instance of the selection class. In one embodiment, the assistant 1002 shows how the item satisfies constraints. For example, a Zagat rating of 5 satisfies the constraint Recommended=yes, and "1 mile away" satisfies the constraint "within walking distance of an address". In one embodiment, the assistant 1002 allows the user to delve into further details about an item, resulting in the display of more item features.

[0479] Assistant 1002 determines whether the user has selected an item (4714). If the user has selected an item, the task is completed. If there is an item, any subsequent tasks are executed (4715), and the method terminates (4799).

[0480] If the user does not select an item in step 4714, the assistant 1002 provides the user with a way to select other criteria and constraints (4716), and returns to step 4717. For example, assuming the currently specified criteria and constraints, the assistant 1002 provides the criterion that is most likely to constrain the set of choices to the desired size. If the user selects a constraint value, the constraint value is added to the previously determined constraint when steps 4703-4713 are repeated.

[0481] Since one or more criteria have unique preference values, selecting criteria adds information to the request. For example, by allowing the user to indicate that positive reviews are important, Assistant 1002 can sort by that criterion. Such information is taken into consideration when steps 4703-4713 are repeated.

[0482] In one embodiment, the assistant 1002 increases the priority of criteria that have already been identified, allowing the user to increase their importance. For example, if the user requests a fast, inexpensive, and particularly recommended restaurant within a block of their location, the assistant 1002 will ask the user to select more important criteria. Such information will be taken into consideration when steps 4703-4713 are repeated.

[0483] In one embodiment, the user provides additional input at any point while the method of Figure 47 is being performed. In one embodiment, the assistant 1002 periodically or continuously checks for such input and returns to step 4703 to process the input in response.

[0484] In one embodiment, when outputting an item or a list of items, the assistant 1002 indicates the features used to select and order the items when presenting them. For example, if the user requests a nearby Italian restaurant, the distance and features of such items in relation to cuisine would be indicated when presenting the items. This includes highlighting matches and listing the selection criteria involved when presenting the items.

[0485] Examples of domains Table 1 provides examples of conditional selection domains addressed by the assistant 1002 according to various embodiments. TIFF2026090448000015.tif161169Table 1

[0486] Filtering and sorting results In one embodiment, a filter / sort methodology is employed when presenting items that satisfy currently identified criteria and constraints. In one embodiment, selection constraints serve as filter and sort parameters for the underlying service. Thus, any selection criteria are used to determine the items in the list and to calculate the order in which those items are presented when paginated. The sort order for this task is analogous to the relevance rank in a search. For example, proximity is a criterion that includes the general concept of sorting by distance and symbolic constraint values ​​such as "within driving distance". The constraint "driving distance" is used to select a group of candidate items. Within that group, closer items are sorted higher in the list.

[0487] In one embodiment, selection constraints are at a separate "level" from associated filtering and sorting, and are functions of the underlying data and user input. For example, proximity is grouped into levels such as "walking distance," "taxi distance," and "driving distance." When sorting, one or more items within walking distance are treated as being the same distance apart. User input begins to act in the way that the user specifies constraints. For example, if the user enters "in palo alto," one or more items within the Palo Alto metropolitan area are exact matches and equivalent. If the user enters "near the University Avenue train station," the matches depend on the distance from the address, and the degree of match depends on the selection class (for example, being close to a restaurant is different from being close to a hotel). Discretization is applied even within constraints specified by continuous values. This is important for the sorting operation, so multiple criteria are involved in determining the best-priority order.

[0488] In one embodiment, the item list may be shorter or longer than the number of items shown on one "page" of output. These items are considered to be "matching" or "sufficiently relevant." Generally, the items on the first page are of the most interest, but there is a conceptually longer list, and pagination is simply a function of the form factor of the output medium. This means, for example, if a user is provided with a way to sort or browse items according to some criterion, it would be the entire set of items being sorted or browsed (corresponding to more than one page of items).

[0489] In one embodiment, there is a prioritization among selection criteria; that is, some criteria are more important than others in filtering and sorting. In one embodiment, criteria that are well-chosen by the user are given higher priority than others, and there is a default ordering for one or more criteria. This enables general lexicographical classification. Assume there is a meaningful deductive priority. For example, unless otherwise specified by the user, proximity is more important than not being high for a restaurant. In one embodiment, the deductive priority order is domain-based. The model allows, upon request, user-specific preferences to override domain default values.

[0490] Since the constraint values ​​represent several types of internal data, there are various methods for constraint matching, and these are constraint-specific. For example, in one embodiment, this is as follows:

[0491] A binary constraint means that one or more conditions must be met, or none of them must be met. For example, is the restaurant "Fast" true, or is it not true?

[0492] • The constraints on the belonging relationships of sets are based on characteristic values, meaning that one or more must match, or none must match. For example, cuisine=Greek means that the set of dishes for a restaurant includes Greece.

[0493] • Enumerated constraints match at the threshold. For example, an evaluation criterion has constraint values ​​that are evaluated, highly evaluated, or highly evaluated. Being constrained to being highly evaluated is equivalent to being highly evaluated.

[0494] Numerical constraints are matched by thresholds based on criteria. For example, "open late" is a criterion, and the user requests places that are open after 10 p.m. This type of constraint is not a symbolic constraint value and therefore falls slightly outside the scope of conditional selection tasks. However, in one embodiment, Assistant 1002 recognizes examples of such numerical constraints and maps them to thresholds using symbolic constraints (e.g., "restaurants in palo alto open now" -> "here are 2 restaurants in palo alto that are open late").

[0495] • Location and time are specifically processed. Constraints on proximity are locations of interest specified at a certain granularity level, and matches are determined. If the user specifies a city, city-level matching is appropriate; that is, areas with postal codes are allowed. Assistant 1002 understands locations that are other "nearby" locations of interest based on specific processing. Time is relevant as a constraint value of criteria with thresholds based on table availability or service calls such as flights during a given time period.

[0496] In one embodiment, the constraint is modeled such that there is a single threshold for selection and a small set of distinct values ​​to sort. For example, the affordability criterion is modeled as a rough binary constraint, where an affordable restaurant is within a certain price range threshold. If the data tolerates multiple distinct levels for selection, the constraint is modeled using a matching gradient. In one embodiment, two levels of match (e.g., high match and low match) are provided. However, it will be understood by those skilled in the art that in other embodiments, one or more levels of match can also be provided. For example, proximity is matched with a fuzzy boundary so that closer locations of interest have a higher match. The result of high or low match operation is a filter / sort algorithm, as described below.

[0497] For at least one criterion, if relevant, methods for matching and default thresholds are constructed. Users can only specify the constraint name, symbolic constraint value, or, in particular, a high-precision constraint representation (such as time and location) if applicable.

[0498] The ideal situation for conditional selection occurs when the user explicitly specifies constraints that result in a short list of candidates where one or more candidates satisfy the constraints. The user selects the appropriate one based on the characteristics of the items. However, often the problem is either too many or too few constraints. If there are too many constraints, there are few or no items that satisfy the constraints. If there are too few constraints, there are so many candidates that it is not practical to examine the list. In one embodiment, the general conditional selection model of the present invention can handle multiple constraints through robust matching and generally generates what to choose. The user chooses to improve the criteria and constraints or simply complete the task with a "sufficiently good" solution.

[0499] method In one embodiment, the following method is used to filter and sort the results.

[0500] 1. Assuming an ordered list of selection criteria selected by the user, determine the constraints on at least one of them. a. If the user specifies a constraint value, use it. For example, if the user says "greek food," the constraint is cuisine=Greek. If the user says "san Francisco," the constraint is In the City of San Francisco. If the user says "south of market," the constraint is In the Neighborhood of SoMa. b. Use domain-specific and criterion-specific default values. For example, if a user says "a table at some thai place()", it indicates that the availability criterion is relevant, but no constraint value is specified. Default constraint values ​​for availability are a date and time range such as "tonight" and default party size 2.

[0501] 2. Select the minimum N results based on the specified constraints. a. We aim to obtain N results with a high degree of agreement. b. If that fails, try to relax the constraint in the reverse order of priority: that is, one or more criteria that match at a high level, excluding the last criterion that matches at a low level. If there are no criteria that match at a low level for that constraint, try to find criteria that match at a low level along the line from the lowest priority to the highest priority. c. Repeat the loop, allowing for matching failures at constraints, from the lowest priority to the highest priority.

[0502] 3. After obtaining the smallest set of choices, perform lexicographical classification across one or more criteria (including user-specific criteria and other criteria) in order of priority. a. Consider the set of user-specific criteria with the highest priority, and then consider one or more remaining criteria with deductive priority. For example, if the deductive priority is (availability, cuisine, proximity, rating) and the user imposes constraints on proximity and cuisine, the sort priority would be (cuisine, proximity, availability, rating). b. When sorting the criteria using separate match levels (high, low, no match), the full criterion list is applied in the same way as when relaxing the constraints. i. When a set of choices is obtained without relaxing constraints, one or more of the choices in the set match at a high level and are therefore "linked" in sorting. The following criteria in the priority list then work to sort them. For example, if the user says cuisine=Italian, proximity=In San Francisco, and the sort priority is (cuisine, proximity, availability, rating), then one or more places in the list have equal matching values ​​for cuisine and proximity. Therefore, the list is sorted by availability (places with available tables appear at the top). Of the available places, the highest rated place is at the top. ii. If the set of choices is obtained by relaxing constraints, one or more items that are exact matches will be at the top of the list, followed by items that are partially matches. Within the matching group, those items are sorted by the remaining criteria, and the same is done for the partially matching group. For example, if there are only two Italian restaurants in San Francisco, the restaurants that are available will be shown first, followed by the restaurants that are unavailable. Then the remaining restaurants in San Francisco will be shown, sorted by availability and rating.

[0503] priority order The techniques described herein enable high robustness when partially specified constraints and incomplete data are presented to the assistant 1002. In one embodiment, the assistant 1002 uses these techniques to generate a user list of items in best priority, i.e., according to relevance.

[0504] In one embodiment, such relevance sorting is based on a deductive priority order; that is, a set of criteria is selected from those that are important with respect to the domain and arranged in order of importance. If one or more are equivalent, a criterion with a higher priority is more strongly associated with conditional selection between items than a criterion with a lower priority. Assistant 1002 may operate with one or more criteria. Furthermore, the criteria may be changed over time without interrupting existing behavior.

[0505] In one embodiment, the order of priority between criteria is adjusted by domain-specific parameters, as the way criteria interact depends on the selection class. For example, when selecting a hotel, availability and price are the primary constraints, but when selecting a restaurant, cuisine and proximity are more important.

[0506] In one embodiment, the user disables the default ordering criteria in a dialog. This allows the system to use ordering to determine which constraints should be relaxed, guiding the user if the search is too constrained. For example, if the user provides constraints on dishes, proximity, recommendations, and food items, and there are no perfectly matching items, the user can say that food items are more important than recommendations and can change the combinations so that matching desired food items are sorted at the top.

[0507] In one embodiment, when priority order is determined, user-specific constraints take precedence over others. For example, in one embodiment, proximity is a requested constraint, is always specified, and takes precedence over other unselected constraints. Therefore, proximity does not need to be the highest priority constraint to be entirely primary. Also, since many criteria will not match one or more unless constraints are given by the user, the priority of these criteria is only important within the criteria selected by the user. For example, if the user specifies a dish, the dish is important; if not, it is irrelevant to sorting the items.

[0508] For example, the following is a candidate priority sorting paradigm for the restaurant domain.

[0509] 1. Dishes* (Cannot be sorted unless constraints are given) 2. Availability* (Sortable using a default constraint value, such as time) 3. Recommended 4. Proximity * (Constraint values ​​are always given) 5. Affordability 6. Delivery available 7. Food items (for example, they cannot be sorted unless a constraint value such as a keyword is given) 8. Keywords (For example, sorting is not possible unless a constraint value that is a keyword is given) 9. Restaurant Name

[0510] The following is an example of a design principle for the sort paradigm described above. • If the user specifies a dish, we want to stick with that dish. If one or more items are equivalent, sort them by their evaluation level (the highest priority among the criteria used for sorting without constraints). In at least one embodiment, proximity is more important than most. However, if proximity matches at a separate level (e.g., within a city and within walking distance) and is always specified, most items that match best in time will be "bound" by proximity. • Availability (determined by searching websites such as open-table.com, for example) is a useful sorting criterion and is based on the default value for sorting if not specified. If the user indicates a reservation time, only available slots will be listed, and sorting will be based on recommendations.

[0511] • If a user specifically requests to know recommended locations, the search results are sorted by proximity and availability, and these criteria are relaxed before recommendations are made. Assume that a user is looking for a satisfactory location and is willing to drive a little further, which is more important than the default table availability. If a specific time for availability is specified and the user requests recommended locations, available and recommended locations appear first, and recommendations are relaxed to lower-level matches until one or more availability matches. The remaining constraints, aside from the name constraint, are one or more constraints based on incomplete data or matches. Therefore, by default, these constraints are weak sort heuristics, and when a constraint is specified, there are either one or more matches, or there are no matches. • The name is used as a constraint to handle an example where a user says a restaurant name, for example, finding one or more Hobee's restaurants near Palo Alto. In this case, one or more items match that name and are sorted by proximity (another specified constraint in this example).

[0512] Domain modeling: Mapping selection criteria to underlying data It is desirable to distinguish between the data available for calculations by Assistant 1002 and the data used to make selections. In one embodiment, Assistant 1002 uses a data model that reduces complexity for the user by combining one or more types of data used to distinguish items into a single selection criteria model. Internally, these data take several forms. Instances of a selection class have unique characteristics and attributes (e.g., restaurant dishes), are compared along aspects (e.g., distance from a certain location), and are discovered by certain queries (e.g., matching a text pattern or being available at a given time). These data are calculated from other data that are not revealed to the user as selection criteria (e.g., a weighted combination of evaluations from multiple sources). These data are one or more pieces of data related to the task, but the distinction between these three types of data is irrelevant to the user. Since the user considers the characteristics of the desired choice rather than characteristics and aspects, Assistant 1002 makes these various criteria into item characteristics and makes them operational. Assistant 1002 provides a domain data model presented to the user and maps it to data that can be found in the web service.

[0513] A type of mapping is an isomorphism from underlying data to criteria presented to the user. For example, the availability of tables for reservation from the user's perspective is precisely what online reservation websites like opentable.com provide, using the same granularity of time and party size.

[0514] Another type of mapping involves normalizing data from one or more services into a common set of values ​​by unifying potentially equivalent values. For example, dishes from one or more restaurants are represented as a single ontology in Assistant 1002, which maps to various glossaries used in the various services. This ontology is hierarchical and has leaf nodes that point to specific values ​​from at least one service. For example, one service has a dish value for "Chinese," another value for "Szechuan," and a third value for "Asian." The ontology used by Assistant 1002 semantically matches references to "Chinese food" or "Szechuan" with one or more of those nodes. The confidence level reflects the degree of match.

[0515] Normalization is relevant when resolving differences in accuracy. For example, the location of a restaurant may be given at the road level in one service, but at the city level in another. In one embodiment, the assistant 1002 uses a deep structured representation of place and time that maps to different surface data values.

[0516] In one embodiment, the assistant 1002 uses a special type of mapping for open-ended modifiers (e.g., romantic, quiet) that are mapped to full-text search matches, tags, or other open texture features. In this case, the name of the selection constraint is a modifier such as "is described as".

[0517] In at least one embodiment, constraints are mapped to an order of behavioral preferences. That is, given the names of the selection constraints and their constraint values, the assistant 1002 can interpret the criteria as an ordering across possible items. There are several technical issues to address in such mapping, such as the following:

[0518] • The order of preferences conflicts. The order given by one constraint may contradict or be inversely correlated with the order given by another constraint. For example, price and quality tend to conflict. In one embodiment, the assistant 1002 interprets the constraints selected by the user in a weighted or combined order that reflects the user's desires but is faithful to the data. For example, the user may ask for "cheap fast food French restaurants within walking distance rated highly." In many places, such restaurants may not exist. However, in one embodiment, the assistant 1002 presents a list of items to optimize for at least one constraint and explains why at least one is listed. For example, one item might be "highly rated French cuisine" and another item might be "cheap fast food within walking distance."

[0519] • Data can be used as hard or soft constraints. For example, the price range of a restaurant is important when choosing a restaurant, but it is difficult to specify a price threshold beforehand. Constraints that appear to be hard constraints, such as cuisine, may actually be soft constraints due to partial matches. In one embodiment, Assistant 1002 uses a data modeling strategy that attempts to smooth one or more criteria into symbolic values ​​("cheap" or "close," etc.), so that these constraints are mapped to a function that makes the criteria and order appropriate without being strict with respect to thresholds for each match. For symbolic criteria that include explicit objective truth values, Assistant 1002 gives the objective criteria a greater weight than other criteria and clarifies in the explanation that it is known that some items do not exactly match the requested criteria.

[0520] • Items must match several constraints, not just one or more, and the "best-fitting" item will be displayed. Generally, the assistant 1002 determines the characteristics of an item that are the main characteristics of the item for the domain, and which serve as selection criteria and possible constraints for at least one criterion. Such information is provided, for example, via behavioral data and API calls.

[0521] Paraphrasing and prompt text As described above, in one embodiment, the assistant 1002 acts toward the user's objective by providing and indicating feedback, understanding the user's intent, and generating paraphrases of the current understanding. In the conversational dialog model of the present invention, paraphrases are output by the assistant 1002 after user input as a preface (e.g., paraphrase 4003 in Figure 40) or as a summary of subsequent results (e.g., list 3502 in Figure 35). Prompts are suggestions to the user regarding what can be done to improve the request or to explore the selection space along several aspects.

[0522] In one embodiment, the purpose of paraphrasing and prompt text includes, for example, the following: • Demonstrate that Assistant 1002 understands not only text but also the concept of user input. • Demonstrate the scope of Assistant 1002's understanding. • Guide the user to enter the text required for the assumed task. • To make it easier for users to explore the space of possible options in conditional selection. • Describe the current results obtained from the service in relation to the user's specified criteria and the assumptions of Assistant 1002 (e.g., describe the results of requests with insufficient constraints and requests with excessive constraints).

[0523] For example, the following paraphrases and prompts illustrate some of their purposes.

[0524] User input: Indonesian food in Menlo Park System interpretation: Task=constrained selection SelectionClass=restaurant Constraints: Location=Menlo Park, CA Cuisine = Indonesian (known in the context of food culture) Results from the service: None of them matched at a high level. In other words: Sorry, I can't find any Indonesian Restaurant Spot Menlo Park .(sorry. Menlo Park No Indonesian restaurants were found nearby. Prompt: You could try other cuisines or location (You can try other dishes and places.) Prompt below the hypertext link: Indonesian :You can try other food categories such as Chinese,or a favorite food item such as steak.( Indonesian cuisine (For example, you can try dishes from other categories, such as Chinese food, or your favorite dishes, such as steak.) Menlo Park :Enter a location such as a city,neighborhood,street address,or "near" followed by a landmark.( Menlo Park Please enter a city, neighborhood, address, or a nearby landmark. Cuisine:Enter a food category such as Chinese or Pizza.( cooking Please enter the type of cuisine, such as Chinese food or pizza. Locations :Enter a location:a city,zip code,or "near" followed by the name of a place.( place Please enter your location: city, zip code, or "near" the location name.

[0525] In one embodiment, the assistant 1002 responds relatively quickly to user input by paraphrasing. The paraphrasing is updated after the result is known. For example, the initial response might be, "Looking for Indonesian Restaurant Spot Menlo Park ...(I'm looking for an Indonesian restaurant near Menlo Park)." Once results are retrieved, Assistant 1002 updates the text, "Sorry, I can't find any Indonesian Restaurant Spot Menlo Park You could try other cuisines or location (sorry. Menlo Park No Indonesian restaurants were found near [location]. cooking , or, place You can try it out.) This displays ". Certain items are highlighted (in this case, underlined) to indicate that they represent constraints that can be relaxed or modified.

[0526] In one embodiment, a special format / emphasis is used for paraphrasing keywords. This is useful for facilitating user training for interaction with the intelligent automated assistant 1002 by showing the user the words that are most important to the assistant 1002 and are more likely to be understood by the assistant 1002. The user is more likely to use such words in the future.

[0527] In one embodiment, paraphrasing and prompts are generated using any relevant contextual data. For example, any of the following data items may be used individually or in combination:

[0528] • Parsing – A tree of ontology nodes associated with matching input tokens, along with annotations and exceptions. This includes, for each node in the parsing, all tokens of the input that provide metadata for the node and / or the basis for the node's value. • Task if known. • Selectable class. • Location constraints that do not depend on the selected class. • Requested parameters that are unknown to the given selection class (for example, location is a constraint required for a restaurant). • The name of the designated entity in parsing, which is an instance of the Select class if a name exists (for example, a specific restaurant or movie title). Is this a subsequent improvement or the start of a new conversation? (Reset starts a new conversation). What parsing constraints are associated with the input value after the constraint values ​​have been changed? In other words, what constraints have been changed by the latest input? • Is the choice class inferred or explicitly specified? • Was it sorted by quality, relevance, or proximity? • To what extent did it match each specified constraint? • Is the improvement entered as text or clicked?

[0529] In one embodiment, the paraphrasing algorithm considers the query, the domain model 1056, and the service results. The domain model 1056 includes features and classes containing metadata used to determine how to generate text. Examples of such metadata for paraphrasing generation include:

[0530] ·IsConstraint={true|false} ·IsMultiValued={true|false} ·ConstraintType={EntityName,Location,Time,CategoryConstraint,AvailabilityConstraint,BinaryConstraint,SearchQualifier,GuessedQualifer} ·DisplayName=String · DisplayTemplateSingular=String ·DisplayTemplatePlural=String ·GrammaticalRole={AdjectiveBeforeNoun,Noun,ThatClauseModifier}

[0531] For example, parsing includes these elements. Class: Restaurant IsConstraint=false DisplayTemplateSingular="restaurant" DisplayTemplatePlural="restaurants" GrammaticalRole=Noun Feature: RestaurantName (e.g. "Il Fornaio") IsConstraint=true IsPrintValued=false ConstraintType=EntityName DisplayTemplateSingular="named$1" DisplayTemplatePlural="named$1" GrammaticalRole=Noun Characteristics: Restaurant Cuisine (e.g., "Chinese") IsConstraint=true IsPrintValued=false ConstraintType=CategoryConstraint GrammaticalRole=AdjectiveBeforeNoun Feature: RestaurantSubtype (e.g., "cafe") IsConstraint=true IsMultiValued=false ConstraintType=CategoryConstraint DisplayTemplateSingular="$1" DisplayTemplatePlural="$1s" GrammaticalRole=Noun Feature: RestaurantQualifiers (e.g., "romantic") IsConstraint=true IsMultiValued=true ConstraintType=SearchQualifier DisplayTemplateSingular="is described as $1" DisplayTemplatePlural="are described as$1" DisplayTemplateCompact="matching $1" GrammaticalRole=Noun Feature: FoodType (e.g., "burritos") IsConstraint=true IsMultiValued=false ConstraintType=SearchQualifier DisplayTemplateSingular="serves $1" DisplayTemplatePlural="serve $1" DisplayTemplateCompact="serving $1" GrammaticalRole=ThatClauseModifer Feature: IsRecommended (e.g., true) IsConstraint=true IsMultiValued=false ConstraintType=BinaryConstraint DisplayTemplateSingular="recommended" DisplayTemplatePlural=" recommended" GrammaticalRole=AdjectiveBeforeNoun Feature: RestaurantGuessedQualifiers (e.g., “spectacular”) IsConstraint=true IsMultiValued=false ConstraintType=GuessedQualifer DisplayTemplateSingular="matches $1 in reviews" DisplayTemplatePlural="matches $1 in reviews" DisplayTemplateCompact="matching $1" GrammaticalRole=ThatClauseModifer

[0532] In one embodiment, the assistant 1002 can handle non-matching inputs. To handle such inputs, the domain model 1056 can provide nodes of the type of GuessedQualifier for each selection class, and words that do not match if they are in a matching rule or correct grammatical context. That is, GuessedQualifiers are handled in parsing as a wide variety of nodes that indicate a match and are likely to be a modifier of a selection class in the correct context when there are words that cannot be found in the ontology. The difference between GuessedQualifiers and SearchQualifiers is that the latter are matched against the glossary of terms in the ontology. This distinction means that the assistant 1002 may be more reliably able to identify intent with SearchQualifiers and more hesitant to echo back GuessedQualifiers.

[0533] In one embodiment, the assistant 1002 performs the following steps when generating paraphrased text.

[0534] 1. If the task is unknown, Assistant 1002 explains what it can do and requests further input.

[0535] 2. If the task is a conditional selection task and the location is known, the assistant 1002 describes the domain it recognizes and requests input for the selection class.

[0536] 3. If the selection class is known but the required constraints are missing, request input for those constraints. (For example, requesting location for a conditional selection about restaurants.)

[0537] 4. If the input contains the EntityName of the selected class, output "looking up" <name> in <location>.

[0538] 5. If this is the first question in the conversation, output a complex noun clause explaining the constraint after "looking for".

[0539] 6. If this is a subsequent improvement step in the dialog a. If the user completes the requested input, output "thanks" and rephrase it as successful. (This is done if there are requested constraints that map to the user input.) b. If the user has changed the constraints, verify this and rephrase it correctly. c. If the user types the correct name for an instance of the selected class, handle this specially. d. If the user adds a phrase that is not understood, show how it will be combined as part of the search. If appropriate, the input will be dispatched to the search service. e. If the user has added a valid constraint, output "OK" and rephrase it as successful.

[0540] 7. Use the same method of paraphrasing to explain the results. However, if the results are surprising or unexpected, use your knowledge of the data and services to explain them. Also, if the query is too restrictive or under-restricted, request further input.

[0541] Grammar for constructing complex noun clauses In one embodiment, when paraphrasing a conditional selection task query (734), the basis is a complex noun clause around a selection class that refers to the current constraint. Each constraint has a grammatical position based on its type. For example, in one embodiment, the assistant 1002 constructs the following paraphrase:

[0542] Recommended romantic Italian Restaurant Spot Menlo Park with open tables for 2 that serve osso buco and are described as " quiet "

[0543] The grammar for constructing this is as follows: <paraphrasenounclause> :== <binaryconstraint> <searchqualifier> <categoryconstraint> <itemnoun> <locationconstraint> <availabiltyconstraint> <adjectivalclauses> <binaryconstraint>:==A single adjective indicating the presence or absence of a binary constraint (e.g., recommended (best), affordable (cheap)) You can display two or more items in a list using the same query.

[0544] <searchqualifier>:==One or more words that match the ontology for the modifier of the selection class passed to the search engine service (e.g., romantic restaurants, funny movies) Used when ConstraintType=SearchQualifier.

[0545] <categoryconstraint>:==Adjectives that identify the genre, cuisine, or category of the selected class (e.g., Chinese restaurant or R-rated file) This is the most specific type of adjective, and therefore the last prefix. It is used for characteristics of the Category Constraint and Grammarical Role = Adjective Before Noun types.

[0546] <itemnoun> :== <namedentityphrase> | <selectionclass> | <selectionclasssubtype> Find the most appropriate way to display that noun. (NamedEntity) <SubType<Class <selectionclass>:== A noun that is a common name for the selected class (e.g., restaurant, movie, place) <selectionclasssubtype>:==A noun clause that is a subtype of the select class when known (e.g., diner, museum, store, bar for the select class local business). Use for features with ConstraintType=CategoryConstraint and GrammaticRole=AdjectiveBeforeNoun.

[0547] <namedentityphrase> :== <entityname>|"the"( <selectionclass> | <selectionclasssubtype> ) <entityname>:== Appropriate names for instances of the selected class (e.g., "Il Fornaio", "Animal House", "Harry's Bar") <locationconstraint> :== <locationpreposition> <locationname> <locationpreposition>:=="in", "near", and "at", etc. <locationname>:==Data related to GPS, such as city, address, landmark, or "your current location" <availabilityconstraint>Availability constraints expressed as prepositional phrases following a noun (e.g., "with open tables," "with seats available," "available online"). These come immediately after the noun for emphasis.

[0548] <adjectivalclauses> :== <modiferverbphrase>|"that" <modiferverbphrase>"and" <modiferverbphrase> <modiferverbphrase>:= A verb phrase representing a constraint of the search keyword type on a selection class (e.g., a restaurant that is "are described as quiet", "serve meat after 11", and "match 'tragically hip' in reviews", a movie that "contains violence", and "star Billy Bob Thornton"). If there are two or more, use the "that...and" variant to include all constraints in the phrase GrammaticalRole=ThatClauseModifer. Use DisplayTemplatePlural to generate the "that" clause and place GuessedQualifier at the end. If only one such constraint exists, use the DisplayTemplateCompact variant.

[0549] Table 2 provides some examples of paraphrases provided in response to a first input to a task according to one embodiment. TIFF2026090448000016.tif230158TIFF2026090448000017.tif238158TIFF2026090448000018.tif239158TIFF2026090448000019.tif246159TIFF2026090448000020.tif141157 Table 2: Paraphrasing in response to the first input

[0550] Improving queries related to places to eat Table 3 provides several examples of paraphrasing that respond to a situation where the user's intention to find a place to eat is understood, but a specific place to eat is not selected. These present a list of restaurants and can be improved. TIFF2026090448000021.tif198169 Table 3: Paraphrasing in response to improvements

[0551] Table 4 provides some examples of the results summaries that are provided once the results are obtained. TIFF2026090448000022.tif212169TIFF2026090448000023.tif225169TIFF2026090448000024.tif211169TIFF2026090448000025.tif213169TIFF2026090448000026.tif242169TIFF2026090448000027.tif204169TIFF2026090448000028.tif138169Table 4: Summary of Results

[0552] Table 5 provides some examples of prompts that are provided when a user clicks a valid link.

[0553] Prompt when a user clicks an active link TIFF2026090448000029.tif240169TIFF2026090448000030.tif247169 Table 5: Prompts when a user clicks a valid link

[0554] Suggestions for possible responses in a dialogue In one embodiment, the assistant 1002 provides contextual suggestions. This is a proposed method by which the assistant 1002 provides user options to move forward from the current state of the dialog. The set of suggestions provided by the assistant 1002 is context-dependent, and the number of suggestions provided is context-dependent, and the number of suggestions provided is medium and form factor-dependent. For example, in one embodiment, the most representative suggestions are provided in a single column in the dialog, and an extended list of suggestions ("more") is provided in a scrollable menu, and further suggestions can be obtained by typing a few characters and selecting from auto-completion options. It will be understood by those skilled in the art that other mechanisms may be used to provide suggestions.

[0555] Various proposals are provided in various embodiments. Examples of the types of proposals include the following:

[0556] Options to improve queries, including adding, removing, or modifying constraint values. Options to repair or recover from a bad situation, such as "not what I mean," "start over," or "search the web." • An option to remove ambiguity. • Interpretation of sounds. • Correction of spelling and interpretation of texts containing semantic ambiguity. Context-specific commands such as "show these on a map," "send directions to my date," or "explain these results." • Providing suggested reciprocal sales, such as the next steps in an example of meal or event planning. • An option to reuse previous commands or parts thereof.

[0557] In various embodiments, the context for determining the most relevant proposal can be derived, for example, from the following: Dialogue state • User states including, for example, the following • Static characteristics (name, address, etc.) • Dynamic characteristics (location, time, network speed) • Dialogue history including, for example, the following Query history • History of results • Auto-completion of previously entered text

[0558] In various embodiments, the proposal is generated by any mechanism, such as the following: Rephrase domains, tasks, or constraints based on the ontology model. • Auto-complete input prompts based on the current domain and constraints. • To rephrase an ambiguous alternative interpretation. • Alternative interpretation of audio-text • Handwrite based on special dialog conditions.

[0559] According to one embodiment, a proposal is generated as an action for a command in a state of completion. The command is an explicit regular expression of the request and includes assumptions and inferences based on attempted interpretations of user input. In situations where user input is incomplete or ambiguous, a proposal is an attempt to facilitate the user adjusting the input to clarify the command.

[0560] In one embodiment, each command is an unconditionally complete statement having one of the following combinations:

[0561] • Command verbs (such as "find" or "where is") • Domain (selectable class such as "restaurants") Constraints such as location=Palo Alto and cuisine=Italian These parts of the command (verb, domain, constraint) correspond to nodes in the ontology.

[0562] The proposals are thought to be actions on commands, such as setting commands, modifying commands, or declaring whether a command is related or not. Examples include the following:

[0563] • Set the command verb or domain (e.g., "find restaurants"). • Change the verb in a command (e.g., "book it", "map it", "save it"). • Change the domain ("looking for a restaurant, not a local business"). • Clearly indicate that constraints are relevant (e.g., "try refining by cuisine"). • Select a value for the constraint (e.g., "Italian" and "French"). • Select both constraints and values ​​("near here", "tables for 2") • Clearly indicate that the constraint value is incorrect ("not that Boston"). • Clearly indicate that the constraint is not relevant ("ignore the expense"). • Clearly state the intention to change the constraint value (e.g., "try a different location"). • Change the constraint value ("Italian, not Chinese"). • Add to constraint values ​​("and with a pool, too") • Snap values ​​to the grid ("Los Angeles, not los angelos") • Start a new command and reuse the context ([after movies] "find nearby restaurants", "send directions to my friend") • Start a command that is "meta" for the context ("explain these results"). • Start a new command to reset or ignore the context ("start over", "help with speech").

[0564] The proposal may further include some of the above combinations. For example, the following: • "the movie Milk not [the food item milk provided by the restaurant]" ·「restaurants serving pizza,not just pizza joints」 ·"The place called Costco in Mountain View, I don't care whether you think it is a restaurant or local buisness" • "Chinese in mountain view" [Recent queries]

[0565] In one embodiment, the assistant 1002 includes a general mechanism for maintaining a list of suggestions ordered by relevance. The format in which the suggestions are provided varies depending on the current context, mode, and form factor of the device.

[0566] In one embodiment, the assistant 1002 determines the constraints to be modified by considering any or all of the following factors. Consider whether the constraint has a value. Consider whether the constraints were inferred or explicitly stated. • Consider key points (suggestion index)

[0567] In one embodiment, the assistant 1002 determines the output format of the proposal. Examples of output formats include the following: Change domain • If the auto-completion option is "find restaurants", it will be "try something different". • Otherwise [inferred], "not looking for restaurants" • Change name constraints • If the name is inferred, it provides another ambiguous interpretation. • Auto-complete the current result entity name. • Different names • Consider that it wasn't a name reference (remove the constraint) - provide a category instead. "not named" "Not in Berkeley" "some other day" This does not mean (to use an ambiguous alternative example). • Inferred date: "any day, I don't need a reservation"

[0568] In one embodiment, the assistant 1002 attempts to resolve ambiguity through suggestions. For example, if the current set of interpretations of the user's intent is too ambiguous (310), a suggestion is one way 322 to request further information. In one embodiment, for a conditional selection task, the assistant 1002 removes common constraints from ambiguous interpretations of the intent 290 and presents the differences in interpretation to the user. For example, if the user input includes the word "cafe" and that word matches a restaurant name or type of restaurant, the assistant 1002 asks, "did you mean restaurants named 'cafe' or 'cafe restaurants'?"

[0569] In one embodiment, the assistant 1002 infers constraints under certain circumstances. That is, for a conditional selection task, not all constraints need to be explicitly indicated in user input. Some constraints are inferred from other information available to the active ontology 1050, short-term memory 1052, and / or the assistant 1002 from other sources of information available to it. For example, this includes:

[0570] • Infer domain or location • Default assumptions such as location • Constraints that match at a low level (fuzzy logic, locations of low importance, etc.) • Ambiguous criteria (matches constraint values ​​without prefixes (name vs. category, which are often ambiguous)) When Assistant 1002 infers constraint values, it provides those assumptions as suggestions for the user to invalidate. For example, the Assistant might say to the user, "I assumed you meant around here. Would you like to look at a different location?"

[0571] The present invention has been described in particular detail with reference to possible embodiments. It will be understood by those skilled in the art that the present invention may be carried out in other embodiments. Firstly, the specific names of the components, the use of capitalization of terms, attributes, data structures, or any other programming or structural aspects are neither essential nor important, and the mechanisms that realize the present invention or its features may have various names, forms, or protocols. Furthermore, the system may be implemented by a combination of hardware and software as described above, or it may be implemented entirely by hardware elements, or entirely by software elements. Also, the division of functionality among the various system components described herein is merely illustrative and not essential. Functions performed by a single system component may be performed by multiple components, and functions performed by multiple components may be performed by a single component.

[0572] In various embodiments, the present invention is realized as a system or method for performing the above-described techniques alone or in any combination. In another embodiment, the present invention is realized as a computer program comprising a non-temporary computer-readable storage medium and computer program code encoded in the medium for causing a processor of a computing device or other electronic device to perform the above-described techniques.

[0573] In this specification, any reference to “one embodiment” means that a particular function, structure, or feature described in connection with that embodiment is included in at least one embodiment of the present invention. Where the phrase “in one embodiment” appears in various places in this specification, it does not necessarily refer to the same embodiment in all instances.

[0574] Some of the above sections present symbolic representations and algorithms of operations on data bits in the memory of computing devices. These descriptions and representations of algorithms are means used by engineers in the field of data processing to most effectively communicate the intent of their research to other engineers. An algorithm is generally considered here as a self-consistent sequence of steps (instructions) that produce a desired result. These steps require the physical manipulation of physical quantities. While not generally required, these quantities may take the form of electrical, magnetic, or optical signals that can be stored, transferred, combined, compared, and manipulated. For general use, it may be convenient to refer to these signals as bits, values, elements, symbols, characters, terms, or numbers, etc. Furthermore, without loss of generality, it may be convenient to refer to specific configurations of steps requiring the physical manipulation of physical quantities as modules or encoding devices.

[0575] Furthermore, all of those terms and similar terms are merely convenient labels applied to and associated with appropriate physical quantities. Unless otherwise indicated, as will be evident from the following explanation, any explanation using terms such as “processing,” “calculating,” “calculating,” “displaying,” or “determining” is understood to refer to the actions and processing of a computer system or similar electronic computing module and / or device that manipulates and transforms data represented as physical (electronic) quantities in computer system memory or registers, or other such information storage devices, information transmission devices, or information display devices.

[0576] Certain aspects of the present invention include the processing steps and instructions described herein in the form of algorithms. The processing steps and instructions of the present invention are embodied in software, firmware and / or hardware, and if embodied in software, are downloaded and reside in the system and run on various platforms used by various operating systems.

[0577] Furthermore, the present invention relates to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for a required purpose, or it may comprise a general-purpose computing device that is selectively invoked or reconfigured by a computer program stored in the general-purpose computing device. Such computer programs are stored in computer-readable storage media, including, but not limited to, any type of disk, such as floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, application-specific integrated circuits (ASICs), or any type of medium suitable for storing electronic instructions and coupled to a computer system bus. Furthermore, the computing devices referenced herein may comprise a single processor, or they may be architectures employing a multi-processor design to enhance computing capabilities.

[0578] The algorithms and displays presented herein are not inherently related to any particular computing device, virtualization system, or other apparatus. Various general-purpose systems may be used with the program in accordance with the art herein, or it may be convenient to construct more specialized apparatus to perform the required method steps. The required structures for these various systems will become apparent from the descriptions provided herein. Furthermore, the invention is not described with reference to any particular programming language. Various programming languages ​​may be used to implement the art of the invention as described herein, and it will be understood that any above references to specific languages ​​are provided to disclose the possibilities and optimal forms of the invention.

[0579] Accordingly, in various embodiments, the present invention may be realized as software, hardware and / or other elements that control a computer system, a computing device or other electronic device, or any combination or combination thereof. Such an electronic device may include, for example, a processor, input devices (such as a keyboard, mouse, touchpad, trackpad, joystick, trackball, microphone and / or any combination thereof), output devices (such as a screen and / or speaker), memory, long-term storage devices (such as magnetic storage devices and / or optical storage devices), and / or network connectivity, in accordance with the art well known in the art. Such an electronic device may be a portable device or a non-portable device. Examples of electronic devices used to realize the present invention include mobile phones, personal digital assistants, smartphones, kiosks, desktop computers, laptop computers, tablet computers, consumer electronic devices, consumer entertainment devices, music players, cameras, televisions, set-top boxes or electronic game units, etc. An electronic device realizing the present invention may use any operating system, such as iOS or macOS, commercially available from Apple Inc. in Cupertino, California, or any other operating system available for use on the device.

[0580] While a limited number of embodiments of the present invention have been described, it will be understood by those skilled in the art who are interested in the above description that other embodiments not departing from the scope of the invention as described herein may be devised. Furthermore, the language used herein has been selected primarily for readability and teaching purposes and not to indicate or limit the subject matter of the invention. Accordingly, the disclosure of the present invention is intended to illustrate the scope of the invention as set forth in the claims, and not to limit the scope of the invention.< / modiferverbphrase> < / modiferverbphrase> < / modiferverbphrase> < / modiferverbphrase> < / adjectivalclauses> < / availabilityconstraint> < / locationname> < / locationpreposition> < / locationname> < / locationpreposition> < / locationconstraint> < / entityname> < / selectionclasssubtype> < / selectionclass> < / entityname> < / namedentityphrase> < / selectionclasssubtype> < / selectionclass> < / selectionclasssubtype> < / selectionclass> < / namedentityphrase> < / itemnoun> < / categoryconstraint> < / searchqualifier> < / binaryconstraint> < / adjectivalclauses> < / availabiltyconstraint> < / locationconstraint> < / itemnoun> < / categoryconstraint> < / searchqualifier> < / binaryconstraint> < / paraphrasenounclause>

Claims

1. An automated assistant that runs on computing devices, An input device that receives user input, A language interpreter component that interprets the received user input in order to derive an expression of the user's intent, A dialog flow processor component that identifies at least one domain, at least one task, and at least one parameter for the task, based at least in part on the derived expression of the user's intent, A service orchestration component that invokes at least one service that performs the identified task, An output processor component that renders an output based on data received from the at least one called service and at least partially based on the current output mode, An output device that outputs the rendered output, An automated assistant characterized by having the following features.

2. The system further comprises an active input derivation component that generates at least one prompt to actively derive user input, The automated assistant according to claim 1, wherein the output device outputs the at least one generated prompt.

3. The active input derivation component generates the at least one prompt in order to actively derive input from the user via the conversation interface. The automated assistant according to claim 2, characterized in that the input device receives the user input via the conversation interface.

4. The aforementioned active input derivation component is, Get a list of ordered criteria, For at least one of the criteria obtained above, at least one constraint is obtained. By generating a set of candidate items based on the at least one criterion and the at least one constraint, The automated assistant according to claim 2, characterized in that it performs a conditional selection to generate the at least one prompt.

5. It further includes an active ontology that includes representations of concepts and relationships between concepts, The automated assistant according to claim 2, wherein the active input derivation component generates the at least one prompt using at least a subset of the representations in the active ontology.

6. The active input derivation component generates the at least one prompt in order to actively derive input from the user via a plurality of input modes. The automatic assistant according to claim 2, characterized in that the input device receives the user input via the plurality of input modes.

7. The aforementioned multiple input modes are, Typing input and, Voice input and, Input provided via a graphical user interface, The automated assistant according to claim 6, characterized by including at least one input selected from the group consisting of the following.

8. It further comprises an event detector that detects at least one event, The automated assistant according to claim 2, characterized in that the active input derivation component generates the at least one prompt in response to at least one detected event.

9. The task flow model component further includes a task flow model component that identifies at least one task flow model representing a step of performing the identified at least one task, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the task flow model component.

10. It further includes an active ontology that includes representations of concepts and relationships between concepts, The automated assistant according to claim 9, characterized in that the at least one task flow model is organized according to the active ontology.

11. The automated assistant according to claim 9, wherein the dialog flow processor component performs a conditional selection to identify the at least one task flow model.

12. The domain model further comprises a domain model component that includes at least one representation of the domain, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the domain model component.

13. Each representation of the domain is, An expression of at least one concept, A representation of at least one entity, At least one expression of a relationship, The automated assistant according to claim 12, characterized in that it includes at least one expression selected from the group consisting of the following.

14. It further includes an active ontology that includes representations of concepts and relationships between concepts, The automated assistant according to claim 12, characterized in that at least one representation of the domain is organized according to the active ontology.

15. The automated assistant according to claim 12, characterized in that the domain model component uses a modeling abstraction from a conditional selection to define at least one representation of the domain.

16. The dialog flow model further includes a representation of the steps taken in the conversation between the assistant and the user, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the dialog flow model component.

17. The service function model further includes components that include a representation of the service's functions, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the service function model component.

18. The automated assistant according to claim 17, characterized in that the service orchestration component selects at least one service to be called based on at least one service model.

19. The automated assistant according to claim 17, wherein the service function model component includes a declarative model of at least one service, and the service orchestration component dynamically selects from available services in accordance with the declarative model and further in accordance with the derived expression of user intent.

20. Further equipped with a glossary database that includes associations between words and concepts, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the glossary database.

21. It further includes a domain entity database containing data describing domain entities, The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the domain entity database and obtains data describing domain entities.

22. The system further includes a short-term memory component that stores data describing at least one user interaction with the automated assistant, The automated assistant according to claim 1, wherein at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the short-term memory component and obtains data describing at least one user interaction with the automated assistant.

23. The automated assistant according to claim 22, characterized in that the language interpreter component interprets the received user input at least partially based on information from the short-term memory component.

24. Personal information associated with the user, The information collected by the aforementioned user, A saved list of entities associated with the aforementioned user, The transaction history associated with the aforementioned user, The long-term memory component further comprises storing at least one piece of information selected from the group consisting of the following: The automated assistant according to claim 1, characterized in that at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component interfaces with the long-term memory component and retrieves stored data.

25. The automated assistant according to claim 24, characterized in that the language interpreter component interprets the received user input at least partially based on information from the long-term memory component.

26. The automated assistant according to claim 1, wherein the dialog flow processor component performs a conditional selection to identify at least one task and at least one parameter for the task.

27. The dialog flow processor component is Get a list of ordered criteria, For at least one of the criteria obtained above, at least one constraint is obtained. By identifying at least one task and at least one parameter for the task based on the at least one criterion and the at least one constraint, The automated assistant according to claim 26, characterized by performing a conditional selection.

28. The automated assistant according to claim 1, wherein the service orchestration component performs a conditional selection to invoke at least one service.

29. The aforementioned service orchestration components are: Get a list of ordered criteria, For at least one of the criteria obtained above, at least one constraint is obtained. By calling at least one service based on the at least one criterion and the at least one constraint, The automated assistant according to claim 28, characterized by performing a conditional selection.

30. The automated assistant according to claim 1, further comprising an active ontology that includes representations of concepts and relationships between concepts.

31. The automated assistant according to claim 30, characterized in that the language interpreter component interprets the received user input using at least a subset of the representations in the active ontology.

32. The automated assistant according to claim 30, wherein the dialog flow processor component identifies the at least one domain, task, and parameter using at least a subset of the representations in the active ontology.

33. The automated assistant according to claim 30, characterized in that the service orchestration component invokes at least one service using at least a subset of the representations in the active ontology.

34. The automated assistant according to claim 30, characterized in that the output processor component renders the output using at least a subset of the representations in the active ontology.

35. The automated assistant according to claim 30, characterized in that the active ontology is used as a semantic unifier for at least one database used by the automated assistant.

36. The automated assistant according to claim 30, characterized in that the active ontology is instantiated for each user session.

37. The aforementioned active ontology is At least one domain model, At least one dialogue flow model, At least one service model, At least one task model, The automated assistant according to claim 30, characterized by including at least one model selected from the group consisting of the following.

38. It also includes an event monitor that receives data related to the event, The automated assistant according to claim 1, wherein the dialog flow processor component identifies at least one task and at least one parameter for the task based at least partially on data relating to at least one received event.

39. The aforementioned input device is Equipped with at least one microphone to receive voice input, The aforementioned silver is, The automatic assistant according to claim 1, further comprising a voice-to-text component that converts the received voice input into text.

40. The automatic assistant according to claim 1, characterized in that the input device receives text input.

41. The automated assistant according to claim 1, wherein the language interpreter removes ambiguity between different parsing results for the received user input.

42. The automated assistant according to claim 1, characterized in that the service orchestration component invokes at least one service that performs an operation on the computing device.

43. The automated assistant according to claim 1, characterized in that the service orchestration component invokes at least one service that performs an operation via a computing network.

44. The automated assistant according to claim 1, characterized in that the service orchestration component calls at least one service that performs an operation via the Internet.

45. The automated assistant according to claim 1, characterized in that the service orchestration component receives results from at least one of the called services and unifies the received results.

46. The aforementioned automatic assistant is A telephone and A smartphone and Tablet computers and A laptop computer and Personal digital assistants and Desktop computers and kiosk and, Consumer electronic devices and Consumer entertainment devices and Music player and Camera and, TV and, Electronic game unit and Set-top box and The automated assistant according to claim 1, characterized in that it operates on at least one device selected from the group consisting of the following.

47. The output processor component paraphrases at least one expression of the user's intent. The automated assistant according to claim 1, characterized in that the output device outputs the paraphrased expression of the user's intent.

48. The automated assistant according to claim 1, wherein at least one of the language interpreter component, the dialog flow processor component, the service orchestration component, and the output processor component obtains parameters from the detected context of the computing device.

49. The detected context of the computing device is The current software application and its current state, Current location and Current environmental conditions detected via at least one environmental sensor, Usage history and Information describing the aforementioned user, The automated assistant according to claim 48, characterized by including at least one piece of information selected from the group consisting of the following.

50. The automated assistant according to claim 1, wherein the output device outputs the rendered output using a combination of at least two output modes.

51. The above-mentioned at least two output modes are, Text output and, Graphics output and, Sound output and, Synthesized speech output, Sampled audio output and, Actuator output and, The automatic assistant according to claim 50, characterized by including at least two outputs selected from the group consisting of the following.

52. An automated assistant that runs on computing devices, An input device that receives user input, An active ontology that includes representations of concepts and relationships between concepts, A language interpreter component that interprets the received user input to derive an expression of the user's intent, and obtains expressions of concepts and relationships between concepts from the active ontology to derive an expression of the user's intent, An output processor component that renders the output based on the derived expression of the user's intent and at least partially based on the current output mode, An output device that outputs the rendered output, An automated assistant characterized by having the following features.

53. The automated assistant according to claim 52, wherein the language interpreter component parses the received user input based on representations of concepts and relationships between concepts from the active ontology in order to derive an expression of the user's intent.

54. The automated assistant according to claim 52, wherein the language interpreter component removes ambiguity between at least two possible interpretations of the user input based on the representation of the concepts and relationships between concepts from the active ontology.

55. The system further comprises an active input derivation component that generates at least one prompt to actively derive user input, The automated assistant according to claim 52, wherein the output device outputs at least one of the generated prompts.

56. The automated assistant according to claim 55, wherein the active input derivation component generates the at least one prompt based on representations of concepts and relationships between concepts obtained from the active ontology.

57. The automated assistant according to claim 55, characterized in that the active input derivation component automatically completes the user input based on representations of concepts and relationships between concepts obtained from the active ontology.

58. The system further comprises a service orchestration component that invokes at least one service that performs a task based on the derived expression of the user's intent, The automated assistant according to claim 52, characterized in that the service orchestration component identifies a service to be invoked based on the representation of the concepts and relationships between concepts from the active ontology.

59. An automated assistant that runs on computing devices, An input device that receives user input, A language interpreter component that interprets the received user input in order to derive an expression of the user's intent, An active input derivation component that generates at least one prompt to actively derive user input via a conversational interface and generates a paraphrase of the user's intent, An output processor component that summarizes a plurality of results based on the derived expression of the user's intent and generates an output representing the summarized plurality of results based at least partially on the current output mode, An output device that outputs the paraphrased version of the user's intent and outputs the generated output, An automated assistant characterized by having the following features.

60. The input device includes an audio input device that receives audio input, The automatic assistant according to claim 59, wherein the output device further outputs text representing the voice input.

61. A method for realizing an automated assistant in a computing device having at least one processor, Receiving user input in an input device, In the processor, the received user input is interpreted in order to derive an expression of the user's intent, The processor identifies at least one domain, at least one task, and at least one parameter for the task, based at least in part on the derived expression of the user's intent. The processor invokes at least one service that performs the identified task, The processor renders the output based on data received from the at least one called service and at least partially based on the current output mode. Outputting the rendered output in the output device, A method characterized by comprising:

62. The processor generates at least one prompt in order to actively derive user input, The output device outputs the generated at least one prompt, The method according to claim 61, further comprising the following:

63. Receiving user input means Requesting input from the user via a conversational interface, Receiving user input via the aforementioned conversation interface, The method according to claim 61, characterized by including the following:

64. To rephrase at least one expression of the user's intent, The output device outputs the paraphrased expression of the user's intent, The method according to claim 61, further comprising the following:

65. The process involves generating multiple different syntactic analysis results for the received user input, To eliminate ambiguity between the aforementioned other syntactic analysis results, The method according to claim 61, further comprising the following:

66. Removing ambiguity between the aforementioned other parsing results is, Identifying at least two conflicting interpretations of the user input, The output device outputs a prompt requesting additional information from the user, The input device receives additional information from the user, To resolve the ambiguity based on the aforementioned additional information, The method according to claim 65, characterized by including the following:

67. Processing the received user input means Data from short-term memory describing at least one subsequent dialogue in the current session, Data from long-term memory describing at least one characteristic of the user, The method according to claim 61, characterized in that it includes processing the user input using at least one data selected from the group consisting of the following.

68. A computer program that provides an automated assistant on a computing device having at least one processor, Non-temporary computer-readable storage media and At least one processor Steps include receiving user input and The steps include interpreting the received user input in order to derive an expression of the user's intent, A step of identifying at least one domain, at least one task, and at least one parameter for the task based at least partially on the derived expression of the user's intent, The steps include calling at least one service that performs the identified task, The steps include rendering an output based on data received from the at least one called service and at least partially based on the current output mode, The steps include outputting the rendered output, Computer program code encoded in a medium that executes it, A computer program characterized by having the following features.

69. The at least one processor A step of generating at least one prompt to actively derive user input, The steps include outputting at least one of the generated prompts, The computer program according to claim 68, further comprising computer program code that causes to execute.

70. The computer program code that receives user input is configured in at least one processor The steps include: requesting input from the user via a conversational interface; The steps include receiving user input via the aforementioned conversation interface, The computer program according to claim 68, characterized in that it includes computer program code that causes to execute.

71. The at least one processor A step of rephrasing at least one expression of the user's intent, The output device includes the step of outputting the paraphrased expression of the user's intent, The computer program according to claim 68, further comprising computer program code that causes to execute.

72. The at least one processor The steps include generating multiple different syntactic analysis results for the received user input, The steps include removing ambiguity between the aforementioned other syntactic analysis results, The computer program according to claim 68, further comprising computer program code that causes to execute.

73. The computer program code that removes ambiguity between the aforementioned other parsing results is provided by the at least one processor. The steps include identifying at least two conflicting interpretations of the user input, The output device outputs a prompt requesting additional information from the user, The input device includes the step of receiving additional information from the user, A step of resolving the ambiguity based on the aforementioned additional information, The computer program according to claim 72, characterized in that it includes computer program code that causes the execution of a computer program.

74. The computer program code that processes the received user input is configured in at least one processor. Data from short-term memory describing at least one subsequent dialogue in the current session, Data from long-term memory describing at least one characteristic of the user, The computer program according to claim 68, characterized in that it includes computer program code that causes the computer program to perform the step of processing the user input using at least one data selected from the group consisting of the following.