Intelligent Automated Assistant in a Media Environment
The digital assistant system in media environments detects user inputs and adjusts output formats to provide contextually relevant assistance, addressing the challenge of interruptions during media consumption.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-12-05
- Publication Date
- 2026-04-30
AI Technical Summary
Integrating digital assistants into media environments, such as televisions and streaming devices, poses a challenge as user interactions through voice and visual outputs can interrupt media content consumption.
A digital assistant system that detects user input formats, performs tasks, and adjusts output formats to minimize interruptions by displaying contextually relevant information while allowing seamless interaction with media content.
The system provides comprehensive assistance to users while minimizing disruptions to media consumption, enhancing user experience by intelligently balancing assistance and content viewing.
Smart Images

Figure 0007854032000105 
Figure 0007854032000106 
Figure 0007854032000107
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority from U.S. Provisional Patent Application No. 62 / 215,676, filed on September 8, 2015, entitled "Intelligent Automated Assistant in a Media Environment", and U.S. Non - Provisional Patent Application No. 14 / 963,094, filed on December 8, 2015, entitled "Intelligent Automated Assistant in a Media Environment". These applications are hereby incorporated by reference in their entirety for all purposes.
[0002] This application is related to the following co - pending applications: U.S. Non - Provisional Patent Application No. 14 / 963,089, filed on December 8, 2015, entitled "Intelligent Automated Assistant for Media Search and Playback" (Attorney Docket No. 106842137900 (P27499US1)), U.S. Non - Provisional Patent Application No. 14 / 498,503, filed on September 26, 2014, entitled "Intelligent Automated Assistant for TV User Interactions" (Attorney Docket No. 106842065100 (P18133US1)), and U.S. Non - Provisional Patent Application No. 14 / 498,391, filed on September 26, 2014, entitled "Real - time Digital Assistant Knowledge Updates" (Attorney Docket No. 106842097900 (P22498US1)). These applications are hereby incorporated by reference in their entirety for all purposes. [Technical Field]
[0003] This application generally relates to intelligent automated assistants, and more particularly, to intelligent automated assistants operating in a media environment.
Background Art
[0004] Intelligent automated assistants (or digital assistants) can provide an intuitive interface between a user and an electronic device. These assistants may enable users to interact with the device or system using natural language in the form of speech and / or text. For example, a user may access the services of an electronic device by providing a virtual assistant associated with the electronic device with speech user input in natural language form. The virtual assistant can perform natural language processing on the speech user input to infer the user's intent and translate the user's intent into a task. The task can then be performed by executing one or more functions of the electronic device, and in some embodiments, the relevant output can be returned to the user in natural language form.
[0005] Integrating digital assistants into media environments (e.g., televisions, TV set-top boxes, cable boxes, gaming devices, streaming media devices, digital video recorders, etc.) can be desirable to assist users with media consumption-related tasks. For example, a digital assistant can be used to help users find desirable media content to consume. However, user interaction with a digital assistant may involve voice and visual output, which could interrupt the consumption of media content. Therefore, integrating digital assistants into media environments in a way that provides users with sufficient assistance while minimizing interruptions to media content consumption can be a challenge. [Overview of the Initiative]
[0006] Systems and processes for operating a digital assistant within a media environment are disclosed. In some exemplary processes, user input can be detected while displaying content. The process can determine whether the user input corresponds to a first input format. Based on the determination that the user input corresponds to a first input format, several exemplary natural language requests can be displayed. The several exemplary natural language requests can be contextually related to the displayed content.
[0007] In some embodiments, following a determination that user input does not correspond to a first input format, the process may determine whether user input corresponds to a second input format. Following the determination that user input corresponds to a second input format, audio data may be sampled. The process may determine whether the audio data contains a user request. Following the determination that the audio data contains a user request, a task may be performed that at least partially satisfies the user request. In some embodiments, the task may include obtaining a result that at least partially satisfies the user request and displaying a second user interface having a portion of the result. The portion of the content may remain displayed while the second user interface is displayed, and the display area of the second user interface may be smaller than the display area of the portion of the content.
[0008] In some embodiments, a third user input can be detected while a second user interface is displayed. Upon detection of the third user input, the display of the second user interface can be replaced with a display of the third user interface having a portion of the result. The third user interface may occupy at least a majority of the display area of the display unit. In addition, a second result can be obtained that at least partially satisfies the user request. The second result may differ from the first result. The third user interface may include at least a portion of the second result.
[0009] In some embodiments, a fourth user input can be detected while a third user interface is displayed. The fourth user input may indicate a direction. In response to the detection of the fourth user input, the focus of the third user interface can be switched from a first item within the third user interface to a second item within the third user interface. The second item may be positioned in the indicated direction relative to the first item.
[0010] In some embodiments, a fifth user input can be detected while a third user interface is displayed. Upon detection of the fifth user input, a search field can be displayed. In addition, a virtual keyboard interface can be displayed, and input received via the virtual keyboard interface can result in text entry into the search field. Furthermore, in some embodiments, selectable affordances can appear on the display of the second electronic device, and the selection of these affordances allows text input to be received by the electronic device via the keyboard of the second electronic device.
[0011] In some embodiments, a sixth user input can be detected while a third user interface is displayed. Depending on the detection of the sixth user input, second audio data encompassing a second user request can be sampled. The process can determine whether the second user request is a request to refine the results of the user request. In accordance with the determination that the second user request is a request to refine the results of the user request, a subset of the results can be displayed via the third user interface. In accordance with the determination that the second user request is not a request to refine the results of the user request, a third result that at least partially satisfies the second user request can be obtained. A portion of the third result can be displayed via the third user interface.
[0012] In some embodiments, sampled audio data may include user utterances, and a user intent corresponding to the user utterance can be determined. The process can determine whether the user intent includes a request to adjust the state or settings of the application. In accordance with the determination that the user intent includes a request to adjust the state or settings of the application, the state or settings of the application can be adjusted to satisfy the user intent.
[0013] In some embodiments, upon determination that the user intent does not include a request to adjust the state or settings of an application on an electronic device, the process can determine whether the user intent is one of a plurality of predetermined request types. Upon determination that the user intent is one of a plurality of predetermined request types, a text-only result that at least partially satisfies the user intent can be displayed.
[0014] In some embodiments, following a determination that the user intent is not one of a plurality of predetermined request types, the process can determine whether the displayed content includes media content. Following the determination that the displayed content includes media content, the process can further determine whether the media content can be paused. Following the determination that the media content can be paused, the media content is paused, and a result that at least partially satisfies the user intent can be displayed via a third user interface. The third user interface can occupy at least a majority of the display area of the display unit. Following a determination that the media content cannot be paused, the result can be displayed via a second user interface while the media content is displayed. The display area occupied by the second user interface can be smaller than the display area occupied by the media content. Furthermore, in some embodiments, following a determination that the displayed content does not include media content, the result can be displayed via a third user interface. [Brief explanation of the drawing]
[0015] [Figure 1] This is a block diagram showing systems and environments for implementing a digital assistant, relating to various embodiments.
[0016] [Figure 2] This is a block diagram showing media systems relating to various embodiments.
[0017] [Figure 3] This is a block diagram showing user devices relating to various embodiments.
[0018] [Figure 4A] This is a block diagram showing a digital assistant system or its server portion according to various embodiments.
[0019] [Figure 4B] Shows the functions of the digital assistant shown in FIG. 4A according to various embodiments.
[0020] [Figure 4C] Shows a part of the ontology according to various embodiments.
[0021] [Figure 5A] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5B] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5C] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5D] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5E] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5F] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5G] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5H] Shows a process for operating a digital assistant of a media system according to various embodiments. [Figure 5I] Shows a process for operating a digital assistant of a media system according to various embodiments.
[0022] In the following figure numbers, FIG. 6O is intentionally omitted to avoid any confusion between the capital letter O and the digit 〇 (zero). [Figure 6A]Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6B] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6C] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6D] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6E] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6F] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6G] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6H] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6I] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6J] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6K]Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6L] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6M] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6N] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6P] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments. [Figure 6Q] Figures 5A to 5I show screenshots displayed on the display unit by a media device at various stages of the process according to various embodiments.
[0023] [Figure 7A] This document describes the process for operating a digital assistant in a media system, relating to various embodiments. [Figure 7B] This document describes the process for operating a digital assistant in a media system, relating to various embodiments. [Figure 7C] This document describes the process for operating a digital assistant in a media system, relating to various embodiments.
[0024] In the following figure numbers, Figure 8O has been intentionally omitted to avoid any confusion between the capital letter O and the number 0 (zero). [Figure 8A]Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8B] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8C] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8D] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8E] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8F] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8G] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8H] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8I] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8J] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8K]Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8L] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8M] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8N] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8P] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8Q] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8R] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8S] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8T] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8U] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8V]Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments. [Figure 8W] Figures 7A to 7C show screenshots displayed on the display unit by a media device at various stages of the process shown in various embodiments.
[0025] [Figure 9] This document describes the process for operating a digital assistant in a media system, relating to various embodiments.
[0026] [Figure 10] This shows a functional block diagram of an electronic device configured to operate a digital assistant for a media system, relating to various embodiments.
[0027] [Figure 11] This shows a functional block diagram of an electronic device configured to operate a digital assistant for a media system, relating to various embodiments. [Modes for carrying out the invention]
[0028] The following description of embodiments refers to the attached drawings, which illustrate specific embodiments that can be implemented. It should be understood that other embodiments can be used and structural modifications can be made without departing from the scope of the various embodiments.
[0029] This application relates to a system and process for operating a digital assistant in a media environment. In one exemplary process, user input can be detected while displaying content. The process can determine whether the user input corresponds to a first input format. In accordance with the determination that the user input corresponds to a first input format, a number of exemplary natural language requests can be displayed. The number of exemplary natural language requests can be contextually related to the displayed content. Contextually related exemplary natural language requests may be desirable to conveniently inform the user of the digital assistant's functions that are most relevant to the user's current usage on the media device. This can encourage the user to use the digital assistant's services and can also improve the user's interaction experience with the digital assistant.
[0030] In some embodiments, upon determining that user input does not correspond to a first input format, the process can determine whether user input corresponds to a second input format. Upon determining that user input corresponds to a second input format, audio data can be sampled. The process can determine whether the audio data contains a user request. Upon determining that the audio data contains a user request, a task that at least partially satisfies the user request can be performed.
[0031] In some embodiments, the task performed may depend on the nature of the user request and the content displayed while the user input in the second input format is detected. If the user request is a request to adjust the state or settings of an application on an electronic device (e.g., to turn on subtitles for displayed media content), the task may include adjusting the state or settings of the application. If the user request is one of several predetermined request types associated with text-only output (e.g., a request for the current time), the task may include displaying text that satisfies the user request. If the displayed content includes media content and the user request requests to retrieve and display results, the process may determine whether the media content can be paused. If it is determined that the media content can be paused, the media content is paused and results that satisfy the user request can be displayed on an enlarged user interface (e.g., a third user interface 626, shown in Figure 6H). If it is determined that the media content cannot be paused, results that satisfy the user request can be displayed on a reduced user interface (e.g., a second user interface 618, shown in Figure 6G) while the media content continues to be displayed. The display area of the second user interface can be smaller than the display area of the media content. Furthermore, if the displayed content does not include media content, the enlarged user interface can display results that satisfy the user's request. By adjusting the output format according to the type of displayed content and the user's request, the digital assistant can intelligently balance providing comprehensive assistance while minimizing interruptions to the user's consumption of media content. This results in an improved user experience. 1. System and Environment
[0032] Figure 1 shows an exemplary system 100 for operating a digital assistant according to various embodiments. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automated digital assistant” may refer to any information processing system that interprets natural language input in the form of utterances and / or text to infer user intent and takes action based on the inferred user intent. For example, to take action based on the inferred user intent, the system may do one or more of the following: identify a task flow that includes steps and parameters designed to realize the inferred user intent; input specific requirements from the inferred user intent into the task flow; execute the task flow by calling a program, method, service, application programming interface (API), or similar; and generate an output response to the user in the form of an audible (e.g., utterance) and / or visual form.
[0033] Specifically, a digital assistant may have the ability to receive user requests, at least partially, in the form of natural language commands, requests, statements, descriptions, and / or inquiries. Typically, a user request may ask the digital assistant to either provide information or perform a task. A satisfactory response to a user request may be the provision of the requested information, the performance of the requested task, or a combination of both. For example, a user may ask the digital assistant, "What time is it in Paris?" The digital assistant may retrieve the requested information and respond, "It is 4:00 PM in Paris." The user may also request the performance of a task, for example, "Find me movies starring Reese Witherspoon." In response, the digital assistant may perform the requested search query and display relevant movie titles for the user to select. While performing the requested task, the digital assistant may interact with the user occasionally in a continuous dialogue involving multiple information exchanges over a longer period. There are many other ways to interact with a digital assistant to request information or the performance of various tasks. In addition to providing text responses and taking programmed actions, digital assistants can also provide responses in other visual or audio formats, such as words, alarms, music, images, videos, animations, etc. Furthermore, as described herein, exemplary digital assistants can control the playback of media content (for example, on a television set-top box) and display media content or other information on a display unit (for example, a television).
[0034] As shown in Figure 1, in some embodiments, the digital assistant can be implemented according to a client-server model. The digital assistant may include a client-side portion 102 (hereinafter, "DA client 102") that runs on a media device 104, and a server-side portion 106 (hereinafter, "DA server 106") that runs on a server system 108. Furthermore, in some embodiments, the client-side portion may also run on a user device 122. The DA client 102 can communicate with the DA server 106 through one or more networks 110. The DA client 102 can provide client-side functionality such as user-responsive input and output processing, as well as communication with the DA server 106. The DA server 106 can provide server-side functionality for any number of DA clients 102 residing on each device (e.g., media device 104 and user device 122).
[0035] The media device 104 can be any suitable electronic device configured to manage and control media content. For example, the media device 104 can include a television set-top box, such as a cable box device, a satellite box device, a video player device, a video streaming device, a digital video recorder, a game system, a DVD player, a Blu-ray Disc® player, a combination of such devices, or the like. As shown in Figure 1, the media device 104 can be part of a media system 128. In addition to the media device 104, the media system 128 can include a remote control unit 124 and a display unit 126. The media device 104 can display media content on the display unit 126. The display unit 126 can be any type of display, such as a television display, a monitor, a projector, or the like. In some embodiments, the media device 104 can be connected to an audio system (e.g., an audio receiver) and a speaker (not shown), which may be integrated with or separate from the display unit 126. In other embodiments, the display unit 126 and the media device 104 can be integrated together within a single device, such as a smart television with advanced processing and network connectivity. In such embodiments, the functions of the media device 104 can be run as an application on the combined device.
[0036] In some embodiments, the media device 104 can function as a media control center for multiple types and sources of media content. For example, the media device 104 can facilitate user access to live television (e.g., wireless, satellite, or cable TV). Therefore, the media device 104 may include a cable tuner, satellite tuner, or similar. In some embodiments, the media device 104 may also record TV programs for later time-shifted viewing. In other embodiments, the media device 104 may provide access to one or more streaming media services, such as on-demand TV programs, videos, and music delivered by cable, as well as TV programs, videos, and music delivered by the Internet (e.g., from various free, paid, and subscription-based streaming services). In yet another embodiment, the media device 104 may facilitate playback or display of media content from any other source, such as displaying photographs from a mobile user device, playing videos from a combined storage device, playing music from a combined music player, or the like. The media device 104 may also include various other combinations of the media control mechanisms described herein, as desired. A detailed description of the media device 104 is provided below with reference to Figure 2.
[0037] The user device 122 can be any personal electronic device, such as a mobile phone (e.g., a smartphone), a tablet computer, a portable media player, a desktop computer, a laptop computer, a PDA, a wearable electronic device (e.g., digital glasses, a wristband, a watch, a brooch, an armband, etc.), or similar. A detailed description of the user device 122 is provided below with reference to Figure 3.
[0038] In some embodiments, the user can interact with the media device 104 through a user device 122, a remote control device 124, or an interface element integrated with the media device 104 (e.g., buttons, microphones, cameras, joysticks, etc.). For example, utterances including media-related queries or commands for a digital assistant can be received by the user device 122 and / or the remote control device 124, and these utterances can be used to perform media-related tasks on the media device 104. Similarly, tactile commands for controlling media on the media device 104 can be received by the user device 122 and / or the remote control device 124 (as well as from other devices not shown). Thus, the various functions of the media device 104 can be controlled in various ways, giving the user multiple options for controlling media content from multiple devices.
[0039] Examples of communication networks (one or more) 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. Communication networks (one or more) 110 can be implemented using any well-known network protocol, including various wired or wireless protocols such as Ethernet®, Universal Serial Bus (USB), FireWire®, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth®, Wi-Fi®, Voice over Internet Protocol (VoIP), Wi-MAX®, or any other suitable communication protocol.
[0040] The DA server 106 may include a client-responsive input / output (I / O) interface 112, one or more processing modules 114, data and models 116, and an I / O interface 118 to external services. The client-responsive I / O interface 112 can facilitate client-responsive input and output processing for the DA server 106. One or more processing modules 114 can utilize the data and models 116 to process utterance input and determine user intent based on natural language input. Furthermore, one or more processing modules 114 can perform task execution based on the inferred user intent. In some embodiments, the DA server 106 may communicate with external services 120, such as telephone services, calendar services, information services, messaging services, navigation services, television program services, streaming media services, media search services, and the like, via a network(s) 110 to complete tasks or retrieve information. The I / O interface 118 to external services can facilitate such communication.
[0041] The server system 108 can be implemented on a distributed network of one or more independent data processing devices or computers. In some embodiments, the server system 108 can also utilize the services of various virtual devices and / or third-party service providers (e.g., third-party cloud service providers) to provide the basic computing resources and / or infrastructure resources of the server system 108.
[0042] The digital assistant shown in Figure 1 may include both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), but in some embodiments, the functionality of the digital assistant can be implemented as a standalone application installed on a user device or media device. In addition, the distribution of functionality between the client and server portions of the digital assistant may vary depending on the embodiment. For example, in some embodiments, the DA client running on user device 122 or media device 104 may be a thin client that provides only user-responsive input and output processing functions, delegating all other functions of the digital assistant to a backend server. 2. Media System
[0043] Figure 2 shows a block diagram of a media system 128 according to various embodiments. The media system 128 may include a display unit 126, a remote control device 124, and a media device 104 which is communicatively coupled to a speaker 268. The media device 104 can receive user input via the remote control device 124. Media content from the media device 104 can be displayed on the display unit 126.
[0044] In this example, as shown in Figure 2, the media device 104 may include a memory interface 202, one or more processors 204, and a peripheral device interface 206. Various components within the media device 104 may be coupled to one or more communication buses or signal lines. The media device 104 may further include various subsystems and peripheral devices coupled to the peripheral device interface 206. These subsystems and peripheral devices can collect information and / or facilitate various functionalities of the media device 104.
[0045] For example, the media device 104 may include a communication subsystem 224. Communication functions may be facilitated through one or more wired and / or wireless communication subsystems 224, which may include various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters.
[0046] In some embodiments, the media device 104 may further include an I / O subsystem 240 coupled to the peripheral interface 206. The I / O subsystem 240 may include an audio / video output controller 270. The audio / video output controller 270 may be coupled to the display unit 126 and the speaker 268, or may provide audio and video output in another way (e.g., via an audio / video port, wireless transmission, etc.). The I / O subsystem 240 may further include a remote controller 242. The remote controller 242 may be communicably coupled to the remote control device 124 (e.g., via a wired connection, Bluetooth®, Wi-Fi®, etc.).
[0047] The remote control unit 124 may include a microphone 272 for capturing audio data (e.g., user utterances), one or more buttons 274 for capturing tactile input, and a transceiver 276 for facilitating communication with the media device 104 via the remote controller 242. Furthermore, the remote control unit 124 may include a touch-sensitive surface 278, a sensor, or a set of sensors that accept user input based on touch and / or tactile contact. The touch-sensitive surface 278 and the remote controller 242 can detect contact (and any movement or interruption of contact) on the touch-sensitive surface 278 and translate the detected contact (e.g., gestures, touch movements, etc.) into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the display unit 126. In some embodiments, the remote control unit 124 may also include other input mechanisms, such as a keyboard, joystick, or the like. In some embodiments, the remote control unit 124 may further include an output mechanism, such as a light, display, speaker, or the like. Inputs received by the remote control device 124 (e.g., user utterances, button presses, touch movements, etc.) can be transmitted to the media device 104 via the remote control device 124. The I / O subsystem 240 may also include one or more other input controllers 244. These other input controllers 244 may be coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses.
[0048] In some embodiments, the media device 104 may further include a memory interface 202 coupled to the memory 250. The memory 250 may be any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, portable computer diskette (magnetic), random access memory (RAM) (magnetic), read-only memory (ROM) (magnetic), erasable programmable read-only memory (EPROM) (magnetic), portable optical discs such as CD, CD-R, CD-RW, DVD, DVD-R, or DVD-RW, or flash memory such as CompactFlash cards, Secure Digital Cards, USB memory devices, Memory Sticks, and the like. In some embodiments, the non-temporary computer-readable storage medium of memory 250 can be used to store instructions (for example, those that execute parts or all of the various processes described herein) for use by or related to instruction execution systems, devices, or other systems that can fetch instructions from and execute those instructions, such as computer-based systems, systems including processors, or instruction execution systems, devices, or other systems. In other embodiments, instructions (for example, those that execute parts or all of the various processes described herein) can be stored on the non-temporary computer-readable storage medium of server system 108, or distributed between the non-temporary computer-readable storage medium of memory 250 and the non-temporary computer-readable storage medium of server system 108. In the context of this specification, “non-temporary computer-readable storage medium” can be any medium capable of containing or storing programs for use by or with instruction execution systems, devices, or other such systems.
[0049] In some embodiments, memory 250 can store an operating system 252, a communication module 254, a graphical user interface (GUI) module 256, an on-device media module 258, an off-device media module 260, and an application module 262. The operating system 252 may include instructions for processing basic system services and instructions for performing hardware-dependent tasks. The communication module 254 may facilitate communication with one or more additional devices, one or more computers, and / or one or more servers. The graphical user interface module 256 can facilitate graphical user interface processing. The on-device media module 258 can facilitate the storage and playback of media content stored locally on the media device 104. The off-device media module 260 can facilitate streaming playback or downloading of media content obtained from external sources (e.g., on a remote server, on user device 122, etc.). Furthermore, the off-device media module 260 can facilitate the reception of broadcast and cable content (e.g., channel tuning). Application module 262 can facilitate various functionalities of media-related applications, such as web browsing, media processing, games, and / or other processes and functions.
[0050] As described herein, memory 250 may also store, for example, client-side digital assistant commands (e.g., within the digital assistant client module 264) to provide client-side functionality for the digital assistant, as well as various user data 266 (e.g., user-specific vocabulary data, preference data, and / or other data such as the user's media search history, media watchlist, recently viewed list, favorite media items, etc.). User data 266 may also be used when performing speech recognition to assist the digital assistant or for any other application.
[0051] In various embodiments, the digital assistant client module 264 may have the ability to accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of the media device 104 (e.g., I / O subsystem 240 or similar). The digital assistant client module 264 may also have the ability to provide output in the form of voice (e.g., speech output), visual, and / or haptic output. For example, the output may be provided as voice, sound, alarm, text message, menu, graphic, video, animation, vibration, and / or a combination of two or more of the above. When operating, the digital assistant client module 264 may communicate with the digital assistant server (e.g., DA server 106) using the communication subsystem 224.
[0052] In some embodiments, the digital assistant client module 264 can utilize various subsystems and peripheral devices to collect additional information related to the media device 104 and from the environment surrounding the media device 104, in order to establish context associated with the user, the current user interaction, and / or the current user input. Such context may also include information from other devices, such as the user device 122. In some embodiments, the digital assistant client module 264 can provide context information or a subset thereof along with user input to the digital assistant server to help infer the user's intent. The digital assistant can also use the context information to determine how to prepare and deliver output to the user. The context information may be further used by the media device 104 or the server system 108 to assist in accurate speech recognition.
[0053] In some embodiments, contextual information associated with user input may include sensor information such as lighting, ambient noise, ambient temperature, distance to another object, and similar information. Contextual information may further include information associated with the physical state of the media device 104 (e.g., device location, device temperature, power level, etc.) or the software state of the media device 104 (e.g., running processes, installed applications, past and present network activity, background services, error logs, resource usage, etc.). Contextual information may further include information received from the user (e.g., utterances), information requested by the user, and information presented to the user (e.g., information currently or previously displayed by the media device). Contextual information may further include information associated with the state of connected devices or other devices associated with the user (e.g., content displayed on user device 122, playable content on user device 122, etc.). Any of these types of contextual information may be provided to the DA server 106 (or used on the media device 104 itself) as contextual information associated with user input.
[0054] In some embodiments, the digital assistant client module 264 can selectively provide information stored on the media device 104 (e.g., user data 266) in response to a request from the DA server 106. In addition, or alternatively, the information can be used on the media device 104 itself when performing speech recognition and / or digital assistant functions. The digital assistant client module 264 can also elicit additional input from the user via a natural language dialog or other user interface in response to a request from the DA server 106. The digital assistant client module 264 can pass additional input to the DA server 106 to assist the DA server 106 in inferring intent and / or achieving the user's intent expressed in the user request.
[0055] In various embodiments, the memory 250 may include additional instructions or fewer instructions. Furthermore, various functions of the media device 104 can be implemented in hardware form and / or firmware form, including in the form of one or more signal processing circuits and / or application-specific integrated circuits. 3. User Devices
[0056] Figure 3 shows a block diagram of an exemplary user device 122 according to various embodiments. As shown in the figure, the user device 122 may include a memory interface 302, one or more processors 304, and a peripheral device interface 306. Various components within the user device 122 may be coupled to one or more communication buses or signal lines. The user device 122 may further include various sensors, subsystems, and peripheral devices coupled to the peripheral device interface 306. The sensors, subsystems, and peripheral devices can collect information and / or facilitate various functionalities of the user device 122.
[0057] For example, the user device 122 may include a motion sensor 310, a light sensor 312, and a proximity sensor 314, which are coupled to the peripheral interface 306 to facilitate orientation, light, and proximity detection functions. To facilitate related functions, one or more other sensors 316, such as a positioning system (e.g., a GPS receiver), a temperature sensor, a biometric sensor, a gyroscope, a compass, an accelerometer, and the like, may also be connected to the peripheral interface 306.
[0058] In some embodiments, the camera subsystem 320 and optical sensor 322 may be used to facilitate camera functions such as taking photographs and recording video clips. Communication functions may be facilitated through one or more wired and / or wireless communication subsystems 324, which may include various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters. To facilitate voice-enabled functions such as voice recognition, voice duplication, digital recording, and telephone functions, a voice subsystem 326 may be coupled to a speaker 328 and microphone 330.
[0059] In some embodiments, the user device 122 may further include an I / O subsystem 340 coupled to a peripheral interface 306. The I / O subsystem 340 may include a touchscreen controller 342 and / or other input controllers(s) 344. The touchscreen controller 342 may be coupled to a touchscreen 346. The touchscreen 346 and the touchscreen controller 342 can detect contact and its movement or interruption using any of several touch sensing technologies, such as capacitive, resistive, infrared, and surface acoustic wave technologies, proximity sensor arrays, and the like. The other input controllers(s) 344 may be coupled to other input / control devices 348, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses.
[0060] In some embodiments, the user device 122 may further include a memory interface 302 coupled to the memory 350. Examples of the memory 350 include any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device; portable computer diskette (magnetic); random access memory (RAM) (magnetic); read-only memory (ROM) (magnetic); erasable programmable read-only memory (EPROM) (magnetic); portable optical discs such as CDs, CD-Rs, CD-RWs, DVDs, DVD-Rs, or DVD-RWs; or flash memory such as compact flash cards, secure digital cards, USB memory devices, memory sticks, and similar devices. In some embodiments, the non-temporary computer-readable storage medium of memory 350 can be used to store instructions (for example, those that execute parts or all of the various processes described herein) for use by or related to instruction execution systems, devices, or other systems that can fetch instructions from and execute those instructions, such as computer-based systems, systems including processors, or instruction execution systems, devices, or other systems. In other embodiments, instructions (for example, those that execute parts or all of the various processes described herein) can be stored on the non-temporary computer-readable storage medium of server system 108, or distributed between the non-temporary computer-readable storage medium of memory 350 and the non-temporary computer-readable storage medium of server system 108. In the context of this specification, “non-temporary computer-readable storage medium” can be any medium capable of containing or storing programs for use by or with instruction execution systems, devices, or other such systems.
[0061] In some embodiments, memory 350 can store an operating system 352, a communication module 354, a graphical user interface (GUI) module 356, a sensor processing module 358, a telephone module 360, and an application module 362. The operating system 352 may include instructions for processing basic system services and instructions for performing hardware-dependent tasks. The communication module 354 may facilitate communication with one or more additional devices, one or more computers, and / or one or more servers. The graphical user interface module 356 can facilitate graphical user interface processing. The sensor processing module 358 can facilitate sensor-related processing and functions. The telephone module 360 may facilitate telephone-related processes and functions. The application module 362 can facilitate various functionalities of user applications, such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions.
[0062] As described herein, the memory 350 may also store, for example, client-side digital assistant commands (e.g., in the digital assistant client module 364) to provide client-side functionality for the digital assistant, as well as various user data 366 (e.g., user-specific vocabulary data, preference data, and / or other data such as the user's electronic address book, to-do list, shopping list, TV show favorites, etc.). The user data 366 may also be used to perform speech recognition to assist the digital assistant or for any other application. The digital assistant client module 364 and the user data 366 may be the same as or identical to the digital assistant client module 264 and user data 266 described above with reference to Figure 2.
[0063] In various embodiments, the memory 350 may include additional instructions or fewer instructions. Furthermore, various functions of the user device 122 can be implemented in hardware form and / or firmware form, including in the form of one or more signal processing circuits and / or application-specific integrated circuits.
[0064] In some embodiments, the user device 122 can be configured to control various aspects of the media device 104. For example, the user device 122 can function as a remote control device (e.g., remote control device 124). User input received through the user device 122 can be transmitted to the media device 104 (e.g., using a communication subsystem) to cause the media device 104 to perform corresponding actions. In addition, the user device 122 can be configured to receive commands from the media device 104. For example, the media device 104 can hand over tasks to the user device 122 to execute and display objects (e.g., selectable affordances) on the user device 122.
[0065] It should be understood that System 100 and Media System 128 are not limited to the components and configurations shown in Figures 1 and 2, and similarly, User Device 122, Media Device 104, and Remote Control Device 124 are not limited to the components and configurations shown in Figures 2 and 3. System 100, Media System 128, User Device 122, Media Device 104, and Remote Control Device 124 may all include fewer components or other components in multiple configurations relating to various embodiments. 4. Digital Assistant System
[0066] Figure 4A shows a block diagram of the digital assistant system 400 according to various embodiments. In some embodiments, the digital assistant system 400 can be implemented on a standalone computer system. In some embodiments, the digital assistant system 400 can be distributed across multiple computers. In some embodiments, some of the modules and functions of the digital assistant can be divided into a server portion and a client portion. In this case, the client portion resides on one or more user devices (e.g., devices 104 or 122) and communicates with the server portion (e.g., server system 108) through one or more networks, for example, as shown in Figure 1. In some embodiments, the digital assistant system 400 can be one implementation of the server system 108 (and / or DA server 106) shown in Figure 1. It should be noted that the digital assistant system 400 is merely one example of a digital assistant system, and the digital assistant system 400 may have more or fewer components than shown, may combine two or more components, or may have different configurations or arrangements of components. The various components shown in Figure 4A can be implemented in the form of hardware, software instructions executed by one or more processors, firmware, or a combination thereof, including one or more signal processing circuits and / or application-specific integrated circuits.
[0067] The digital assistant system 400 may include a memory 402, one or more processors 404, an I / O interface 406, and a network communication interface 408. These components can communicate with each other through one or more communication buses or signal lines 410.
[0068] In some embodiments, the memory 402 may include a non-transient computer-readable medium, such as a high-speed random-access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid memory devices).
[0069] In some embodiments, the I / O interface 406 can connect I / O devices 416 of the digital assistant system 400, such as a display, keyboard, touchscreen, and microphone, to a user interface module 422. The I / O interface 406 can work with the user interface module 422 to receive and process user input (e.g., voice input, keyboard input, touch input, etc.) as appropriate. In some embodiments, for example, when the digital assistant is implemented on a standalone user device, the digital assistant system 400 may include any of the components and I / O communication interfaces described in relation to device 104 or 122 in Figure 2 or 3, respectively. In some embodiments, the digital assistant system 400 may represent a server portion of the digital assistant implementation and can interact with the user through a client-side portion residing on a client device (e.g., user device 104 or 122).
[0070] In some embodiments, the network communication interface 408 may include one or more wired communication ports 412 and / or wireless transmission and reception circuitry 414. The one or more wired communication ports can receive and transmit communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FireWire®, etc. The wireless circuitry 414 can receive and transmit RF signals and / or optical signals to and from the communication network and other communication devices. Wireless communication can use any of several communication standards, protocols and technologies, such as GSM®, EDGE, CDMA, TDMA, Bluetooth®, Wi-Fi®, VoIP, Wi-MAX®, or any other suitable communication protocol. The network communication interface 408 can enable communication between the digital assistant system 400 and devices using a network, such as the Internet, an intranet, and / or a wireless network, such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN).
[0071] In some embodiments, memory 402, or the computer-readable storage medium of memory 402, can store programs, modules, instructions, and data structures, including all or a subset of the operating system 418, communication module 420, user interface module 422, one or more applications 424, and digital assistant module 426. In particular, memory 402, or the computer-readable storage medium of memory 402, can store instructions for executing the processes 800 described later. One or more processors 404 can execute these programs, modules, and instructions, and can read from and write to the data structures.
[0072] An operating system 418 (for example, embedded operating systems such as Darwin®, RTXC®, LINUX®, UNIX®, iOS, OS X®, WINDOWS®, or VxWorks) may include various software components and / or drivers for controlling and managing common system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware, firmware, and software components.
[0073] The communication module 420 can facilitate communication between the digital assistant system 400 and other devices via the network communication interface 408. For example, the communication module 420 can communicate with the communication subsystems (e.g., 224, 324) of electronic devices (e.g., 104, 122). The communication module 420 may also include various components for processing data received by the wireless circuit mechanism 414 and / or the wired communication port 412.
[0074] The user interface module 422 can receive commands and / or input from the user via the I / O interface 406 (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) and generate user interface objects on the display. The user interface module 422 can also prepare outputs (e.g., speech, sound, animation, text, icons, vibration, haptic feedback, light, etc.) and deliver them to the user via the I / O interface 406 (e.g., through a display, audio channel, speaker, touchpad, etc.).
[0075] Application 424 may include programs and / or modules configured to run on one or more processors 404. For example, if the digital assistant system 400 is implemented on a standalone user device, application 424 may include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 400 is implemented on a server, application 424 may include, for example, resource management applications, diagnostic applications, or scheduling applications.
[0076] Memory 402 can also store the digital assistant module 426 (or the server portion of the digital assistant). In some embodiments, the digital assistant module 426 may include the following submodules, or subsets or supersets thereof: I / O processing module 428, speech-to-text (STT) processing module 430, natural language processing module 432, dialogue flow processing module 434, task flow processing module 436, service processing module 438, and speech synthesis module 440. Each of these modules may have access to one or more of the following systems or data and models of the digital assistant module 426, or subsets or supersets thereof: ontology 460, vocabulary index 444, user data 448, task flow model 454, service model 456, and automatic speech recognition (ASR) system 431.
[0077] In some embodiments, using processing modules, data, and models implemented within the digital assistant module 426, the digital assistant can perform at least some of the following: converting utterance input into text; identifying the user's intent expressed in the natural language input received from the user; actively extracting and obtaining the information necessary to fully infer the user's intent (e.g., by removing ambiguity such as words, games, and intentions); determining a task flow to achieve the inferred intent; and executing the task flow to achieve the inferred intent.
[0078] In some embodiments, as shown in Figure 4B, the I / O processing module 428 may interact with the user through the I / O device 416 in Figure 4A, or with an electronic device (e.g., device 104 or 122) through the network communication interface 408 in Figure 4A, in order to acquire user input (e.g., utterance input) and to provide a response to the user input (e.g., utterance output). The I / O processing module 428 may optionally acquire contextual information associated with the user input from the electronic device along with the user input, or immediately after its reception. The contextual information may include user-specific data, vocabulary, and / or preferences related to the user input. In some embodiments, the contextual information may also include the software and hardware state of the electronic device at the time the user request is received, and / or information about the user's surrounding environment at the time the user request is received. In some embodiments, the I / O processing module 428 may also send follow-up questions to the user regarding the user request and receive answers from the user. When a user request is received by the I / O processing module 428 and the user request includes speech input, the I / O processing module 428 can transfer the speech input to the STT processing module 430 (or speech recognition device) for speech-to-text conversion.
[0079] The STT processing module 430 may include one or more ASR systems (e.g., ASR system 431). One or more ASR systems can process utterance input received through the I / O processing module 428 and generate recognition results. Each ASR system may include a front-end utterance preprocessor. The front-end utterance preprocessor can extract representative features from the utterance input. For example, the front-end utterance preprocessor can perform a Fourier transform on the utterance input and extract spectral features that characterize the utterance input as a set of representative multidimensional vectors. Furthermore, each ASR system may include one or more utterance recognition models (e.g., acoustic models and / or language models) and implement one or more utterance recognition engines. Examples of utterance recognition models include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of utterance recognition engines include dynamic time-warping based engines and weighted finite-state transducer (WFST) based engines. Using one or more speech recognition models and one or more speech recognition engines, the extracted representative features of the front-end speech preprocessor can be processed to generate intermediate recognition results (e.g., phonemes, phoneme strings, and subwords), and finally, text recognition results (e.g., words, word strings, or sequences of tokens). In some embodiments, the speech input can be processed at least partially by a third-party service or on an electronic device (e.g., device 104 or 122) to generate recognition results. Once the STT processing module 430 generates recognition results containing text strings (e.g., words, sequences of words, or sequences of tokens), the recognition results can be passed to the natural language processing module 432 for intent inference.
[0080] In some embodiments, one or more language models in one or more ASR systems can be configured to be biased towards media-related results. In one embodiment, one or more language models can be trained using a corpus of media-related text. In another embodiment, the ASR system can be configured to prioritize media-related recognition results. In some embodiments, one or more ASR systems can include static and dynamic language models. Static language models can be trained using a general corpus of text, while dynamic language models can be trained using user-specific text. For example, text corresponding to previous utterances received from a user can be used to generate a dynamic language model. In some embodiments, one or more ASR systems can be configured to generate recognition results based on static and / or dynamic language models. Furthermore, in some embodiments, one or more ASR systems can be configured to prioritize recognition results corresponding to more recently received previous utterances.
[0081] Further details regarding the speech-to-text processing are described in U.S. Utility Patent Application No. 13 / 236,942, “Consolidating Speech Recognition Results,” filed on September 20, 2011. The entire disclosure of that application is incorporated herein by reference.
[0082] In some embodiments, the STT processing module 430 may include and / or access a vocabulary of recognizable words via the phonetic transcription module 431. Each vocabulary word may be associated with one or more candidate pronunciations of a word represented by speech recognition phonetic symbols. In particular, the vocabulary of recognizable words may include words associated with multiple candidate pronunciations. For example, the vocabulary may include: The candidate pronunciation of TIFF0007854032000001.tif6128 may include the word "tomato". Furthermore, lexical words can be associated with custom candidate pronunciations based on previous utterance input from the user. Such custom candidate pronunciations can be stored in the STT processing module 430 and can be associated with a particular user via that user's profile on the device. In some embodiments, candidate pronunciations for a word can be determined based on the spelling of the word, as well as one or more linguistic rules and / or phonetic rules. In some embodiments, candidate pronunciations can be generated manually, for example, based on known standard pronunciations.
[0083] In some embodiments, candidate pronunciations can be ranked based on the generality of the candidate pronunciations. For example, candidate pronunciations It can be ranked higher than TIFF0007854032000002.tif6128 because the former is a more commonly used pronunciation (for example, among all users, for users in a particular geographical area, or for any other appropriate subset of users). In some embodiments, candidate pronunciations can be ranked based on whether they are custom candidate pronunciations associated with a user. For example, custom candidate pronunciations can be ranked higher than standard candidate pronunciations. This can be useful for recognizing proper nouns that have distinctive pronunciations that deviate from standard pronunciations. In some embodiments, candidate pronunciations can be associated with one or more articulation characteristics, such as place of origin, nationality, or ethnicity. For example, candidate pronunciations TIFF0007854032000003.tif6128 can be associated with the United States, and for that, candidate pronunciation TIFF0007854032000004.tif6128 can be associated with the United Kingdom. Furthermore, the ranking of candidate pronunciations can be based on one or more user characteristics (e.g., place of origin, nationality, ethnicity, etc.) stored in the user's profile on the device. For example, from the user's profile, it can be determined that the user is associated with the United States. Based on the user being associated with the United States, the candidate pronunciations TIFF0007854032000005.tif6128 (associated with the United States) is a suggested pronunciation. It can be ranked higher than TIFF0007854032000006.tif6128 (associated with the UK). In some embodiments, one of the ranked candidate pronunciations can be selected as the predicted pronunciation (e.g., the most likely pronunciation).
[0084] When utterance input is received, the STT processing module 430 can be used to determine the phonemes corresponding to the utterance input (for example, using an acoustic model), and then attempt to determine the words that match the phonemes (for example, using a language model). For example, the STT processing module 430 first determines the phoneme sequence corresponding to a portion of the utterance input. If we can identify TIFF0007854032000007.tif6128, then we can determine, based on vocabulary index 444, that this column corresponds to the word "tomato".
[0085] In some embodiments, the STT processing module 430 can use approximate matching techniques to determine words in a utterance. For example, the STT processing module 430 may, even if a particular phoneme sequence is not one of the candidate phoneme sequences for that word, use the phoneme sequence TIFF0007854032000008.tif6128 can be determined to correspond to the word "tomato".
[0086] The digital assistant's natural language processing module 432 ("natural language processor") can acquire a sequence of words or tokens ("token sequence") generated by the STT processing module 430 and attempt to associate the token sequence with one or more "actable intents" recognized by the digital assistant. An "actable intent" can represent a task that can be executed by the digital assistant and may have an associated task flow implemented within the task flow model 454. The associated task flow may be a series of programmed actions and steps that the digital assistant takes to execute the task. The scope of the digital assistant's capabilities may depend on the number and types of task flows implemented and stored within the task flow model 454, or in other words, on the number and types of "actable intents" recognized by the digital assistant. However, the effectiveness of the digital assistant may also depend on the assistant's ability to infer the exact "actable intent(s)" from user requests expressed in natural language.
[0087] In some embodiments, in addition to the sequence of words or tokens obtained from the STT processing module 430, the natural language processing module 432 may also receive contextual information associated with the user request, for example, from the I / O processing module 428. The natural language processing module 432 may optionally use the contextual information to clarify, complement, and / or further clarify the information contained within the sequence of tokens received from the STT processing module 430. The contextual information may include, for example, user preferences, the hardware and / or software state of the user device, sensor information collected before, during, or immediately after the user request, previous interactions (e.g., dialogues) between the digital assistant and the user, and similar. As described herein, the contextual information may be dynamic and may change with time, location, dialogue content, and other factors.
[0088] In some embodiments, natural language processing can be based on, for example, Ontology 460. Ontology 460 is a hierarchical structure encompassing numerous nodes, each node can represent either an "implementable intent" or an "attribute" related to one or more of the "implementable intents" or other "attributes." As mentioned above, an "implementable intent" can represent a task that a digital assistant is capable of performing; that is, it is "implementable" or can be made an object of implementation. An "attribute" can represent a parameter associated with an implementable intent or a sub-part of another attribute. Links between implementable intent nodes and attribute nodes within Ontology 460 can define how the parameter represented by the attribute node relates to the task represented by the implementable intent node.
[0089] In some embodiments, ontology 460 can consist of actionable intent nodes and attribute nodes. Within ontology 460, each actionable intent node can be linked directly to one or more attribute nodes or via one or more intermediate attribute nodes. Similarly, each attribute node can be linked directly to one or more actionable intent nodes or via one or more intermediate attribute nodes. For example, as shown in Figure 4C, ontology 460 may include a “Media” node (i.e., an actionable intent node). The attribute nodes “Actor(singular or plural)”, “Media Genre”, and “Media Title” can each be directly linked to an actionable intent node (i.e., a “Media Search” node). In addition, the attribute nodes “Name”, “Age”, “Ulmer Scale Ranking”, and “Nationality” can be subnodes of the attribute node “Actor”.
[0090] In another embodiment, as shown in Figure 4C, ontology 460 may also include a “Weather” node (i.e., another actionable intent node). The attribute nodes “Date / Time” and “Location” may each be linked to a “Weather Search” node. It should be noted that in some embodiments, one or more attribute nodes may be associated with two or more actionable intents. In these embodiments, one or more attribute nodes may be linked to the respective nodes corresponding to two or more actionable intents within ontology 460.
[0091] An actionable intent node, along with its linked conceptual node, can be described as a “domain.” In this description, each domain can be associated with a particular actionable intent and can refer to a group of nodes (and relationships between nodes) associated with a specific actionable intent. For example, ontology 460 shown in Figure 4C may include an example media domain 462 and an example weather domain 464 within ontology 460. Media domain 462 may include the actionable intent node “media search,” as well as attribute nodes “actor(single or plural),” “media genre,” and “media title.” Weather domain 464 may include the actionable intent node “weather search,” as well as attribute nodes “location” and “date / time.” In some embodiments, ontology 460 may consist of many domains. Each domain may share one or more attribute nodes with one or more other domains.
[0092] Figure 4C shows two exemplary domains within Ontology 460, but other domains could include, for example, “Athletes,” “Stock Prices,” “Directions,” “Media Settings,” “Sports Teams,” and “Time,” “Tell a Joke.” The “Athletes” domain can be associated with the “Search Athlete Information” actionable intent node and may further include attribute nodes such as “Athlete Name,” “Athlete Team,” and “Athlete Statistics.”
[0093] In some embodiments, ontology 460 may include all domains (and therefore implementable intentions) that the digital assistant is capable of understanding and acting upon. In some embodiments, ontology 460 may be modified by adding or removing domains or entire nodes, or by changing the relationships between nodes within ontology 460.
[0094] In some embodiments, each node in ontology 460 may be associated with a set of words and / or phrases related to the attribute or actionable intent represented by that node. Each set of words and / or phrases associated with each node may be the so-called "vocabulary" associated with that node. Each set of words and / or phrases associated with each node may be stored in the vocabulary index 444 in relation to the attribute or actionable intent represented by the node. For example, returning to Figure 4C, the vocabulary associated with the node for the attribute "actor" may include words such as "A list," "Reese Witherspoon," "Arnold Schwarzenegger," "Brad Pitt," etc. As another example, the vocabulary associated with the node for the actionable intent "weather search" may include words and phrases such as "weather," "what's it like in," "forecast," etc. The vocabulary index 444 may optionally include words and phrases from different languages.
[0095] The natural language processing module 432 receives a token sequence (e.g., a text string) from the STT processing module 430 and can determine which nodes are implied by the words in the token sequence. In some embodiments, if it is found that a word or phrase in the token sequence is associated with one or more nodes in the ontology 460 (via the lexical index 444), then that word or phrase can “trigger” or “activate” those nodes. Based on the number and / or relative importance of the activated nodes, the natural language processing module 432 can select one of the possible intentions as the task the user intended the digital assistant to perform. In some embodiments, the domain with the most “triggered” nodes can be selected. In some embodiments, the domain with the highest confidence value (e.g., based on the relative importance of its various triggered nodes) can be selected. In some embodiments, the domain can be selected based on a combination of the number and importance of the triggered nodes. In some embodiments, additional factors such as whether the digital assistant has previously accurately interpreted a similar request from the user are also considered when selecting nodes.
[0096] User data 448 may include user-specific information such as user-specific vocabulary, user preferences, user address, user's default and second language, user's contact list, and other short-term or long-term information about each user. In some embodiments, the natural language processing module 432 may use user-specific information to complement information contained in user input and further clarify user intent. For example, for the user request "What's the weather like this week?", instead of asking the user to explicitly provide such information in its request, the natural language processing module 432 may access user data 448 to determine where the user is located.
[0097] Further details of token string-based ontology searching are described in U.S. Utility Patent Application No. 12 / 341,743, filed December 22, 2008, for “Method and Apparatus for Searching Using an Active Ontology.” The entire disclosure of that application is incorporated herein by reference.
[0098] In some embodiments, once the natural language processing module 432 identifies an actionable intent (or domain) based on a user request, the natural language processing module 432 can generate a structured query to represent the identified actionable intent. In some embodiments, the structured query may include parameters for one or more nodes within the domain for the actionable intent, with at least some of the parameters being populated with specific information and requirements specified in the user request. For example, the user might say, "Find other seasons of this TV series." In this case, the natural language processing module 432 can accurately identify the actionable intent as "media search" based on the user input. According to the ontology, a structured query for the "media" domain may include parameters such as {media actors}, {media genre}, {media title}, and so on. In some embodiments, based on utterance input and the text derived from the utterance input using the STT processing module 430, the natural language processing module 432 can generate a partially structured query for a restaurant reservation domain. In this case, the partially structured query includes the parameter {media genre = "TV series"}. However, in this example, the user's utterance does not contain enough information to complete the structured query associated with the domain. Therefore, other necessary parameters such as {media title} do not need to be specified in the structured query based on the currently available information. In some embodiments, the natural language processing module 432 can populate some parameters of the structured query with received contextual information. For example, the TV series "Mad Men" may currently be playing on a media device. Based on this contextual information, the natural language processing module 432 can populate the {media title} parameter in the structured query with "Mad Men".
[0099] In some embodiments, the natural language processing module 432 can pass the generated structured query (including any completed parameters) to the task flow processing module 436 ("task flow processor"). The task flow processing module 436 can receive the structured query from the natural language processing module 432, complete the structured query as needed, and perform the actions required to "complete" the user's final request. In some embodiments, the various steps required to complete these tasks can be provided within the task flow model 454. In some embodiments, the task flow model 454 may include steps for obtaining additional information from the user and task flows for performing actions associated with the actionable intent.
[0100] As described above, in order to complete a structured query, the task flow processing module 436 may need to initiate an additional dialogue with the user to obtain additional information and / or to remove ambiguity from potentially ambiguous statements. When such an dialogue is necessary, the task flow processing module 436 can call the dialogue flow processing module 434 to engage in a dialogue with the user. In some embodiments, the dialogue flow processing module 434 can determine how (and / or when) to ask the user for additional information, receive user responses, and process them. It can provide questions to the user through the I / O processing module 428 and receive answers from the user. In some embodiments, the dialogue flow processing module 434 can present the dialogue output to the user via audio and / or visual output and receive input from the user via verbal or physical (e.g., click) responses. For example, the user might ask, "What's the weather like in Paris?" When the task flow processing module 436 calls the dialog flow processing module 434 to determine the "location" information for a structured query associated with the domain "weather search", the dialog flow processing module 434 can generate a question, such as "Which Paris?", to pass to the user. In addition, the dialog flow processing module 434 can present affordances associated with "Paris, Texas" and "Paris, France" for user selection. Once a response is received from the user, the dialog flow processing module 434 can then pass that information to the task flow processing module 436 to either populate the structured query with missing information or to complete the structured query with missing information.
[0101] Once the task flow processing module 436 completes the structured query for an actionable intent, it can proceed to execute the final task associated with the actionable intent. Accordingly, the task flow processing module 436 can execute steps and instructions within the task flow model 454, depending on the specific parameters contained within the structured query. For example, a task flow model for the actionable intent of "media search" may include steps and instructions to execute a media search query and retrieve relevant media items. For example, using a structured query such as {media search, media genre=TV series, media title=Mad Men}, the task flow processing module 436 may execute (1) a media search query using a media database and retrieve relevant media items, (2) a step to rank the retrieved media items according to relevance and / or popularity, and (3) a step to display the media items sorted according to relevance and / or popularity.
[0102] In some embodiments, the task flow processing module 436 may use the assistance of the service processing module 438 ("service processing module") to complete a task requested in user input or to provide a response to information requested in user input. For example, the service processing module 438 may act on behalf of the task flow processing module 436 to perform media searches, retrieve weather information, call or interact with applications installed on other user devices, and call or interact with third-party services (e.g., social networking websites, media review websites, media subscription services, etc.). In some embodiments, the protocols and APIs required by each service may be specified by the respective service models in the service model 456. The service processing module 438 may access the appropriate service model for a service and generate a request for the service in accordance with the protocols and APIs required by the service relating to the service model.
[0103] For example, a third-party media search service may submit a service model that specifies the parameters required to perform a media search, and an API for communicating the values of the required parameters to the media search service. When requested by the task flow processing module 436, the service processing module 438 can establish a network connection with the media search service and send the required parameters for the media search (e.g., media actors, media genre, media title) to the online booking interface in a format corresponding to the media search service's API.
[0104] In some embodiments, the natural language processing module 432, the dialogue flow processing module 434, and the task flow processing module 436 can be used collectively and iteratively to infer and clarify the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (e.g., output to the user or completion of a task) to achieve the user's intent. The generated response may be a dialogue response to an utterance input that at least partially achieves the user's intent. Furthermore, in some embodiments, the generated response may be output as an utterance output. In these embodiments, the generated response may be sent to an utterance synthesis module 440 (e.g., an utterance synthesizer), where it can be processed to synthesize a dialogue response in utterance form. In yet another embodiment, the generated response may be data content related to satisfying the user request in the utterance input.
[0105] The speech synthesis module 440 can be configured to synthesize speech output for presentation to the user. The speech synthesis module 440 synthesizes speech output based on text provided by a digital assistant. For example, the generated dialogue response may be in the form of a text string. The speech synthesis module 440 can convert the text string into audible speech output. The speech synthesis module 440 can use any suitable speech synthesis technique to generate speech output from text, but is not limited to, waveform concatenation synthesis, unit selection synthesis, diphone synthesis, domain-limited synthesis, formant synthesis, articulation synthesis, hidden Markov model (HMM) based synthesis, and sine wave synthesis. In some embodiments, the speech synthesis module 440 can be configured to synthesize individual words based on phoneme strings corresponding to words. For example, phoneme strings may be associated with words in the generated dialogue response. Phoneme strings may be stored in metadata associated with the words. The speech synthesis module 440 can be configured to directly process phoneme strings in metadata in order to synthesize words in utterance form.
[0106] In some embodiments, instead of using (or in addition to) the speech synthesis module 440, speech synthesis can be performed on a remote device (e.g., a server system 108), and the synthesized speech can be sent to a user device for output to the user. For example, this can be done in some implementations where the output for the digital assistant is generated on the server system. Also, since the server system generally has more processing power or resources than the user device, it may be possible to obtain higher quality speech output than would be possible using client-side synthesis.
[0107] Further details regarding the digital assistant can be found in U.S. Utility Patent Application No. 12 / 987,982, filed on January 10, 2011, entitled "Intelligent Automated Assistant," and U.S. Utility Patent Application No. 13 / 251,088, filed on September 30, 2011, entitled "Generating and Processing Task Items That Represent Tasks to Perform." The entire disclosures of these applications are incorporated herein by reference. 4. Process for interacting with digital assistants within a media environment
[0108] Figures 5A to 5I illustrate process 500 for operating a digital assistant in a media system, according to various embodiments. Process 500 can be performed using one or more electronic devices that implement the digital assistant. For example, process 500 can be performed using one or more of the above-described systems 100, media system 128, media device 104, user device 122, or digital assistant system 400. Figures 6A to 6Q show screenshots displayed on a display unit by the media device at various stages of process 500, according to various embodiments. Process 500 will be described below with simultaneous reference to Figures 5A to 5I and Figures 6A to 6Q. It should be understood that some operations within process 500 can be combined, the order of some operations can be changed, and some operations can be omitted.
[0109] In block 502 of process 500, content can be displayed on a display unit (e.g., display unit 126). In this embodiment shown in Figure 6A, the displayed content may include media content 602 (e.g., movies, videos, television programs, video games, etc.) that is being played on a media device (e.g., media device 104). In other embodiments, the displayed content may include content associated with an application running on the media device, or other content associated with the media device, such as a user interface for interacting with the media device's digital assistant. Specifically, the displayed content may include a main menu user interface, or a user interface (e.g., a second user interface 618 or a third user interface 626) that has an object or result previously requested by the user.
[0110] In block 504 of process 500, user input can be detected. User input can be detected while the content of block 502 is being displayed. In some embodiments, user input can be detected on a remote control device of the media device (e.g., remote control device 124). Specifically, user input can be a user interaction with the remote control device, such as pressing a button on the remote control device (e.g., button 274) or touching a touch-sensitive surface (e.g., touch-sensitive surface 278). In some embodiments, user input can be detected via a second electronic device (e.g., device 122) configured to interact with the media device. Depending on the detection of user input, one or more of blocks 506 to 592 can be executed.
[0111] In block 506 of process 500, a determination can be made as to whether user input corresponds to a first input format. The first input format can be a default input to a media device. For example, the first input format may include pressing a specific button on a remote control device and releasing the button within a predetermined period of time after pressing it (e.g., a short press). The media device can determine whether user input matches the first input format. Based on the determination that user input corresponds to the first input format, one or more of blocks 508 to 514 can be executed.
[0112] In block 508 of process 500, and referring to Figure 6B, a text instruction 604 can be displayed for invoking and interacting with the digital assistant. Specifically, the instruction 604 can describe the user input required to invoke and interact with the digital assistant. For example, the instruction 604 can describe how to perform the second input format, which will be described later in block 516.
[0113] In block 510 of process 500, and as shown in Figure 6B, a passive visual indicator 606 can be displayed on the display unit. The passive visual indicator 606 can indicate that the digital assistant has not yet been invoked. Specifically, the microphone of the media device (e.g., microphone 272) may not be activated in response to the detection of user input. Therefore, the passive visual indicator 606 can serve as a visual signal that the digital assistant is not processing voice input. In this example, the visual indicator 606 can be a passive, flat waveform that does not respond to user utterances. Furthermore, the passive visual indicator 606 may include achromatic colors (e.g., black, gray, etc.) to indicate its passive status. It should be noted that other visual patterns or images may also be intended for the passive visual indicator. The passive visual indicator 606 can be displayed simultaneously with the instruction 604. Furthermore, the passive visual indicator 606 can be displayed continuously while one or more of blocks 512-514 are being executed.
[0114] Referring to block 512 of process 500 and Figure 6C, instructions 608 for performing a typed search can be displayed on the display unit. Specifically, instructions 608 can describe the user input necessary to display a virtual keyboard interface that can be used to perform a typed search. In some embodiments, instructions 604 for summoning and interacting with a digital assistant, and instructions 608 for performing a typed search, can be displayed sequentially at different times. For example, the display of instructions 608 may supersede the display of instructions 604, or vice versa. In this example, instructions 604 and 608 are in text format. It should be noted that in other embodiments, instructions 604 and 608 can be in graphical format (e.g., pictures, symbols, animations, etc.).
[0115] In block 514 of process 500, one or more exemplary natural language requests can be displayed on the display unit. For example, Figures 6D to 6E show two different exemplary natural language requests 610 and 612 displayed on the display unit. In some embodiments, exemplary natural language requests can be displayed on the display unit via a first user interface. The first user interface can be overlaid on the displayed content. Exemplary natural language requests can provide the user with guidance for interacting with the digital assistant. Furthermore, exemplary natural language requests can inform the user of various functions of the digital assistant. In response to receiving a user utterance corresponding to one of the exemplary natural language requests, the digital assistant can perform the respective action. For example, if the digital assistant of a media device is invoked (e.g., by user input in a second input format in block 504) and the user utterance "Jump 30 seconds ahead" is provided (e.g., in block 518), the digital assistant can jump forward 30 seconds to the media content playing on the media device.
[0116] The displayed exemplary natural language requests can be contextually related to the displayed content (e.g., media content 602). For example, a set of exemplary natural language requests may be stored on the media device or on a separate server. Each exemplary natural language request in the set of exemplary natural language requests may be associated with one or more contextual attributes (e.g., media content being played, homepage, iTunes® media store, actors, movies, weather, sports, stock prices, etc.). In some embodiments, block 514 may include identifying from the set of exemplary natural language requests an exemplary natural language request that has contextual attributes corresponding to the content displayed on the display unit. The identified exemplary natural language request can then be displayed on the display unit. Therefore, different exemplary natural language requests may be displayed depending on the content displayed on the display unit. Displaying contextually related exemplary natural language requests can help the user conveniently know which digital assistant functions are most relevant to the user's current usage on the media device. This can improve the overall user experience.
[0117] In the embodiment shown in Figures 6D to 6E, exemplary natural language requests 610 and 612 can each be contextually related to media content 602 on the display unit. Specifically, exemplary natural language requests 610 and 612 can be requests to change or control one or more settings associated with media content being played on the media device. Such exemplary natural language requests may include requests to turn closed captions on or off, turn on subtitles in a specific language, rewind or skip forward, pause playback of media content, restart playback of media content, decrease or increase the playback speed of media content, increase or decrease the volume of media content (e.g., audio gain), and similar requests. Furthermore, other exemplary natural language requests that are contextually relevant to media content 602 may include requests to add a media item corresponding to media content 602 to the user's watchlist, to display information related to media content 602 (e.g., actor information, synopsis, release date, etc.), to display other media items or content related to media content 602 (e.g., the same series, the same season, the same actor / director, the same genre, etc.), and similar requests.
[0118] In embodiments where the displayed content includes content associated with an application on a media device, a contextually relevant exemplary natural language request may include a request to change one or more settings or states of the application. Specifically, an exemplary natural language request may include a request to open or close the application, or to operate one or more functions of the application.
[0119] In some embodiments, the displayed content may include a user interface (e.g., a second user interface 618 or a third user interface 626) for searching, browsing, or selecting items. Specifically, the displayed user interface may include one or more media items. Furthermore, the focus of the user interface may be on one of the media items (e.g., media item 623 highlighted by cursor 624 in Figure 6G). In these embodiments, contextually relevant exemplary natural language requests may include information relating to one or more media items within the displayed user interface or requests for other media items. Specifically, exemplary natural language requests may include requests relating to the media item that is currently in focus of the user interface. In these embodiments, exemplary natural language requests may include requests such as "What is this?", "What is the rating of this?", "Who is in this?", "When is the next episode out?", "Can you recommend more movies similar to this?", and "Can you recommend movies starring the same actors?". In certain embodiments, the user interface may display information relating to a media item or a set of media items, such as the television series Mad Men. In this embodiment, a contextually relevant exemplary natural language request may include a request based on one or more attributes of a media item or a set of media items (e.g., cast, plot, rating, release date, director, provider, etc.) (e.g., "Other shows featuring January Jones."). In addition, a contextually relevant exemplary natural language request may include a request to play, select, or obtain the focused media item or another media item displayed within the user interface (e.g., "Rent this," "Play this," "Buy this," or "Play How to Train Your Dragon 2.").The requests may include, for example, "Go to the comedy section," or "Jump to the horror movie section," or requests to navigate between media items within the user interface. Furthermore, in these embodiments, the contextually relevant exemplary natural language requests may include requests to search for other media items (for example, "Find a new comedy," "Tell me some free classic movies," or "What are some shows starring Nicole Kidman?").
[0120] In some embodiments, the displayed content may include media items organized according to a specific category or topic. In these embodiments, a contextually relevant exemplary natural language request may include requests related to that specific category or topic. For example, in an embodiment where the displayed content includes media items organized according to various actors, a contextually relevant exemplary natural language request may include requests for information or media items related to the actors (e.g., "What movies star Jennifer Lawrence?", "How old is Scarlett Johansson?", or "What is Brad Pitt's latest movie?"). In another embodiment where the displayed content includes media items organized according to a programming channel or content provider (e.g., a channel page or TV guide page), a contextually relevant exemplary natural language request may include requests for information or media items related to the programming channel or content provider (e.g., "What shows are on in an hour?", "What's on HBO in prime time?", "Tune in to ABC," or "Which channel is showing basketball?"). In yet another embodiment, where the displayed content includes media items recently selected by the user (e.g., a "Recently Played" list) or media items identified as of user interest (e.g., a "Watchlist"), a contextually relevant exemplary natural language request may include a request to watch or continue watching one of the media items (e.g., "Resume where I left off," "Continue watching Birdman," or "Play this again from the beginning").
[0121] In some embodiments, the displayed content may include a user interface that contains results or information corresponding to a specific topic. Specifically, the results may be associated with a previous user request (e.g., a request to a digital assistant) and may include information corresponding to a topic such as weather, stock prices, or sports. In these embodiments, contextually relevant exemplary natural language requests may include requests to refine the results or requests for additional information related to a specific topic. For example, in an embodiment where the displayed content includes weather information for a specific location, contextually relevant exemplary natural language requests may include requests to display additional weather information for a different location or for a different time of day (e.g., "What's it like in New York City?", "What's it going to be like next week?", "And what about Hawaii?", etc.). In another embodiment where the displayed content includes information related to a sports team or athlete, a contextually relevant exemplary natural language request may include a request for additional information related to the sports team or athlete (e.g., "How tall is Shaquille O'Neal?", "When was Tom Brady born?", "When is the 49ers' next game?", "How did Manchester United do in their last game?", "Who is the point guard for the LA Lakers?"). In yet another embodiment where the displayed content includes information related to stock prices, a contextually relevant exemplary natural language request may include a request for additional stock-related information (e.g., "What was the opening price of the S&P 500?", "How is Apple doing?", "What was the closing price of the Dow Jones yesterday?"). Furthermore, in some embodiments, the displayed content may include a user interface that encompasses media search results associated with a previous user request.In these embodiments, contextually relevant exemplary natural language requests may include requests to refine the displayed media search results (e.g., "only from last year," "only those rated G," "only those that are free") or requests to perform a different media search (e.g., "find me a good action movie," "tell me some Jackie Chan movies," etc.).
[0122] In some embodiments, the displayed content may include the main menu user interface of a media device. The main menu user interface may be, for example, the home screen or root directory of the media device. In these embodiments, contextually relevant exemplary natural language requests may include requests that represent various functions of the digital assistant. Specifically, the digital assistant may have a set of core capabilities associated with the media device, and contextually relevant exemplary natural language requests may include requests related to each of the digital assistant's core capabilities (e.g., "Can you recommend some good free movies?", "What's the weather like?", "Play the next episode of Breaking Bad", or "What's Apple's stock price?").
[0123] Example natural language requests can be in natural language form. This can help inform the user that the digital assistant has the ability to understand natural language requests. Furthermore, in some embodiments, example natural language requests can be contextually ambiguous to inform the user that the digital assistant has the ability to infer the correct user intent associated with the user's request based on the displayed content. Specifically, as shown in the embodiments described above, example natural language requests may include contextually ambiguous terms such as "this" or "ones," or contextually ambiguous phrases such as "only free things" or "How about in New York?" These example natural language requests can inform the user that the digital assistant has the ability to determine the correct context associated with such requests based on the displayed content. This encourages the user to rely on the context of the displayed content when interacting with the digital assistant. This can be desirable to promote a more natural conversational experience with the digital assistant.
[0124] In some embodiments, block 514 can be executed after blocks 508-512. Specifically, an exemplary natural language request can be displayed on the display unit after a predetermined time has elapsed since block 506 determined that the user input corresponds to a first input format. It should be noted that in some embodiments, blocks 508-514 can be executed in any order, and in some embodiments, two or more of blocks 508-514 can be executed simultaneously.
[0125] In some embodiments, exemplary natural language requests are displayed alternately in a predetermined order. Each exemplary natural language request can be displayed separately at different times. Specifically, the display of the current exemplary natural language request can be replaced by the display of a subsequent exemplary natural language request. For example, as shown in Figure 6D, exemplary natural language request 610 may be displayed first. After a predetermined time, the display of exemplary natural language request 610 ("Please skip ahead 30 seconds") can be replaced by the display of exemplary natural language request 612 ("Please play the next episode"), as shown in Figure 6E. Therefore, in this embodiment, exemplary natural language requests 610 and 612 are displayed one at a time, not simultaneously.
[0126] In some embodiments, exemplary natural language requests can be grouped into multiple lists, each containing one or more exemplary natural language requests. In these embodiments, block 514 may include displaying lists of exemplary natural language requests on a display unit. Each list may be displayed at different times in a predetermined order. Furthermore, the lists may be displayed alternately.
[0127] While one or more of blocks 508-514 are being executed, the displayed content can continue to be displayed on the display unit. For example, as shown in Figures 6B-6E, media content 602 can continue to be played on the media device and displayed on the display unit while blocks 508-512 are being executed. Furthermore, while the media content is being played, audio associated with the media content can be output by the media device. In some embodiments, the amplitude of the audio is not reduced in response to the detection of user input or in accordance with the determination that the user input corresponds to a first input format. This may be desirable to reduce interruptions in the consumption of the media content 602 being played. Therefore, the user can continue to follow the media content 602 via the audio output even though elements 604-612 are displayed on the display unit.
[0128] In some embodiments, the brightness of the displayed content can be reduced (for example, by 20-40%) in response to detection of user input or determination that the user input corresponds to a first input format, as represented by the hollow font of the media content 602 in Figures 6B-6D. In these embodiments, the displayed elements 604-612 can be superimposed on the displayed media content 602. Reducing the brightness can help make the displayed elements 604-612 stand out. At the same time, the media content 602 can still be recognized on the display unit, thereby allowing the user to continue consuming the media content 602 while elements 604-612 are displayed.
[0129] While executing one of blocks 508-512, the digital assistant can be invoked (for example, by detecting a second input form of user input in block 504) and a user utterance corresponding to one of the exemplary natural language requests can be received (for example, in block 518). The digital assistant can then perform a task in response to the received request (for example, in block 532). Further details regarding invoking and interacting with the digital assistant are provided below with reference to Figures 5B-5I. Furthermore, while executing one of blocks 508-512, a virtual keyboard interface for performing a type search can be invoked (for example, by detecting a fifth user input in block 558). Further details regarding invoking the virtual keyboard interface and performing a type search are provided below with reference to Figure 5G.
[0130] Referring again to block 506, if it is determined that the user input does not correspond to a first input format, one or more of blocks 516-530 in Figure 5B can be performed. In block 516, a determination can be made as to whether the user input corresponds to a second input format. The second input format may be a default input to a different media device than the first input format. In some embodiments, the second input format may include pressing a specific button on the remote control of the media device and holding the button down for a longer period than a predetermined time (e.g., a long press). The second input format may be associated with invoking a digital assistant. In some embodiments, the first and second input formats may be performed using the same button on the remote control (e.g., a button configured to invoke a digital assistant). This may be desirable to intuitively integrate invoking a digital assistant and providing instructions for invoking and interacting with the digital assistant into a single button. Furthermore, inexperienced users may intuitively perform a short press rather than a long press. Therefore, by providing instructions in response to a short press, it becomes possible to primarily direct the instructions to less experienced users rather than experienced users. This improves the user experience by allowing experienced users the option to bypass the instructions, while making them easily accessible to less experienced users who most need guidance.
[0131] According to the determination in block 516 that the user input corresponds to a second input format, one or more of blocks 518 to 530 can be executed. In some embodiments, the media content 602 can continue to play on the media device while one or more of blocks 518 to 530 are being executed. Specifically, the media content 602 can continue to play on the media device and be displayed on the display unit while audio data is being sampled in block 518 and while a task is being executed in block 528.
[0132] In block 518 of process 500, audio data can be sampled. Specifically, the first microphone of the media device (e.g., microphone 272) can be activated and sampling of audio data can be started. In some embodiments, the sampled audio data may include user utterances from the user. User utterances can represent user requests directed to the digital assistant. Furthermore, in some embodiments, the user request may be a request to perform a task. Specifically, the user request may be a media search request. For example, referring to Figure 6F, the sampled audio data may include the user utterance, "Find a romantic comedy starring Reese Witherspoon." In other embodiments, the user request may be a request to play a media item or to provide specific information (e.g., weather, stock prices, sports, etc.).
[0133] User utterances within sampled audio data can be in natural language form. In some embodiments, user utterances can express user requests that are incompletely specified. In this case, the user utterance does not explicitly limit all the information necessary to satisfy the user request. For example, the user utterance could be "Play the next episode." In this embodiment, the user request does not explicitly limit the media series on which the next episode should be played. Furthermore, in some embodiments, the user utterance may contain one or more ambiguous terms.
[0134] The period over which audio data is sampled can be based on endpoint detection. Specifically, audio data can be sampled from the start time when the second form of user input is first detected until the end time when the endpoint is detected. In some embodiments, the endpoint can be based on user input. Specifically, the first microphone can be activated when the second form of user input (e.g., pressing a button for a longer period than predetermined) is first detected. The first microphone can remain activated to sample audio data as long as the second form of user input continues to be detected. The first microphone can be deactivated when the detection of the second form of user input ceases (e.g., the button is released). Therefore, in these embodiments, the endpoint is detected when the end of user input is detected. Thus, audio data is sampled while the second form of user input is being detected.
[0135] In other embodiments, endpoint detection can be based on one or more speech characteristics of sampled speech data. Specifically, one or more speech characteristics of sampled speech data can be monitored, and the endpoint can be detected a predetermined time after it is determined that one or more speech characteristics do not satisfy one or more predetermined criteria. In yet another embodiment, the endpoint can be detected based on a fixed period of time. Specifically, the endpoint can be detected a predetermined period of time after the first detection of user input in a second input format.
[0136] In some embodiments, while block 504 or 516 is being executed, audio associated with the displayed content can be output (for example, using speaker 268). Specifically, the audio may be the sound of a media item played on a media device and displayed on the display unit. The audio may be output from the media device via an audio signal. In these embodiments, when it is determined that user input corresponds to a second input format and when audio data is sampled, the audio associated with the displayed content can be ducked (for example, by reducing the amplitude of the audio). For example, the audio can be ducked by reducing the gain associated with the audio signal. In other embodiments, while audio data is being sampled in block 518, the output of audio associated with the media content can be stopped. For example, the audio can be stopped by cutting off or interrupting the audio signal. Ducking or stopping the output of the audio may be desirable to reduce background noise in the sampled audio data and increase the relative intensity of the utterance signal associated with the user utterance. Furthermore, the ducking or stopping of the audio can serve as an audio cue for the user to begin providing utterance input to the digital assistant.
[0137] In some embodiments, background audio data can be sampled while audio data is being sampled in order to perform noise cancellation. In these embodiments, the remote control or media device may include a second microphone. The second microphone may be oriented in a different direction from the first microphone (e.g., opposite to the first microphone). The second microphone may be activated to sample background audio data while audio data is being sampled. In some embodiments, background audio data can be used to remove background noise in the audio data. In other embodiments, the media device may generate an audio signal to output audio associated with the displayed content. The generated audio signal can be used to remove background noise from the audio data. Performing noise cancellation of background noise from the audio signal may be particularly suitable for interaction with a digital assistant in a media environment. This may be due to the communal nature of consuming media content, where speech from multiple people may be intermingled in the audio data. By removing background noise in the audio data, a higher signal-to-noise ratio can be obtained in the audio data. This may be desirable when processing audio data for user requests.
[0138] In block 520 of process 500, and referring to Figure 6F, an active visual indicator 614 can be displayed on the display unit. The active visual indicator 614 can indicate to the user that the digital assistant has been invoked and is actively listening. Specifically, the active visual indicator 614 can serve as a visual cue to prompt the user to begin providing speech input to the digital assistant. In some embodiments, the active visual indicator 614 may include colors and / or visual animations to indicate that the digital assistant has been invoked. For example, as shown in Figure 6F, the active visual indicator 614 may include an active waveform that responds to one or more characteristics (e.g., amplitude) of the audio data received by the digital assistant. For example, the active visual indicator 614 may display a waveform with a larger amplitude in proportion to portions of the audio data where the sound is louder, and a waveform with a smaller amplitude in proportion to portions of the audio data where the sound is quieter. Furthermore, in embodiments where the digital assistant is invoked while a passive visual indicator 606 (e.g., Figure 6E) is displayed, the display of the visual indicator 606 can be replaced with the display of the active visual indicator 614. This can provide a smooth transition from the instructional user interface shown in Figures 6B–6E, which illustrates how to invoke and interact with the digital assistant, to the active user interface shown in Figure 6F, which allows for active interaction with the digital assistant.
[0139] In block 522 of process 500, the text representation of user utterances in sampled audio data can be determined. For example, the text representation can be determined by performing speech-to-text (STT) processing on the sampled audio data. Specifically, the sampled audio data can be processed using an STT processing module (e.g., STT processing module 430) to convert user utterances in the sampled audio data into text representations. The text representation can be a token string that represents the corresponding text string.
[0140] In some embodiments, STT processing can be biased towards media-related text results. This bias can be achieved by utilizing a language model trained on a corpus of media-related text. Additionally, or alternatively, bias can be achieved by giving greater weight to media-related text result candidates. In this way, bias can be applied to rank media-related text result candidates higher than if no bias were applied. Bias may be desirable to improve the accuracy of STT processing of media-related user utterances (e.g., movie titles, movie actors, etc.). For example, certain media-related words or phrases such as "Jurassic Park," "Arnold Schwarzenegger," and "Shrek" are rarely found in a typical text corpus and therefore may not be properly recognized during STT processing without bias towards media-related text results.
[0141] In some embodiments, the text representation can be obtained from a separate device (e.g., DA server 106). Specifically, the sampled audio data can be transmitted from the media device to the separate device for STT processing. In these embodiments, the media device can instruct the separate device (e.g., via data transmitted to the separate device along with the sampled audio data) that the sampled audio data is associated with a media application. This instruction can cause the STT processing to be biased towards media-related text results.
[0142] In some embodiments, the text representation can be based on previous user utterances received by the media device before the audio data is sampled. Specifically, candidate text results for the sampled audio data can be weighted more heavily if they correspond to one or more portions of the previous user utterances. In some embodiments, a language model can be generated using the previous user utterances, and the generated language model can be used to determine the text representation of the current user utterance in the sampled audio data. The language model can be dynamically updated as additional user utterances are received and processed.
[0143] Furthermore, in some embodiments, the text representation can be based on the time when previous user utterances were received before the audio data was sampled. Specifically, text result candidates corresponding to more recently received previous user utterances can be weighted more heavily than text result candidates corresponding to earlier received previous user utterances relative to the sampled audio data.
[0144] In block 524 of process 500, a text representation can be displayed on the display unit. For example, Figure 6F shows a text representation 616 corresponding to a user utterance in sampled audio data. In some embodiments, blocks 522 and 524 can be executed while the audio data is being sampled. Specifically, the text representation 616 of the user utterance can be displayed in a streaming manner so that the text representation 616 is displayed in real time as the audio data is sampled and as STT processing is performed on the sampled audio data. Displaying the text representation 616 can provide the user with confirmation that the digital assistant is correctly processing the user's request.
[0145] In block 526 of process 500, the user intent corresponding to the user utterance can be determined. The user intent can be determined by performing natural language processing on the text expression in block 522. Specifically, the text expression can be processed using a natural language processing module (e.g., natural language processing module 432) to derive the user intent. For example, referring to Figure 6F, from the text expression 616 corresponding to "Find romantic comedies starring Reese Witherspoon," it can be determined that the user intent is to request a search for media items that have the genre of romantic comedy and the actor Reese Witherspoon. In some embodiments, block 526 may further include using a natural language processing module to generate a structured query that expresses the determined user intent. In this embodiment, "Find romantic comedies starring Reese Witherspoon," a structured query can be generated that expresses a media search query for media items that have the genre of romantic comedy and the actor Reese Witherspoon.
[0146] In some embodiments, natural language processing for determining user intent can be biased towards media-related user intent. Specifically, a natural language processing module can be trained to identify media-related words and phrases (e.g., media title, media genre, actor, MPAA movie rating label, etc.) that trigger media-related nodes within an ontology. For example, a natural language processing module can identify the phrase "Jurassic Park" in a text representation as a movie title and, as a result, trigger a "media search" node within an ontology associated with the feasible intent of searching for media items. In some embodiments, bias can be implemented by limiting the nodes within an ontology to a predetermined set of media-related nodes. For example, the set of media-related nodes could be nodes associated with applications for media devices. Furthermore, in some embodiments, bias can be implemented by weighting media-related user intent candidates more heavily than non-media-related user intent candidates.
[0147] In some embodiments, user intent can be obtained from a separate device (e.g., DA server 106). Specifically, audio data can be transmitted to a separate device for natural language processing. In these embodiments, the media device can instruct the separate device (e.g., via data transmitted to the separate device along with the sampled audio data) that the sampled audio data is associated with a media application. This instruction can then bias the natural language processing towards media-related user intent.
[0148] In block 528 of process 500, a determination can be made as to whether the sampled audio data contains a user request. This determination can be made from the determined user intent in block 526. If the user intent includes a user request to perform a task, the sampled audio data can be determined to contain the user request. Conversely, if the user intent does not include a user request to perform a task, the sampled audio data can be determined not to contain the user request. Furthermore, in some embodiments, if the user intent is undeterminable from the text representation in block 526, or if the text representation is undeterminable from the sampled audio data in block 522, the sampled audio data can be determined not to contain the user request. Block 530 can be executed in accordance with the determination that the audio data does not contain the user request.
[0149] In block 530 of process 500, a request for clarification of the user's intent can be displayed on the display unit. In one embodiment, the request for clarification may be a request to the user to repeat the user request. In another embodiment, the request for clarification may be a statement that the digital assistant cannot understand the user's statement. In yet another embodiment, an error message can be displayed to indicate that the user's intent could not be determined. Furthermore, in some embodiments, a response may not be provided if it is determined that the voice data does not contain the user request.
[0150] Referring to Figure 5C, block 532 can be executed according to the determination in block 528 that the sampled audio data contains the user request. In block 532 of process 500, a task can be executed that satisfies the user request at least partially. For example, executing a task in block 526 may include executing one or more tasks defined in the generated structured query in block 526. One or more tasks may be executed using the digital assistant's task flow processing module (e.g., task flow processing module 436). In some embodiments, a task may include changing the state or settings of an application on a media device. More specifically, a task may include, for example, selecting or playing a requested media item, opening or closing a requested application, or navigating within a displayed user interface in a requested manner. In some embodiments, in block 532, a task may be executed without outputting any utterances related to the task from the media device. Therefore, in these embodiments, the user may provide a request to the digital assistant in the form of an utterance, but the digital assistant may not provide a response to the user in the form of an utterance. Rather, the digital assistant may simply respond visually by displaying the results on the display unit. This may be desirable in order to maintain a shared experience of consuming media content.
[0151] In other embodiments, a task may include retrieving and displaying requested information. Specifically, executing a task in block 532 may include executing one or more of blocks 534-536. In block 534 of process 500, results that at least partially satisfy the user request may be obtained. The results may be obtained from an external service (e.g., external service 120). In one embodiment, the user request may be a request to execute a media search query, such as "Find romantic comedies starring Reese Witherspoon." In this embodiment, block 534 may include executing the requested media search (e.g., using a media-related database of an external service) and obtaining media items that have the genre of romantic comedy and the actor Reese Witherspoon. In other embodiments, the user request may include requests for other types of information such as weather, sports, and stock prices, and each of these pieces of information may be obtained in block 534.
[0152] In block 536 of process 500, a second user interface can be displayed on the display unit. The second user interface may include a portion of the results obtained in block 534. For example, as shown in Figure 6G, a second user interface 618 may be displayed on the display unit. The second user interface 618 may include media items 622 that satisfy the user request, "Find a romantic comedy starring Reese Witherspoon." In this embodiment, media items 622 may include media items such as "Legally Blonde," "Legally Blonde 2," "Hot Pursuit," and "This Means War." The second user interface 618 may further include a text header 620 that describes the results obtained. The text header 620 may rephrase a portion of the user request to give the impression that the user's request has been directly addressed. This provides a more pleasant and interactive experience between the user and the digital assistant. In this embodiment shown in Figure 6G, the media items 622 are organized in a single column across the second user interface 618. Please note that in other embodiments, the arrangement and presentation of media item 622 may differ.
[0153] The second user interface 618 may further include a cursor 624 for navigating and selecting media items 622 within the second user interface 618. The cursor's position can be indicated by visually highlighting the media item to which the cursor is located relative to other media items. For example, in this example, the media item 623 to which the cursor 624 is located may be made larger and drawn with a thicker outline compared to other media items displayed within the second user interface 618.
[0154] In some embodiments, at least a portion of the displayed content can remain visible while the second user interface is displayed. For example, as shown in Figure 6G, the second user interface 618 may be a small pane displayed at the bottom of the display unit, while the media content 602 continues to play on the media device and remain displayed on the display unit above the second user interface 618. The second user interface 618 can be superimposed on the playing media content 602. In this embodiment, the display area of the second user interface 618 on the display unit may be smaller than the display area of the media content 602 on the display unit. This may be desirable to reduce the intrusiveness of the results displayed by the digital assistant while the user is consuming the media content. In other embodiments, it should be noted that the display area of the second user interface may differ from the display area of the displayed content. Furthermore, when the second user interface 618 is displayed, as indicated by the solid font for "Media Playing" in Figure 6G, the brightness of the media content 602 can be returned to normal (e.g., the brightness in Figure 6A before detecting user input). This can help indicate to the user that the interaction with the digital assistant is complete. Therefore, the user can continue to consume media content 602 while viewing the requested result (e.g., media item 622).
[0155] In embodiments where media items retrieved from a media search are displayed on a second user interface, the number of media items displayed can be limited. This may be desirable to allow the user to focus on the most relevant results and to prevent the user from being overwhelmed by the number of results when making a selection. In these embodiments, block 532 may further include determining whether the number of media items in the retrieved results is less than or equal to a predetermined number (e.g., 30, 28, or 25). According to the determination that the number of media items in the retrieved results is less than or equal to the predetermined number, all media items in the retrieved results can be included in the second user interface. According to the determination that the number of media items in the retrieved results is greater than the predetermined number, only a predetermined number of media items in the retrieved results can be included in the second user interface.
[0156] Furthermore, in some embodiments, only the media items in the retrieved results that are most relevant to the media search request can be displayed in the second user interface. Specifically, each media item in the retrieved results can be associated with a relevance score for the media search request. The displayed media item can have the highest relevance score among the retrieved results. Furthermore, the media items in the second user interface can be arranged according to their relevance scores. For example, referring to Figure 6G, media items with higher relevance scores are more likely to be positioned closer to one side of the second user interface 618 (e.g., the side closer to the cursor 624), while media items with lower relevance scores are more likely to be positioned closer to the opposite side of the second user interface 618 (e.g., the side further from the cursor 624). In addition, each media item in the retrieved results can be associated with a popularity rating. The popularity rating can be based on ratings from film critics (e.g., Rotten Tomatoes ratings) or on the number of users who selected the media item for playback. In some embodiments, media items 622 can be arranged within the second user interface 618 based on their popularity ranking. For example, media items with higher popularity rankings are more likely to be located on one side of the second user interface 618, while media items with lower popularity rankings are more likely to be located close to the opposite side of the second user interface 618.
[0157] As indicated by the different flows following block 532 in Figure 5C (e.g., D, E, F, and G), after block 532, one of blocks 538, 542, 550, or 570 in Figures 5D, 5E, 5F, or 5I can be executed, respectively. Blocks 538, 542, 550, or 570 can be executed while the second user interface is displayed in block 536. In some embodiments, process 500 may alternatively include a decision step after block 536 to determine the appropriate flow to execute (e.g., D, E, F, or G). Specifically, user input can be detected after block 536, and a determination can be made as to whether the detected user input corresponds to a second user input (e.g., block 538), a third user input (e.g., block 542), a fourth user input (e.g., block 550), or a sixth user input (e.g., block 570). For example, based on the determination that user input corresponds to the third user input in block 542, one or more of blocks 544 to 546 can be executed. A similar decision step may be included after block 546.
[0158] Referring to block 538 of process 500 and Figure 5D, a second user input can be detected. As described above, the second user input can be detected while the second user interface is displayed on the display unit. The second user input can be detected on the remote control of the media device. For example, the second user input may include a first predetermined motion pattern on the touch-sensitive surface of the remote control. In one embodiment, the first predetermined motion pattern may include a continuous contact motion in a first direction from a first contact point to a second contact point on the touch-sensitive surface. When the remote control is being held in the intended manner, the first direction can be downward or toward the user. It should be noted that other input formats for the second user input may also be contemplated. In response to the detection of the second user input, block 540 can be executed.
[0159] In block 540 of process 500, the second user interface can be closed, thereby preventing it from being displayed. For example, referring to Figure 6G, the second user interface 618 disappears in response to the detection of a second user input. In this embodiment, closing the second user interface 618 allows the media content 602 to be displayed on the full screen of the display unit. For example, if the display of the second user interface 618 is stopped, the media content 602 can be displayed as shown in Figure 6A.
[0160] Referring to block 542 of process 500 and Figure 5E, a third user input can be detected. The third user input can be detected while the second user interface is displayed on the display unit. The third user input can be detected on the remote control device of the media device. For example, the third user input may include a second predetermined motion pattern on the touch-sensitive surface of the remote control device. The second predetermined motion pattern may include a continuous contact motion in a second direction from a third contact point to a fourth contact point on the touch-sensitive surface. The second direction may be opposite to the first direction. Specifically, when the remote control device is being held in the intended manner, the second direction may be upward or away from the user. Depending on the detection of the third user input, one or more of blocks 544 to 546 can be executed. In some embodiments, as shown in Figure 6G, the second user interface 618 may include a graphic indicator 621 (e.g., an arrow) to indicate to the user that the second user interface 618 can be expanded by providing a third user input. Furthermore, the graphic indicator 621 may indicate to the user a second direction associated with a second predetermined motion pattern on a touch-sensitive surface for the third user input.
[0161] In block 544 of process 500, a second result can be obtained. The obtained second result may be similar to, but not identical to, the result obtained in block 534. In some embodiments, the obtained second result can satisfy a user requirement at least partially. For example, the obtained second result may share one or more characteristics, parameters, or attributes of the result obtained in block 534. In embodiments shown in Figures 6F to 6G, block 544 may include executing one or more additional media search queries related to the media search query executed in block 534. For example, one or more additional media search queries may include searching for media items that belong to the romantic comedy genre, or searching for media items starring Reese Witherspoon. Therefore, the obtained second result may include media items that are romantic comedies (e.g., media item 634), and / or media items starring Reese Witherspoon (e.g., media item 636).
[0162] In some embodiments, the obtained second result may be based on a previous user request received before detecting user input in block 504. Specifically, the obtained second result may include one or more characteristics or parameters of the previous user request. For example, the previous user request may be "Please tell me about movies released in the last five years." In this embodiment, the obtained second result may include a media item which is a romantic comedy movie starring Reese Witherspoon released in the last five years.
[0163] Furthermore, in some embodiments, block 544 may include retrieving a second result that is contextually relevant to the item that the second user interface is focused on when a third user input is detected. For example, referring to Figure 6G, when a third user input is detected, the cursor 624 may be positioned on a media item 623 in the second user interface 618. The media item 623 may be, for example, the movie "Legally Blonde". In this embodiment, the retrieved second result may share one or more characteristics, attributes, or parameters associated with the media item "Legally Blonde". Specifically, the retrieved second result may include media items such as "Legally Blonde" that feature a female protagonist who attends law school or has a professional occupation.
[0164] In block 546 of process 500, a third user interface can be displayed on the display unit. Specifically, the display of the second user interface in block 536 can be replaced with the display of the third user interface in block 546. In some embodiments, the second user interface can be expanded to become the third user interface in response to the detection of a third user input. The third user interface can occupy at least a majority of the display area of the display unit. The third user interface can include a portion of the results obtained in block 534. Furthermore, the third user interface can include a portion of the second results obtained in block 544.
[0165] In one embodiment, as shown in Figure 6H, the third user interface 626 can substantially occupy the entire display area of the display unit. In this embodiment, the previous display of media content 602 and the second user interface 618 can be superseded by the display of the third user interface 626. In response to the detection of a third user input, playback of the media content can be paused on the media device. This may be desirable to prevent the user from missing any portion of the media content 602 while browsing media items within the third user interface 626.
[0166] A third user interface 626 may include a media item 622 that satisfies the user request, "Find a romantic comedy starring Reese Witherspoon." Furthermore, the third user interface 626 may include a media item 632 that satisfies the same user request at least partially. Media item 632 may include multiple sets of media items, each corresponding to different characteristics, attributes, or parameters. In this embodiment, media item 632 may include media item 634, which is a romantic comedy, and media item 636, which stars Reese Witherspoon. Each set of media items may be labeled with a text header (e.g., text headers 628, 630). The text header may describe one or more attributes or parameters associated with each set of media items. Furthermore, each text header may be an exemplary user statement that, when provided to the digital assistant by the user, can cause the digital assistant to retrieve a similar set of media items. For example, by referring to text header 628, the digital assistant can, upon receiving the user utterance "romantic comedy" from the user, retrieve and display a media item (e.g., media item 634) that is a romantic comedy.
[0167] In the embodiment shown in Figure 6H, media item 622 is based on the initial user request, “Find a romantic comedy starring Reese Witherspoon,” but it should be noted that in other embodiments, media item 632 may be based on other factors such as media selection history, media search history, the order in which previous media searches were received, relationships between media-related attributes, popularity of the media item, and similar factors.
[0168] In embodiments where the user request is a media search request, the retrieved second result can be based on the number of media items in the retrieved result of block 534. Specifically, in response to detecting a third user input, a determination can be made as to whether the number of media items in the retrieved result is less than or equal to a predetermined number. Following the determination that the number of media items in the retrieved result is less than or equal to a predetermined number, the retrieved second result may include media items different from those in the second user interface. The retrieved second result can at least partially satisfy the media search request performed in block 534. At the same time, the retrieved second result can be broader than the retrieved result and can be associated with fewer parameters than all of the parameters limited in the media search request performed in block 534. This may be desirable to provide the user with a broader set of results and more choices to select from.
[0169] In some embodiments, a determination can be made as to whether a media search request contains more than one search attribute or parameter, based on the determination that the number of media items in the retrieved results of block 534 is less than or equal to a predetermined number. Based on the determination that a media search request contains more than one search attribute or parameter, the retrieved second result may contain media items associated with more than one search attribute or parameter. Furthermore, the media items in the retrieved second result can be organized in the third user interface according to more than one search attribute or parameter.
[0170] In the embodiments shown in Figures 6F to 6H, the media search request "Find romantic comedies starring Reese Witherspoon" can be determined to contain more than one search attribute or parameter (e.g., "romantic comedy" and "Reese Witherspoon"). Following the determination that the media search request contains more than one search attribute or parameter, the second result obtained may include media item 634 associated with the search parameter "romantic comedy" and media item 636 associated with the search parameter "movies starring Reese Witherspoon". As shown in Figure 6H, media item 634 can be organized under the category "romantic comedy" and media item 636 can be organized under the category "Reese Witherspoon".
[0171] In some embodiments, the third user interface may include a first and a second part of the acquired results, depending on the determination that the number of media items in the acquired results of block 534 is greater than a predetermined number. The first part of the acquired results may include a predetermined number of media items (e.g., those with the highest relevance scores). The second part of the acquired results may differ from the first part and may include more media items than the first part. Furthermore, it can be determined whether the media items in the acquired results include more than one media type (e.g., movies, TV shows, music, applications, games, etc.). Depending on the determination that the media items in the acquired results include more than one media type, the media items in the second part of the acquired results may be sorted according to their media type.
[0172] In the embodiment shown in Figure 6I, the results obtained in block 534 may include a media item that is a romantic comedy starring Reese Witherspoon. If it is determined that the number of media items in the obtained results is greater than a predetermined number, the first part of the obtained results (media item 622) and the second part of the obtained results (media item 638) can be displayed in the third user interface 626. If it is determined that the obtained results include more than one media type (e.g., movies and TV shows), media item 638 can be sorted according to the media type. Specifically, media item 640 can be sorted under the category "movies," and media item 642 can be sorted under the category "TV shows." Furthermore, in some embodiments, each set of media items corresponding to each media type (e.g., movies, TV shows) (e.g., media items 640, 642) can be sorted within each set of media items according to the most common genre, actor / director, or release date. In other embodiments, it should be noted that if it is determined that a media item in the acquired results is associated with one or more media attributes or parameters, the media items in the second part of the acquired results may be organized according to media attributes or parameters (rather than media type).
[0173] In some embodiments, user input representing a scroll command (e.g., a fourth user input described later in block 550) can be detected. Upon receiving user input representing a scroll command, the enlarged user interface (or more specifically, items within the enlarged user interface) can be scrolled. While scrolling, a determination can be made as to whether the enlarged user interface has scrolled beyond a predetermined position within the enlarged user interface. Upon determining that the enlarged user interface has scrolled beyond a predetermined position within the enlarged user interface, media items from the third part of the retrieved results can be displayed on the enlarged user interface. The media items from the third part can be organized according to one or more media content providers associated with the media items from the third part (e.g., iTunes, Netflix, HuluPlus, HBO, etc.). In other embodiments, it should be noted that upon determining that the enlarged user interface has scrolled beyond a predetermined position within the enlarged user interface, other media items can be retrieved. For example, popular media items or media items related to the retrieved results can be retrieved.
[0174] As indicated by different flows (e.g., B, F, G, and H) that proceed from block 546 in Figure 5E, blocks 550, 558, 566, or 570 in Figures 5F, 5G, 5H, or 5I can be executed after block 532. Specifically, in some embodiments, blocks 550, 560, 564, or 570 can be executed while the third user interface is displayed in block 546.
[0175] In block 550 of process 500, and referring to Figure 5F, a fourth user input can be detected. The fourth user input can be detected while a second user interface (e.g., second user interface 618) or a third user interface (e.g., third user interface 626) is displayed on the display unit. In some embodiments, the fourth user input can be detected on the remote control device of the media device. The fourth user input can indicate a direction (e.g., up, down, left, right) on the display unit. For example, the fourth user input can be a touch motion from a first position on the touch-sensing surface of the remote control device to a second position on the touch-sensing surface to the right of the first position. Therefore, the touch motion can correspond to a rightward direction on the display unit. In response to the detection of the fourth user input, block 552 can be executed.
[0176] In block 552 of process 500, the focus of the second or third user interface can be switched from the first item to the second item on the second or third user interface. The second item can be positioned in the direction described above (for example, the same direction corresponding to the fourth user input) relative to the first item. For example, in Figure 6G, the focus of the second user interface 618 can be on media item 623 because the cursor 624 is positioned on media item 623. In response to the detection of a fourth user input corresponding to the right direction on the display unit, the focus of the second user interface 618 can be switched from media item 623 in Figure 6G to media item 625 in Figure 6J, which is located to the right of media item 623. Specifically, the position of the cursor 624 can be changed from media item 623 to media item 625. In another embodiment, referring to Figure 6H, the focus of the third user interface 626 can be on media item 623. In response to the detection of a fourth user input corresponding to a downward direction on the display unit, the focus of the third user interface 626 can be switched from media item 623 in Figure 6H to media item 627 in Figure 6K, which is positioned below media item 623. Specifically, the position of the cursor 624 can be changed from media item 623 to media item 627.
[0177] In block 554 of process 500, the selection of one or more media items can be received via a second or third user interface. For example, referring to Figure 6J, the selection of media item 625 can be received via the second user interface 618 by detecting user input corresponding to the user selection while the cursor 624 is positioned over media item 625. Similarly, referring to Figure 6K, the selection of media item 627 can be received via the third user interface 626 by detecting user input corresponding to the user selection while the cursor 624 is positioned over media item 627. In response to receiving the selection of one or more media items, block 556 can be executed.
[0178] In block 556 of process 500, media content associated with the selected media item can be displayed on the display unit. In some embodiments, the media content may be a movie, video, television program, animation, or similar that is playing on or streaming through a media device. In some embodiments, the media content may be a video game, ebook, application, or program running on a media device. Furthermore, in some embodiments, the media content may be information related to the media item. The information may be product information describing various characteristics of the selected media item (e.g., synopsis, cast, director, author, release date, rating, duration, etc.).
[0179] In block 558 of process 500, and referring to Figure 5G, a fifth user input can be detected. In some embodiments, the fifth user input can be detected while the third user interface (e.g., the third user interface 626) is displayed. In these embodiments, the fifth user input can be detected while the focus of the third user interface is on a media item in the top row of the third user interface (e.g., one of the media items 622 in the third user interface 626 in Figure 6H). In other embodiments, the fifth user input can be detected while the first user interface is displayed. In these embodiments, the fifth user input can be detected while any one of blocks 508 to 514 is being executed. In some embodiments, the fifth user input can be detected on a remote control device of the media device. The fifth user input can be similar to or identical to the third user input. For example, the fifth user input can include a continuous touch motion in a second direction on the touch-sensitive surface (e.g., a swipe-up touch motion). In other embodiments, the fifth user input may be the activation of an affordance. The affordance may be associated with a virtual keyboard interface or a type-search interface. In response to the detection of the fifth user input, one or more of blocks 560 to 564 may be executed.
[0180] In block 560 of process 500, a search field configured to receive typed search input can be displayed. For example, a search field 644 can be displayed on the displayed unit as shown in Figure 6L. In some embodiments, the search field can be configured to receive typed search queries. Typed search queries can be media-related search queries, such as searching for media items. In some embodiments, the search field can be configured to perform a media-related search based on a text string match between the text entered via search field 644 and the stored text associated with the media item. Furthermore, in some embodiments, the digital assistant may not be configured to receive input via search field 644. This may encourage the user to interact with the digital assistant via a speech interface rather than a typed interface, promoting a more user-friendly interface between the media device and the user. It should be noted that in some embodiments, the search field may be displayed from the beginning within a second user interface (e.g., second user interface 618) or a third user interface (e.g., third user interface 626). In these embodiments, it may not be necessary to execute block 566.
[0181] In block 562 of process 500, a virtual keyboard interface can be displayed on the display unit. For example, a virtual keyboard interface 646 can be displayed as shown in Figure 6L. The virtual keyboard interface 646 can be configured so that user input received through the virtual keyboard interface 646 results in text being entered into a search field. In some embodiments, the virtual keyboard interface cannot be used to interact with a digital assistant.
[0182] In block 564 of process 500, the focus of the user interface can be switched to the search field. For example, referring to Figure 6L, the search field 644 can be highlighted in block 568. Furthermore, the text input cursor can be positioned within the search field 644. In some embodiments, text prompting the user to type a search can be displayed within the search field. As shown in Figure 6L, the text 648 includes the prompt "Please type your search".
[0183] In block 566 of process 500, and referring to Figure 5H, a seventh user input can be detected. In some embodiments, the seventh user input can be detected while displaying a third user interface (e.g., third user interface 626). In some embodiments, the seventh user input may include pressing a button on a remote control of the electronic device. The button may be, for example, a menu button for navigating to the main menu user interface of the electronic device. In other embodiments, it should be recognized that the seventh user input may include other forms of user input. Depending on the detection of the seventh user input, block 568 can be executed.
[0184] In block 568 of process 500, the display of the third user interface on the display unit can be stopped. Specifically, the seventh user input can cause the third user interface to be closed. In some embodiments, the seventh user input can cause the main menu user interface menu to be displayed instead of the third user interface. Alternatively, in embodiments where media content (e.g., media content 602) is displayed before the third user interface (e.g., third user interface 626) is displayed, and playback of the media content on the electronic device is paused simultaneously with the display of the third user interface (e.g., paused in response to the detection of the third user input), playback of the media content on the electronic device can be resumed in response to the detection of the seventh user input. Thus, the media content can be displayed in response to the detection of the seventh user input.
[0185] In block 570 of process 500, and with reference to Figure 5I, a sixth user input can be detected. As shown in Figure 6M, the sixth user input can be detected while the third user interface 626 is displayed. However, in other embodiments, the sixth user input can alternatively be detected while the second user interface (e.g., the second user interface 618) is displayed. When the sixth user input is detected, the second or third user interface may include a portion of the result that at least partially satisfies the user request. The sixth user input may include an input for calling the digital assistant of an electronic device. Specifically, the sixth user input may be similar to or identical to the second input form of user input described above with reference to block 516. For example, the sixth user input may include pressing a specific button on the remote control device of a media device and holding the button for a longer period than a predetermined time (e.g., a long press). Depending on the detection of the sixth user input, one or more of blocks 572 to 592 can be performed.
[0186] In block 572 of process 500, second audio data can be sampled. Block 572 can be the same as or identical to block 518 described above. Specifically, the sampled second audio data may include a second user utterance from the user. The second user utterance can express a second user request directed to the digital assistant. In some embodiments, the second user request may be a request to perform a second task. For example, referring to Figure 6M, the sampled second audio data may include a second user utterance, "only those featuring Luke Wilson." In this embodiment, the second user utterance can express a second user request to narrow down the previous media search to include only media items in which Luke Wilson appears as an actor. In this embodiment, the second user utterance is in natural language form. Furthermore, the second user request may be incomplete. In this case, the second user utterance does not clearly specify all the information necessary to define the user request. For example, the second user utterance does not clearly specify what "the ones" refers to. In other embodiments, the second user request may be a request to play a media item or to provide specific information (e.g., weather, stock prices, sports, etc.).
[0187] It should be noted that in some embodiments, blocks 520-526 described above can be similarly executed for a sixth user input. Specifically, as shown in Figure 6M, an active visual indicator 614 can be displayed on the display unit simultaneously with the detection of the sixth user input. A second text representation 650 of the second user utterance can be determined (e.g., using the STT processing module 430) and displayed on the display unit. Based on the second text representation, a second user intent corresponding to the second user utterance can be determined (e.g., using the natural language processing module 432). In some embodiments, as shown in Figure 6M, the content displayed on the display unit when the sixth user input is detected can be faded or its brightness reduced in response to the detection of the sixth user input. This can help to highlight the active visual indicator 614 and the second text representation 650.
[0188] In block 574 of process 500, a determination can be made as to whether the sampled second audio data contains the second user request. Block 574 can be the same as or identical to block 528 described above. Specifically, the determination in block 574 can be made based on the second user intent determined from the second text representation of the second user utterance. If the determination is made that the second audio data does not contain the user request, block 576 can be executed. Alternatively, if the determination is made that the second audio data contains the second user request, one or more of blocks 578 to 592 can be executed.
[0189] In block 576 of process 500, a request for clarification of the user's intent can be displayed on the display unit. Block 576 may be the same as or identical to block 530 described above.
[0190] In block 578 of process 500, a determination can be made as to whether a second user request is a request to refine the results of a user request. In some embodiments, the determination can be made from a second user intent corresponding to the second user utterance. Specifically, the second user request can be determined to be a request to refine the results of a user request based on an explicit instruction to refine the results of a user request identified within the second user utterance. For example, referring to Figure 6M, the second text expression 650 can be parsed during natural language processing to determine whether the second user utterance contains a predetermined word or phrase corresponding to an explicit intent to refine media search results. Examples of words or phrases corresponding to an explicit intent to refine media search results include "just", "only", "filter by", and similar. Therefore, based on the word "just" in the second text expression 650, the second user request can be determined to be a request to narrow down the media search results associated with the user request, "Find romantic comedies starring Reese Witherspoon." It should be noted that other techniques can also be employed to determine whether the second user request is a request to narrow down the results of the user request. In accordance with the determination that the second user request is a request to narrow down the results of the user request, one or more of blocks 580-582 can be performed.
[0191] In block 580 of process 500, a subset of results that at least partially satisfy the user request can be obtained. In some embodiments, the subset of results can be obtained by filtering the existing results according to additional parameters limited in a second user request. For example, the results obtained in block 534 (e.g., including media item 622) can be filled so that media items in which Luke Wilson appears as an actor are identified. In other embodiments, a new media search query can be executed that combines the requirements of the user request and the requirements of a second user request. For example, the new media search query could be a search query for media items in the romantic comedy genre, as well as media items in which Reese Witherspoon and Luke Wilson appear as actors. In this embodiment, the new media search query could yield media items such as "Legally Blonde" and "Legally Blonde 2".
[0192] In embodiments where a sixth user input is detected while a third user interface is displayed, additional results related to the user request and / or the second user request can be obtained. These additional results may include media items having one or more attributes or parameters mentioned within the user request and / or the second user request. Furthermore, the additional results may not include all attributes or parameters mentioned within the user request and the second user request. For example, referring to the embodiments shown in Figures 6H and 6M, the additional results may include media items having at least one (but not all) of the following attributes or parameters: romantic comedy, Reese Witherspoon, and Luke Wilson. Additional results may be desirable to provide the user with a broader set of results and more choices to select from. Furthermore, the additional results may be relevant results that are likely to be of interest to the user.
[0193] In block 582, a subset of results can be displayed on the display unit. For example, as shown in Figure 6N, the subset of results may include media item 652, which may include movies such as "Legally Blonde" and "Legally Blonde 2". In this embodiment, media item 652 is displayed in the top row of the third user interface 626. A text header 656 may describe the attributes or parameters associated with the displayed media item 652. Specifically, the text header 656 may include a paraphrase of the user's intent associated with the second user statement. In embodiments where a sixth user input is detected while the second user interface (e.g., the second user interface 618 shown in Figure 6G) is displayed, media item 652 may instead be displayed within the second user interface. In these embodiments, media item 652 may be displayed as a single column across the second user interface. It should be noted that there are various ways in which media item 652 can be displayed within the second or third user interface.
[0194] In embodiments where a sixth user input is detected while a third user interface is being displayed, additional results related to the user request and / or the second user request may be displayed within the third user interface. For example, referring to Figure 6N, the additional results may include media item 654 having one or more parameters stated in the user request and / or the second user request. Specifically, media item 654 may include media item 658, a romantic comedy starring Luke Wilson, and media item 660, starring Luke Wilson and published in the last decade. Each set of media items (e.g., media items 658, 660) may be labeled with a text header (e.g., text headers 662, 664). The text header may describe one or more parameters associated with each set of media items. The text header may be in natural language format. Furthermore, each text header may be an exemplary user statement that, when provided by the user to the digital assistant, can cause the digital assistant to retrieve a similar set of media items. For example, by referring to text header 662, the digital assistant can, in response to receiving the user statement "A romantic comedy starring Luke Wilson" from the user, retrieve and display a media item (e.g., media item 658) that is a romantic comedy starring Luke Wilson.
[0195] Referring again to block 578, we can determine that the second user request is not a request to narrow down the results of the user request. Such a determination can be made based on the fact that there is no explicit instruction to narrow down the results of the user request at all in the second user utterance. For example, when parsing the second text representation of the second user utterance during natural language processing, a given word or phrase corresponding to an explicit intention to narrow down media search results may not be identified. This may be because the second user request is a request unrelated to the previous user request (e.g., a new request). For example, the second user request could be "Find horror movies," which is an unrelated request to the previous user request, "Find romantic comedies starring Reese Witherspoon." Alternatively, the second user request may contain ambiguous words that can be interpreted as either a request to narrow down the results of the previous user request or a new request unrelated to the previous user request. For example, referring to Figure 6P, the second user utterance could be "Luke Wilson." This can be interpreted as either a request to refine the results of a previous user request (e.g., a request to include only media items in which Luke Wilson appears as an actor), or a new request unrelated to the previous user request (e.g., a new media search for media items in which Luke Wilson appears as an actor). In these embodiments, it can be determined that the second user request is not a request to refine the results of a user request. Depending on the determination that the second user request is a request to refine the results of a user request, one of more of blocks 584-592 can be executed.
[0196] In block 584 of process 500, a second task can be performed that at least partially satisfies a second user request. Block 584 can be the same as block 532 described above, except that the second task of block 584 may differ from the task of block 532. Block 584 may contain one or more of blocks 586 to 588.
[0197] In block 586 of process 500, a third result can be obtained that at least partially satisfies the second user request. Block 586 can be similar to block 534 described above. Referring to the embodiment shown in Figure 6P, the second user statement "Luke Wilson" can be interpreted as a request to perform a new media search query to identify media items in which Luke Wilson appears as an actor. Therefore, in this embodiment, block 586 can include performing the requested media search and obtaining media items in which Luke Wilson appears as an actor. In other embodiments, the user request may include requests for other types of information (e.g., weather, sports, stock prices, etc.), and it should be noted that each type of information can be obtained in block 586.
[0198] In block 588 of process 500, a portion of the third result can be displayed on the display unit. For example, referring to Figure 6Q, the third result, including media items 670 in which Luke Wilson appears as an actor (e.g., movies such as "Playing It Cool," "The Skeleton Twins," and "You Kill Me"), can be displayed within the third user interface 626. In this embodiment, media items 670 can be displayed in the top row of the third user interface 626. A text header 678 can describe the attributes associated with the displayed media item 670. Specifically, the text header 678 may include a paraphrase of the determined user intent associated with the second user statement. In embodiments where a sixth user input is detected while the second user interface (e.g., the second user interface 618 shown in Figure 6G) is displayed, media items 670 can be displayed within the second user interface. In these embodiments, media items 670 can be displayed in a single column across the second user interface. It should be noted that in other embodiments, the organization or configuration of media items 670 within the second or third user interface may differ.
[0199] In block 590 of process 500, a fourth result can be obtained that at least partially satisfies the user requirement and / or the second user requirement. Specifically, the fourth result may include a media item having one or more attributes or parameters limited within the user requirement and / or the second user requirement. Referring to the embodiments shown in Figures 6P and 6Q, the fourth result may include a media item having one or more of the following attributes or parameters: romantic comedy, Reese Witherspoon, and Luke Wilson. For example, the fourth result may include a media item 676 that has the genre of romantic comedy and stars Luke Wilson. Obtaining a fourth result may be desirable to provide the user with a broader set of results and, therefore, more choices to select from. Furthermore, the fourth result may be associated with an alternative predicted user intent derived from the second user requirement and one or more previous user requirements to increase the likelihood that the user's actual intent is satisfied. This can help improve the accuracy and relevance of the results returned to the user, thereby improving the user experience.
[0200] In some embodiments, at least a portion of the fourth result may include a media item having all the parameters limited within the user request and the second user request. For example, the fourth result may include a media item 674 having the genre of romantic comedy and starring Reese Witherspoon and Luke Wilson. Media item 674 may be associated with an alternative intent to refine the results of the previous user request using the second user request. If the user actually intended the second request to be a request to refine the results to be retrieved, then retrieving media item 674 may be desirable to increase the likelihood that the user's actual intent will be satisfied.
[0201] In some embodiments, a portion of the fourth result may be based on the focus of the user interface at the time the sixth user input is detected. Specifically, the focus of the user interface may be on one or more items of the third user interface when the sixth user input is detected. In this embodiment, a portion of the fourth result may be contextually related to one or more items on which the user interface is focused. For example, referring to Figure 6K, the cursor 624 may be positioned on media item 627, and therefore the focus of the third user interface 626 may be on media item 627. In this embodiment, attributes or parameters associated with media item 627 may be used to obtain a portion of the fourth result. For example, the category "Movies starring Reese Witherspoon" associated with media item 627 may be used to obtain a portion of the fourth result, and the obtained portion may include media items starring both Reese Witherspoon and Luke Wilson. In another embodiment, media item 627 may be an adventure movie, and therefore, a portion of the fourth result may include media items that are adventure movies starring Luke Wilson.
[0202] In block 592 of process 500, a portion of the fourth result can be displayed. In embodiments where a sixth user input is detected while the third user interface is being displayed, a portion of the fourth result can be displayed within the third user interface. For example, as shown in Figure 6Q, the portion of the fourth result may include media item 672, which is displayed in a subsequent row following media item 670. Media item 672 may be associated with one or more attributes or parameters (e.g., romantic comedy, Reese Witherspoon, and Luke Wilson) limited within the second user request and / or user request. For example, media item 672 may include media item 676, which is a romantic comedy starring Luke Wilson, and media item 674, which is a romantic comedy starring both Reese Witherspoon and Luke Wilson. Each set of media items (e.g., media items 674, 676) may be labeled with a text header (e.g., text headers 680, 682). A text header can describe one or more attributes or parameters associated with each set of media items. The text header may be in natural language format. Furthermore, each text header may be an exemplary user statement, provided by the user to the digital assistant, which can then cause the digital assistant to retrieve similar sets of media items having similar attributes.
[0203] As described above, the second user statement, "Luke Wilson," can be associated with two possible user intentions: a first user intention to perform a new media search, or a second user intention to refine the results of a previous user request. Displayed media item 670 can satisfy the first user intention, and displayed media item 674 can satisfy the second user intention. In this embodiment, media items 670 and 674 are displayed in the top two rows. In this way, the results for the two most likely user intentions associated with the second user request (e.g., a new search or refinement of a previous search) can be displayed prominently (e.g., in the top two rows) within the third user interface 626. This may be desirable to minimize scrolling or browsing by the user within the third user interface until they find the desired media item to consume. It should be noted that there are various ways in which media items 670 and 674 can be displayed prominently within the third user interface 626 to minimize scrolling and browsing.
[0204] Figures 7A to 7C illustrate a process 700 for operating a digital assistant in a media system, according to various embodiments. Process 700 can be performed using one or more electronic devices that implement the digital assistant. For example, process 700 can be performed using one or more of the aforementioned system 100, media system 128, media device 104, user device 122, or digital assistant system 400. Figures 8A to 8W show screenshots displayed on a display unit by the media device at various stages of process 700, according to various embodiments. Process 700 will be described below with simultaneous reference to Figures 7A to 7C and Figures 8A to 8W. It should be understood that some operations within process 700 can be combined, the order of some operations can be changed, and some operations can be omitted.
[0205] In block 702 of process 700, content can be displayed on a display unit (e.g., display unit 126). Block 702 may be the same as or identical to block 502 described above. Referring to Figure 8A, the displayed content may include media content 802 (e.g., movies, videos, television programs, video games, etc.) being played on a media device (e.g., media device 104). In other embodiments, the displayed content may include other content, such as content associated with an application running on the media device, or a user interface for interacting with the media device's digital assistant. Specifically, the displayed content may include a main menu user interface, or a user interface with objects or results previously requested by the user.
[0206] In block 704 of process 700, user input can be detected. Block 704 may be the same as or identical to block 504 described above. User input can be used to invoke the media device's digital assistant. In some embodiments, user input can be detected while the content of block 702 is being displayed. User input can be detected on a remote control device of the media device (e.g., remote control device 124). For example, user input may correspond to the second input format described in block 516 of process 500. Specifically, user input in block 704 may include pressing a specific button on the media device's remote control device and holding the button for a longer period than a predetermined time (e.g., long press). Depending on the detection of user input, one or more of blocks 706 to 746 can be executed.
[0207] In block 706 of process 700, audio data can be sampled. Block 706 may be the same as or identical to block 518 described above. The sampled audio data may include user utterances. User utterances may represent user requests directed to the digital assistant of a media device. For example, referring to the embodiment shown in Figure 8A, the sampled audio data may include the user utterance, "What time is it in Paris?" User utterances may be in the form of unstructured natural language. In some embodiments, the request expressed by the user utterance may be incompletely specified. In this case, the user utterance (e.g., "Play this") may lack or not explicitly limit the information necessary to perform the request. In other embodiments, the user utterance may not be an explicit request, but rather an indirect question or statement (e.g., "What did he say?") in which the request is inferred. Furthermore, as will be described in more detail below in block 712, the user utterance may include one or more ambiguous terms.
[0208] In block 708 of process 700, the text representation of a user utterance in the sampled audio data can be determined. Block 708 can be the same as or identical to block 522 described above. Specifically, the text representation can be determined by performing STT processing on the user utterance in the sampled audio data. For example, referring to Figure 8A, the text representation 804 "What time is it in Paris?" is determined from the user utterance in the sampled audio data and can be displayed on the display unit. As shown in the figure, the text representation 804 can be superimposed on the media content 802 while the media content 802 continues to play on the media device.
[0209] In some embodiments, the STT processing used to determine the text representation can be biased towards media-related text results. In addition, or alternatively, the text representation can be based on previous user utterances received by the media device before the audio data was sampled. Furthermore, in some embodiments, the text representation can be based on the time when the previous user utterances were received before the audio data was sampled. In embodiments where the text representation is obtained from a separate device (e.g., DA server 106), the media device can instruct the separate device that the sampled audio data is associated with a media application, and this instruction can bias the STT processing on the separate device towards media-related text results.
[0210] In block 710 of process 700, the user intent corresponding to the user utterance can be determined. Block 710 can be the same as block 526 described above. Specifically, the text representation in block 708 can be processed using natural language processing (e.g., by the natural language processing module 432) to derive the user intent. For example, referring to Figure 8A, from the text representation 804 "What time is it in Paris?", the user intent can be determined to be to request the time in a location named "Paris". The natural language processing used to determine the user intent can be biased towards media-related user intent. In embodiments where the user intent is obtained from a separate device (e.g., DA server 106), the media device can instruct the separate device that the sampled audio data is associated with a media application, and this instruction can bias the natural language processing on the separate device towards media-related user intent.
[0211] In some embodiments, user intent can be determined based on prosodic information derived from user utterances in sampled audio data. Specifically, prosodic information (e.g., tone, rhythm, volume, stress, intonation, tempo, etc.) can be derived from user utterances to determine the user's attitude, mood, emotion, or feeling. Then, the user intent can be determined from the user's attitude, mood, emotion, or feeling. For example, the sampled audio data may include the user utterance "What did he say?". In this embodiment, based on the high volume and stress detected in the user utterance, it can be determined that the user is irritated or frustrated. Based on the user utterance and the determined user emotion, the user intent can be determined to include a request to increase the volume of audio associated with media content being played on the media device.
[0212] As shown in Figure 7A, block 710 can contain one or more of blocks 712-718. Specifically, if it is determined that there is a high probability of two or more user intents, and the natural language processing module cannot narrow down the two or more user intents to a single user intent, then one or more of blocks 712-718 can be executed. For example, this situation may occur when a user utterance contains ambiguous terms that cannot be cleared based on available contextual information.
[0213] In block 712 of process 700, a determination can be made as to whether a user utterance (or a textual representation of a user utterance) contains ambiguous terminology. This determination can be made during natural language processing (for example, using the natural language processing module 432) to determine the user intent. Ambiguous terminology can be a word or phrase that has more than one possible interpretation. For example, referring to Figure 8A, the term "Paris" in the user utterance "What time is it in Paris?" can be interpreted as the city of Paris in France, or the city of Paris in Texas, USA. Therefore, the term "Paris" in the user utterance can be determined to be ambiguous terminology.
[0214] In some embodiments, contextual information can be retrieved (e.g., by a digital assistant) to remove ambiguity from potentially ambiguous terms. If ambiguation is successful, it can be determined that the user utterance does not contain ambiguous terms. For example, media content 802 is a film (e.g., "Ratatouille") set in Paris, France, and therefore it can be determined that the user is more likely to be referring to Paris, France than to Paris, Texas. In this embodiment, the term "Paris" can be successfully deambiguated to refer to Paris, France, and therefore it can be determined that the user utterance does not contain ambiguous terms.
[0215] In another embodiment, the user utterance could be "Play this." In this embodiment, the user utterance does not explicitly limit the specific media item to be played, and therefore the term "this," when interpreted alone, could be an ambiguous term that could refer to any media item available to the media device. The ambiguity of the term can be removed using contextual information displayed on the display unit by the media device. For example, the digital assistant can determine whether the focus of the displayed user interface is on a media item. Following the determination that the focus of the user interface is on a media item, the digital assistant can remove the ambiguity of the term "this" and determine that the term refers to the media item that the displayed user interface is focused on. Based on this determination, in block 712, it can be determined that the user utterance does not contain an ambiguous term. Therefore, the user intent can be determined to be a request to play the media item that the displayed user interface is focused on.
[0216] In embodiments where ambiguity of terminology cannot be removed, block 712 can determine that the user utterance contains ambiguous terminology. Depending on the determination that the user utterance contains ambiguous terminology, one or more of blocks 714 to 718 can be executed. In block 714 of process 700, two or more candidate user intents can be obtained based on the ambiguous terminology. The two or more candidate user intents may be the most likely candidate user intents determined from the user utterance for which ambiguity cannot be removed. Referring to the embodiment shown in Figure 8A, the two or more candidate user intents may include a first candidate user intent of requesting a time in Paris, France, and a second candidate user intent of requesting a time in Paris, Texas.
[0217] In block 716 of process 700, two or more user intent candidates can be displayed on the display unit for user selection. For example, referring to Figure 8B, a first user intent candidate 810 and a second user intent candidate 808 can be displayed. Furthermore, a text prompt 806 can be provided to prompt the user to indicate the actual user intent corresponding to the user statement by selecting between the first user intent candidate 810 and the second user intent candidate 808. The text prompt 806, the first user intent candidate 810, and the second user intent candidate 808 can be overlaid on media content 802.
[0218] In block 716 of process 700, a user selection of one of two or more candidate user intents can be received. In some embodiments, the user selection can be received via the selection of an affordance corresponding to one of the candidate user intents. Specifically, as shown in Figure 8B, each of two or more candidate user intents (810, 808) can be displayed on the display unit as a selectable affordance. The media device can receive input from the user to change the display focus to one of the affordances (e.g., via the media device's remote control). Subsequently, the user selection of the candidate user intent corresponding to that affordance can be received (e.g., via the media device's remote control). For example, as shown in Figure 8B, the media device can receive user input to move the cursor 812 over the affordance corresponding to the first candidate user intent 810 (e.g., Paris, France). Subsequently, the user selection of the first candidate user intent 810 can be received.
[0219] In other embodiments, user selection can be received via voice interaction with a digital assistant. For example, a second user input can be detected while displaying two or more candidate user intents. The second user input may be similar to or identical to the user input in block 704. Specifically, the second user input may be an input for calling the digital assistant (e.g., pressing a specific button on a remote control device of a media device and holding the button for a longer period than a predetermined time). In response to the detection of the second user input, second voice data may be sampled. The second voice data may include a second user utterance that represents a user selection of one of two or more interpretations. For example, referring to Figure 8C, the second voice data may include a second user utterance, "Paris, France." As shown in the figure, a text representation 814 of the second user utterance, "Paris, France," can be displayed on the display unit. In this embodiment, the second user utterance, "Paris, France," can represent a user selection of the first candidate user intent 810 (e.g., Paris, France). Based on the second user statement "Paris, France", it can be determined that candidate 810 for the first user intent is the actual user intent corresponding to the user statement "What time is it in Paris?". Therefore, in block 710, it can be determined that the user intent is to request the time in Paris, France. Once the user intent is determined based on the received user selection, one or more of blocks 720-746 can be executed.
[0220] In some embodiments, blocks 710-718 can be executed without outputting utterances from the media device. Specifically, the text prompt 806 and user intent candidates 808, 810 can be displayed without outputting utterances associated with two or more user intent candidates 808, 810. Thus, user input can be received in the form of utterances, while the output of the digital assistant can be presented to the user visually (and not in audio form) on the display unit. This may be desirable to maintain a shared experience associated with consuming media content, thereby improving the user experience of the media device.
[0221] Referring again to block 712, depending on whether it has been determined that the user statement does not contain ambiguous terminology, one or more of blocks 720-746 can be executed. In block 720 of process 700, a determination can be made as to whether the user intent corresponds to one of several core capabilities associated with the media device. For example, the media device may be associated with several predetermined core capabilities, such as searching for media items, playing media items, and providing information related to media items, weather, stock prices, and sports. If the user intent involves performing a task related to one of several predetermined core capabilities, then it can be determined that the user intent corresponds to one of several predetermined core capabilities. For example, if the user intent is a request for a media item starring Reese Witherspoon, then it can be determined that the user intent corresponds to one of several predetermined core capabilities. Depending on whether it has been determined that the user intent corresponds to one of several core capabilities associated with the electronic device, one or more of blocks 724-746 can be executed.
[0222] Conversely, if a user intent involves performing a task other than one of several predetermined core capabilities, it can be determined that the user intent does not correspond to one of those predetermined core capabilities. For example, if the user intent is a request for map guidance, it can be determined that the user intent does not correspond to one of the predetermined core capabilities. Block 722 can be executed in response to the determination that the user intent does not correspond to one of the multiple core capabilities associated with the electronic device.
[0223] In block 722 of process 700, a second electronic device (e.g., device 122) can be made to satisfy the user intent at least partially. Specifically, the second electronic device can be made to perform a task to assist in satisfying the user intent. In one embodiment, it can be determined that the media device is not configured to satisfy the user intent of requesting map guidance, and therefore the user intent can be transmitted to the second electronic device to satisfy the user intent. In this embodiment, the second user device can perform the task of displaying the requested map guidance. In other embodiments, information other than the user intent can be transmitted to the second electronic device in order to cause the second electronic device to perform a task to assist in satisfying the user intent. For example, the digital assistant of the media device can determine a task flow or structured query to satisfy the user intent (e.g., using a natural language processing module 432 or a task flow processing module 436), and the task flow or structured query can be transmitted to the second electronic device. The second electronic device can then perform the task flow or structured query to assist in satisfying the user intent.
[0224] As will become clear in the description provided below, the level of intrusion associated with satisfying user intent can be based on the nature of the user intent. In some cases, the task associated with satisfying user intent can be performed without displaying any additional responses or outputs on the display (e.g., block 726). In other cases, to satisfy user intent, only text responses (e.g., without corresponding visual or audio outputs) are provided (e.g., block 732). In yet other cases, to satisfy user intent, a user interface with relevant results can be displayed (e.g., blocks 738, 742, or 746). The user interface may occupy more or less than half of the display unit. Thus, process 700 can intelligently adjust the level of intrusion of the output depending on the nature of the user intent. This allows for convenient access to the digital assistant's services while reducing undesirable interruptions during the consumption of media content. This improves the overall user experience.
[0225] In block 724 of process 700, a determination can be made as to whether the user intent includes a request to adjust the state or settings of an application on a media device. Depending on whether it is determined that the user intent includes a request to adjust the state or settings of an application on a media device, block 726 can be executed. In block 726 of process 700, the state or settings of the application can be adjusted to satisfy the user intent.
[0226] In some embodiments, a state or setting may be associated with displayed media content being played on a media device. For example, a request to adjust the state or setting of an application may include a request to control the playback of media content by the media device. Specifically, it may include a request to pause, resume, restart, stop, rewind, or fast-forward the playback of displayed media content on the media device. It may also include a request to jump forward or backward within the media content (e.g., by a specified period) to play a desired portion of the media content. Furthermore, a request to adjust the state or setting of an application may include a request to turn on / off subtitles or closed captions (e.g., in a specified language) associated with the displayed media content, increase / decrease the volume of audio associated with the displayed media content, mute / unmute audio associated with the displayed media content, or accelerate / decelerate the playback speed of the displayed media content.
[0227] Figures 8E to 8F illustrate an example of a user intent that includes a request to control the playback of media content by a media device. In this embodiment, a digital assistant can be invoked (e.g., in block 704) while media content 802 is playing. The media content can initially be displayed without subtitles. (e.g., in block 706) Sampled audio data may include the user statement "Please turn on English subtitles." As shown in Figure 8E, a text representation 816 of the user statement can be displayed on the display unit. Based on this user statement, in block 710, it can be determined that the user intent includes a request to turn on the display of English subtitles for media content 802. Furthermore, in block 724, it can be determined that this user intent is a request to adjust the state or setting of an application on an electronic device. In response to this determination, English subtitles for media content 802 can be turned on. As represented by label 817 in Figure 8F, the display of English subtitles associated with media content 802 can be initiated to satisfy the user intent.
[0228] In another exemplary embodiment shown in Figures 8G to 8H, a user utterance in the sampled audio data may be a natural language expression indicating that the user did not hear a portion of the audio associated with media content. Specifically, the user utterance may be "What did he say?" as shown by the text expression 820 in Figure 8G. In this embodiment, it can be determined (for example, in block 710) that the user intent includes a request to replay the portion of the media content corresponding to the portion of audio the user did not hear. It can also be determined that the user intent includes a request to turn on closed captions to help with the difficulty in hearing the audio associated with the media content. Furthermore, based on the prosodic information in the user utterance, it can be determined that the user is frustrated or irritated, and therefore, based on the user's emotions, it can be determined that the user intent includes a request to increase the volume of the audio associated with the media content. In block 724, it can be determined that these user intents are requests to adjust the state or settings of the application on the electronic device. Depending on this determination, the media content can be rewound by a predetermined period (e.g., 15 seconds) to a previous portion of the media content (as represented, for example, by label 822 in Figure 8H), and playback of the media content can be restarted from that previous portion. In addition, before restarting playback of the media content from the previous portion, closed captions can be turned on (e.g., as represented, by label 824 in Figure 8H). Furthermore, before restarting playback of the media content from the previous portion, the volume of the audio associated with the media content can be increased.
[0229] It should be understood that closed captions or subtitles associated with media content can be obtained from a service provider (e.g., a cable provider or media subscription service). However, in embodiments where closed captions or subtitles are not available from a service provider, the media device can generate closed captions or subtitles to assist with hearing difficulties associated with the audio associated with the media content. For example, before receiving user utterances in sampled audio data and while the media content is playing, utterances in the audio associated with the media content can be continuously converted to text (e.g., using the STT processing module 730) and stored in association with the media content. In response to a user request to replay a previous portion of the media content that the user could not hear, while the previous portion of the media content is being replayed, the text corresponding to the previously played portion can be retrieved and displayed.
[0230] In some embodiments, states or settings associated with displayed media content can be adjusted without displaying an additional user interface for performing the adjustment, or without providing any text or graphics to indicate that the state or setting has been adjusted. For example, in the embodiments illustrated in Figures 8E–8H, subtitles (or closed captions) can be simply turned on without explicitly displaying text such as "Subtitles turned on," or without displaying a user interface to control the display of subtitles. Furthermore, states or settings can be adjusted without outputting any audio associated with satisfying user intent. For example, in Figures 8E–8H, subtitles (or closed captions) can be turned on without outputting any audio (e.g., utterances or non-verbal speech signals) to confirm that subtitles have been turned on. Therefore, the requested action can be simply performed without any additional auditory or visual interruption of the media content. In this way, process 700 can minimize interruption to the user's consumption of media content while providing convenient access to the digital assistant's services, thereby improving the user experience.
[0231] In other embodiments, a request to adjust the state or settings of an application on a media device may include a request to navigate within the media device's user interface (e.g., a second user interface 818, a third user interface 826, or a main menu user interface). In one embodiment, a request to navigate within a user interface may include a request to switch the focus of the user interface from a first object (e.g., a first media item) to a second object within the user interface (e.g., a second media item). Figures 8I to 8K illustrate one example of such a request. As shown in Figure 8I, the displayed content may include a third user interface 826 having multiple media items organized into various categories (e.g., "Romantic Comedies," "Romantic Comedies Starring Reese Witherspoon," and "Movies by Luke Wilson"). As indicated by the position of the cursor 828, the focus of the third user interface 826 may be on a first media item 830 under the "Romantic Comedies" category. The second media item 832 may have the title "Legally Blonde" and may be categorized under "Romantic Comedy starring Reese Witherspoon." As shown by the text representation 834 in Figure 8J, the user utterance in the sampled audio data (for example, in block 706) may be "Go to Legally Blonde." Based on this user utterance, it can be determined (for example, in block 710) that the user intent is a request to switch the focus of the third user interface 826 from the first media item 830 to the second media item 832 having the title "Legally Blonde."Depending on whether the user intent is determined to be a request to adjust the state or settings of an application on an electronic device (for example, in block 724), the focus of the third user interface 826 can be switched from the first media item 830 to the second media item 832. For example, as shown in Figure 8K, the position of the cursor 828 can be changed from the first media item 830 to the second media item 832.
[0232] In another embodiment, a request to navigate within a user interface may include a request to change the focus of the user interface to a specific category of results displayed within the user interface. For example, Figure 8I includes media items associated with the categories “Romantic Comedy,” “Romantic Comedy Starring Reese Witherspoon,” and “Movies by Luke Wilson.” Instead of “Go to Legally Blonde,” the user utterance in the sampled audio data could instead be “Jump to Romantic Comedy Starring Reese Witherspoon.” Based on this user utterance, it can be determined (e.g., in block 710) that “Romantic Comedy Starring Reese Witherspoon” limits the category of media items displayed within the third user interface 826, and therefore, the user intent can be determined to be a request to change the focus of the user interface to one or more media items associated with that category. Depending on whether the user intent is determined to be a request to adjust the state or settings of an application on an electronic device (for example, in block 724), the focus of the third user interface 826 can be shifted to one or more media items associated with the category. For example, as shown in Figure 8K, the position of the cursor 828 can be shifted to a second media item 832 associated with "Romantic Comedy starring Reese Witherspoon".
[0233] In yet another embodiment, a request to navigate within the user interface of a media device may include a request to select an object within the user interface. Selecting an object may cause an action associated with that object to be performed. For example, as shown in Figure 8K, the cursor 828 is positioned over a second media item 832 with the title "Legally Blonde". As shown in Figure 8L, a digital assistant can be invoked (e.g., in block 704), and a user utterance in the sampled audio data (e.g., in block 706) may be "Play this" (e.g., displayed as text representation 836). Based on this user utterance, it can be determined (e.g., in block 710) that the user intent is a request to play a specific media item. In this embodiment, the user utterance does not explicitly limit or specify a particular media item to be played. Specifically, the word "this" is ambiguous. However, the digital assistant can obtain contextual information to eliminate the ambiguity of the user intent. For example, it can be determined that the focus of the third user interface 826 was over the second media item 832 at the time the audio data was sampled. Based on this determination, the second media item 832 can be identified as the media item to be played. Depending on whether the user's intention to play the second media item 832 is determined to be a request to adjust the state or settings of the electronic device's application (for example, in block 724), an action can be taken to assist in playing the second media item 832. For example, preview information about the second media item 832 can be displayed on the display unit. The preview information may include, for example, a plot summary, a list of cast members, public data, user ratings, and the like.In addition, or alternatively, a second media item 832 can be played on the media device, and the media content associated with the second media item 832 can be displayed on the display unit (for example, represented by the text 838 “Playing Legally Blonde” in Figure 8M). It should be noted that in other embodiments, the media item to be selected can be explicitly identified. For example, instead of “Play this,” the user statement could specifically say “Play Legally Blonde,” and similar actions could be taken to assist in playing the second media item 832.
[0234] In yet another embodiment, a request to navigate within the user interface of a media device may include a request to browse a specific user interface or application on the media device. For example, a user statement in the sampled audio data may be, "Go to the actor's page." In this case, the user intent includes a request to display the user interface associated with browsing for media items related to a particular actor. In yet another embodiment, a user statement in the sampled audio data may be, "Take me to the homepage." In this case, the user intent includes a request to display the main menu user interface of the media device. In yet another embodiment, a request to navigate within the user interface of a media device may include a request to launch an application on the electronic device. For example, a user statement in the sampled audio data may be, "Go to the iTunes Store." In this case, the user intent includes a request to launch the iTunes Store application. It should be noted that other requests to adjust the state or settings of applications on the media device may also be contemplated.
[0235] Referring again to block 724, it can be determined that the user intent does not include a request to adjust the state or settings of an application on an electronic device. For example, the user intent may instead be a request to present information related to one or more media items. Depending on such a determination, one or more of blocks 728 to 746 may be executed. In block 728 of process 700, a determination can be made as to whether the user intent is one of a plurality of predetermined request types. In some embodiments, the plurality of predetermined request types may be requests associated with text-only responses. More specifically, the plurality of predetermined request types may be requests for information predetermined to require a text-only response. This is in contrast to requests predetermined to require a response containing a media object (e.g., an image, an animation object, a video, etc.). In some embodiments, the plurality of predetermined request types may include requests for the current time in a particular location (e.g., "What time is it in Paris?"), requests for a joke (e.g., "Tell me a funny joke."), or requests for information about media content currently playing on an electronic device (e.g., "When was this movie released?"). Depending on whether the user intent is determined to be one of several predetermined request types, one or more of blocks 730 to 732 can be executed.
[0236] In block 730 of process 700, results that at least partially satisfy the user intent can be obtained. For example, results can be obtained from an external service (e.g., external service 120) by executing a task flow. In block 732 of process 700, the results obtained in block 730 can be displayed in text format on a display unit. Furthermore, the results can be displayed in text format without displaying any corresponding graphics or media-related items corresponding to the results.
[0237] Figures 8M to 8P illustrate exemplary embodiments of blocks 728 to 732. As shown in Figure 8M, the movie "Legally Blonde" may be initially playing on the media device and displayed on the display unit. While "Legally Blonde" is playing, the digital assistant may be invoked (e.g., in block 704), and the user utterance in the sampled audio data may be "Who is the lead actress?". For example, as shown in Figure 8N, a text representation 840 of the user utterance may be displayed on the display unit. Based on this user utterance, it can be determined (e.g., in block 710) that the user intent includes a request to identify the lead actress of a particular media item. The user intent may be ambiguous because the user utterance does not specify any particular media item. However, based on the fact that the movie "Legally Blonde" was displayed at the time the audio data was sampled, it can be determined that the media item associated with the user intent is "Legally Blonde". In this embodiment, it can be determined (e.g., in block 728) that the user intent is one of several predetermined request types. Specifically, it can be determined that a text-only response can be provided to satisfy the user intent of identifying the lead actress in Legally Blonde. Depending on the determination that the user intent is one of several predetermined request types, a search can be performed in the media-related database (for example, in block 730) to retrieve "Reese Witherspoon" as the lead actress in the movie "Legally Blonde". As shown in Figure 8P, the text-only result 842 "Reese Witherspoon" can be displayed on the display unit to satisfy the user intent. The text-only result 842 can be overlaid on the displayed media content of "Legally Blonde". Furthermore, the media content of "Legally Blonde" can continue to play while the text-only result 842 is displayed.By displaying text-only results (for example, without displaying graphic results or additional user interfaces to satisfy user intent), user intent can be satisfied in an unobtrusive manner, minimizing interruptions to user consumption of media content. Simultaneously, users are provided with access to the digital assistant's services, which can be desirable for improving the user experience.
[0238] Referring again to block 728, it can be determined that the user intent is not one of several predetermined request types. Specifically, the user intent may be a predetermined request type that requires results other than text to be satisfied. For example, the user intent may be a request to execute a media search query and display the media items corresponding to the media search query. In other embodiments, the user intent may be a request for information other than media items. For example, the user intent may be a request for information associated with a sports team (e.g., "How did the LaLaportes do in their last game?"), an athlete (e.g., "How tall is LeBron James?"), a stock price (e.g., "What was the closing price of the Dow Jones yesterday?"), or weather (e.g., "What is the weather forecast for Paris, France next week?"). Depending on whether the user intent is determined to be not one of several predetermined request types, one or more of blocks 734-746 may be executed.
[0239] In block 734 of process 700, a second result can be obtained that at least partially satisfies the user intent. Block 734 may be the same as or identical to block 534 described above. In one embodiment, the user intent may include a request to execute a media search query. In this embodiment, the media search query can be executed in block 734 and a second result can be obtained. Specifically, the second result may include media items corresponding to the media search query.
[0240] In some embodiments, the user intent does not have to be a media search query. For example, the user intent could be a request for a weather forecast for Paris, France (e.g., "How about a weather forecast for Paris, France?"). In this embodiment, the second result obtained in block 734 could include a 7-day weather forecast for Paris, France. The second result could include non-media data that at least partially satisfies the user intent. Specifically, the 7-day weather forecast for Paris, France could include text data (e.g., date, temperature, and a brief description of the weather conditions) and graphical images (e.g., images of sunny, cloudy, windy, or rainy). Furthermore, in some embodiments, the scope of the user intent can be expanded in block 710 to include a request for a media item that at least partially satisfies the user intent. In these embodiments, the second result obtained in block 734 could further include one or more media items having media content that at least partially satisfies the user intent. For example, in block 734, a media search query can be executed for weather forecasts in Paris, France for a relevant period, and one or more media items related to weather forecasts in Paris, France can be retrieved. One or more media items may include, for example, video clips from weather channels presenting weather forecasts in Paris, France. In these embodiments, non-media data and / or one or more media items can be displayed within the user interface on the displayed unit (for example, in blocks 738, 742, or 746 described below).
[0241] In block 736 of process 700, a determination can be made as to whether the displayed content includes media content being played on the electronic device. In some embodiments, it can be determined that the displayed content does not include media content being played on the electronic device. For example, the displayed content may instead include a user interface, such as a main menu user interface or a third user interface (e.g., a third user interface 826). The third user interface may occupy at least a majority of the display area of the display unit. Furthermore, the third user interface may include previous results related to previous user requests received before user input was detected in block 704. Block 738 can be executed according to the determination that the displayed content does not include media content.
[0242] In block 738 of process 700, a portion of the second result can be displayed within a third user interface on the display unit. In embodiments where the displayed content already includes a third user interface when user input is received in block 704, a display of a previous result related to a previous user request can be replaced with a portion of the display of the second result within the third user interface. In embodiments where the displayed content does not include a third user interface when user input is received in block 704 (for example, the displayed content includes a main menu user interface), a third user interface can be displayed, and the second result can be included within the displayed third user interface.
[0243] In some embodiments, a determination can be made as to whether the second result includes a predetermined type of result. The predetermined type of result may be associated with a display area that occupies less than half of the display area of the display unit. The predetermined type of result may include, for example, results related to stock prices or weather. It should be noted that in other embodiments, the predetermined type of result may be different. Depending on whether it is determined that the second result includes a predetermined type of result, a portion of the second result may be displayed in a second user interface on the display unit. The second user interface may occupy less than half of the display area of the display unit. In these embodiments, even if it is determined in block 736 that the displayed content does not include media content, a portion of the second result may still be displayed in the second user interface.
[0244] Figures 8Q to 8S illustrate exemplary embodiments of blocks 734 to 738. In this embodiment, as shown in Figure 8Q, the displayed content may first include a third user interface 826. The third user interface 826 may include previous results from previous user requests. Specifically, the third user interface 826 includes a media item 844 from a previously requested media search query. As shown in Figure 8R, while the third user interface 826 is displayed, the digital assistant can be invoked (e.g., in block 704). The user utterance in the sampled audio data may include "Tell me about movies starring Luke Wilson." A text representation 846 of the user utterance can be displayed on the display unit. In this embodiment, the user intent can be determined (e.g., in block 710) to be a request to execute a media search query for movies starring Luke Wilson. The media search query can be executed (e.g., in block 734) to obtain a second result. Specifically, the second result may include a media item 848 corresponding to a movie starring Luke Wilson. Furthermore, additional results related to user intent or previous user intent (e.g., media item 850) can be obtained. These additional results can be obtained in a similar manner to the second results described in block 544.
[0245] In this embodiment shown in Figures 8Q to 8S, the displayed content includes only the third user interface 826, and therefore it can be determined (for example, in block 736) that the displayed content does not include media content being played on the electronic device. In accordance with this determination, the second result can be displayed within the third user interface 826. Specifically, as shown in Figure 8S, the display of media item 844 within the third user interface 826 can be replaced by the display of media item 848 within the third user interface 826. Furthermore, media item 850 can be displayed within the third user interface 826.
[0246] As shown in this embodiment, the second result can be presented in the third user interface only after it has been determined that the media content is not displayed on the display unit. This allows for the display of a wider range of results in a larger area, increasing the likelihood that the user's actual intent will be satisfied. At the same time, by ensuring that the media content is not displayed on the display unit before the second result is presented in the third user interface, the user's consumption of the media content is not interrupted.
[0247] Referring again to block 736, the displayed content may include media content that is currently playing on the media device. In these embodiments, it is possible to determine that the displayed content includes media content that is currently playing on the media device. Based on this determination, one or more of blocks 740 to 746 can be executed.
[0248] In block 740 of process 700, a determination can be made as to whether the media content being played can be paused. Examples of media content that can be paused include on-demand media items such as on-demand movies and television programs. Examples of media content that cannot be paused include broadcast or streaming service media programs, as well as live media programs (e.g., sporting events, concerts, etc.). Therefore, on-demand media items do not necessarily have to include broadcast or live programs. In accordance with the determination in block 740 that the media content being played cannot be paused, block 742 can be executed. In block 742 of process 700, a second user interface having a portion of the second result can be displayed on the display unit. Block 742 can be the same as block 536 described above. The second user interface can be displayed while the media content is being displayed. The display area occupied by the second user interface on the display unit can be smaller than the display area occupied by the media content on the display unit. In accordance with the determination that the media content being played can be paused, one or more of blocks 744 to 746 can be executed. In block 744 of process 700, the media content being played can be paused on the media device. In block 746 of process 700, a third user interface having a portion of the second result can be displayed. The third user interface can be displayed while the media content is paused.
[0249] Figures 8T to 8W illustrate exemplary embodiments of blocks 740 to 746. As shown in Figure 8T, media content 802 playing on a media device can be displayed on the display unit. While media content 802 is displayed, the digital assistant can be activated (e.g., in block 704). The user utterance in the sampled audio data may be, "Tell me about movies starring Luke Wilson." A text representation 846 of the user utterance can be displayed on the display unit. As described above, the user intent can be determined (e.g., in block 710) to be a request to retrieve a media item for a movie starring Luke Wilson. A corresponding media search query can be executed (e.g., in block 734) to retrieve a second result. The second result may include a media item 848 for a movie starring Luke Wilson. In embodiments where it is determined (e.g., in block 744) that media content 802 cannot be paused, the media item 848 can be displayed in a second user interface 818 while media content 802 continues to be displayed on the display unit (e.g., Figure 8U). Displaying the media item 848 within the second user interface 818 may be desirable to ensure that the media content 802 remains continuously available for user consumption while the media item 848 is displayed to satisfy the user's intent. This prevents the user from missing any portion of the media content 802 that cannot be paused or resumed. Alternatively, in embodiments where it is determined that the media content 802 can be paused (e.g., in block 744), playback of the media content 802 on the media device can be paused, and the media item 848 can be displayed within the third user interface 826 on the display unit (e.g., Figure 8S).Displaying a third user interface 826 may be desirable to allow a wider range of media items (e.g., media item 850) associated with various alternative user intents to be displayed alongside the requested media item (e.g., media item 848), thereby increasing the likelihood that the user's actual intent will be satisfied. At the same time, the media content 802 is paused to ensure that the user does not miss any portion of the media content 802. By changing the user interface used to display media item 848 based on whether the media content 802 can be paused, it is possible to comprehensively achieve the user intent associated with the user's utterance while reducing interruptions to the user's consumption of the media content 802. This can enhance the overall user experience.
[0250] In some embodiments, as shown in FIG. 8V, the displayed content can include a second user interface 818 in addition to the media content 802 being played on the media device. In these embodiments, the second user interface 818 can include media items 852 related to previous user requests (e.g., a request for a romantic comedy starring Reese Witherspoon). While the media content 802 and the second user interface 818 are being displayed, a digital assistant can be invoked (e.g., at block 704). As shown in FIG. 8W, the sampled audio data can include the user utterance "Please tell me movies starring Luke Wilson". A text representation 846 of the user utterance can be displayed on the display unit. Based on this user utterance, it can be determined (e.g., at block 710) that the user intent is a request to obtain media items of movies starring Luke Wilson. A corresponding media search query can be executed (e.g., at block 734) to obtain a second result (e.g., media item 848). In these embodiments, the display of the media item 852 within the second user interface 818 can be replaced with the display of the media item 848 (e.g., FIG. 8U).
[0251] FIG. 9 shows a process 900 for interacting with a digital assistant of a media system according to various embodiments. The process 900 can be executed using one or more electronic devices that implement the digital assistant. For example, the process 900 can be executed using one or more of the system 100, media system 128, media device 104, user device 122, or digital assistant system 400 described above. It should be understood that some operations within the process 900 can be combined, the order of some operations can be changed, and some operations can be omitted.
[0252] In block 902 of process 900, content can be displayed on the display unit. Block 902 may be similar to or identical to block 502 described above. In some embodiments, the displayed content may include media content (e.g., movies, videos, television programs, video games, etc.). In addition, or alternatively, the displayed content may include a user interface. For example, the displayed content may include a first user interface having one or more exemplary natural language requests (e.g., as shown in Figures 6D to 6E). In other embodiments, the displayed content may include a third user interface (e.g., a third user interface 626) having results from previous user requests (e.g., previously requested media items). The third user interface may occupy at least a majority of the display area of the display unit.
[0253] In block 904 of process 900, user input can be detected while displaying the content of block 902. The user input may be similar to or identical to the fifth user input described in block 558. Specifically, the user input may be detected on a remote control device of the media device. For example, the user input may include a predetermined motion pattern on the touch-sensitive surface of the remote control device. In some embodiments, the user input may be detected via a second electronic device (e.g., device 122) distinct from the media device. The second electronic device may be configured to wirelessly control the media device. Depending on the detection of user input, one or more of blocks 906 to 914 may be executed.
[0254] In block 906 of process 900, a virtual keyboard interface (e.g., virtual keyboard interface 646) can be displayed on the display unit. Block 906 can be the same as or identical to block 562 described above. The virtual keyboard interface can be overlaid on at least a portion of the first user interface or the third user interface. Further, a search field (e.g., search field 644) can be displayed on the display unit. The virtual keyboard interface can be configured such that user input received via the virtual keyboard interface causes text entry into the search field.
[0255] In block 908 of process 900, selectable affordances can be displayed on a second electronic device (e.g., on touch screen 346 of device 122). The second electronic device can be a device different from the remote control device of the media device. Selection of an affordance can enable text input to be received by the media device via the keyboard of the second electronic device. For example, selection of an affordance can cause a virtual keyboard interface (e.g., one similar to virtual keyboard interface 646) to be displayed on the second electronic device. Input to the virtual keyboard interface of the second electronic device can cause corresponding text to be entered into the search field (e.g., search field 644).
[0256] In block 910 of process 900, text input can be received via the keyboard of a second electronic device (e.g., a virtual keyboard interface). Specifically, a user can input text via the keyboard of the second electronic device, and this text input can be transmitted to and received by a media device. The text input can represent a user request. For example, the text input could be "Jurassic Park," which can represent a request to perform a search for media items associated with the search string "Jurassic Park."
[0257] In block 912 of process 900, results that at least partially satisfy the user request can be obtained. For example, a media search can be performed using text input, and corresponding media items can be obtained. In a particular embodiment where the text input is "Jurassic Park", media items that have the title "Jurassic Park" or have actors or directors in common with the movie "Jurassic Park" can be obtained. In another embodiment where the text input is "Reese Witherspoon", media items in which Reese Witherspoon is an actress can be obtained.
[0258] In block 914 of process 900, a user interface can be displayed on the display unit. The user interface may include at least a portion of the results. For example, the user interface may include media items obtained as a result of a media search performed in block 912.
[0259] While certain blocks of processes 500, 700, and 900 have been described above as being executed by a device or system (e.g., media device 104, user device 122, or digital assistant system 400), it should be noted that in some embodiments, more than one device may be used to execute a block. For example, in a block where a determination is made, a first device (e.g., media device 104) may obtain the determination from a second device (e.g., server system 108). Similarly, in a block where content, objects, text, or a user interface is displayed, a first device (e.g., media device 104) may display the content, objects, text, or user interface on a second device (e.g., display unit 126). 5. Electronic Devices
[0260] Based on several embodiments, Figure 10 shows a functional block diagram of an electronic device 1000 configured according to the principles of various embodiments described, for example, to provide voice control of media playback and real-time updates of a virtual assistant's knowledge. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software to perform the principles of the various embodiments described. Those skilled in the art will understand that the functional blocks described in Figure 10 may be combined or separated into subblocks to perform the principles of the various embodiments described. Accordingly, the description herein supports, at its discretion, any possible combination or division of the functional blocks described herein, or further definitions.
[0261] As shown in Figure 10, the electronic device 1000 may include an input unit 1003 (e.g., remote control unit 124, or similar) configured to receive user input such as tactile input or gesture input; an audio input unit 1004 (e.g., microphone 272, or similar) configured to receive audio data; a speaker unit 106 (e.g., speaker 268, or similar) configured to output audio; and a communication unit 1007 (e.g., communication subsystem 224, or similar) configured to transmit and receive information from external devices over a network. In some embodiments, the electronic device 1000 may optionally include a display unit 1002 (e.g., display unit 126, or similar) configured to display media, interfaces, and other content. The electronic device 1000 may further include an input unit 1003, an audio input unit 1004, a speaker unit 1006, a communication unit 1007, and a processing unit 1008 optionally coupled to a display unit 1002. In some embodiments, the processing unit 1008 may include a display-enabled unit 1010, a detection unit 1012, a determination unit 1014, a sampling unit 1016, an output unit 1018, an execution unit 1020, an acquisition unit 1022, and a switching unit 1024.
[0262] According to some embodiments, the processing unit 1008 is configured to display content on a display unit (e.g., display unit 1002 or a separate display unit) (e.g., using a display-enabled unit 1010). The processing unit 1008 is further configured to detect user input (e.g., using a detection unit 1012). The processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the user input corresponds to a first input format. In accordance with the determination that the user input corresponds to a first input format, the processing unit 1008 is further configured to display a plurality of exemplary natural language requests on the display unit (e.g., using a display-enabled unit 1010). The plurality of exemplary natural language requests are contextually relevant to the displayed content, and receiving a user utterance corresponding to one of the plurality of exemplary natural language requests causes the digital assistant to perform the respective action.
[0263] In some embodiments, user input is detected on a remote control device of an electronic device. In some embodiments, a first form of input includes pressing a button on the remote control device and releasing the button within a predetermined period of time. In some embodiments, a plurality of exemplary natural language requests are displayed on a display unit via a first user interface, the first user interface being overlaid on the displayed content. In some embodiments, the displayed content includes media content, which continues to play while the plurality of exemplary natural language requests are displayed.
[0264] In some embodiments, the processing unit 1008 is further configured to display a visual indicator on the display unit (for example, using the display-enabled unit 1010) indicating that the digital assistant is not processing the voice input, in accordance with the determination that the user input corresponds to a first input format.
[0265] In some embodiments, when it is determined that user input corresponds to a first input format, a set of exemplary natural language requests are displayed on the display unit after a predetermined time. In some embodiments, each of the set of exemplary natural language requests is displayed separately at different times in a predetermined order.
[0266] In some embodiments, the processing unit 1008 is further configured to display multiple lists of exemplary natural language requests (for example, using a displayable unit 1010), each list being displayed alternately at different times.
[0267] In some embodiments, the processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the user input corresponds to a second input format, based on the determination that the user input does not correspond to a first input format. The processing unit 1008 is further configured to sample the audio data (e.g., using a sampling unit 1016 and an audio input unit 1004), based on the determination that the user input corresponds to a second input format. The processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the audio data contains a user request. The processing unit 1008 is further configured to execute a task (e.g., using an execution unit 1020) that at least partially satisfies the user request, based on the determination that the audio data contains a user request.
[0268] In some embodiments, the second input method includes pressing a button on a remote control device of an electronic device and holding the button down for a longer period than a predetermined time.
[0269] In some embodiments, the processing unit 1008 is further configured to display a request for clarification of user intent on the display unit (for example, using the display-enabled unit 1010) in accordance with the determination that the audio data does not contain a user request.
[0270] In some embodiments, the displayed content includes media content, which continues to play on the electronic device while audio data is being sampled and while the task is being performed.
[0271] In some embodiments, the processing unit 1008 is further configured to output audio associated with media content (for example, using the speaker unit 1006) (for example, using the output unit 1018). The processing unit 1008 is further configured to reduce the amplitude of the audio (for example, using the output unit 1018) according to a determination that the user input corresponds to a second input format.
[0272] In some embodiments, the task is performed without outputting any utterances related to the task from the electronic device. In some embodiments, the audio data is sampled while user input is being detected. In some embodiments, the audio data is sampled during a predetermined period after user input has been detected.
[0273] In some embodiments, audio data is sampled via a first microphone (e.g., audio input unit 1004) on a remote control unit of an electronic device. The processing unit 1008 is further configured to sample background audio data (e.g., using sampling unit 1016 and audio input unit 1004) via a second microphone (e.g., a second audio input unit of electronic device 1000) on a remote control unit while sampling the audio data. The processing unit 1008 is further configured to remove background noise in the audio data (e.g., using output unit 1018) using the background audio data.
[0274] In some embodiments, audio associated with the displayed content is output from the electronic device via an audio signal. The processing unit 1008 is further configured to remove background noise in the audio data using the audio signal (for example, using the output unit 1018).
[0275] In some embodiments, the processing unit 1008 is further configured to display, on the display unit, a visual cue (e.g., using the enabling display unit 1010) that prompts the user to provide a speech request in response to detecting a user input.
[0276] In some embodiments, the processing unit 1008 is further configured to obtain (e.g., using the acquisition unit 1022) results that at least partially satisfy a user request. The processing unit 1008 is further configured to display, on the display unit, a second user interface (e.g., using the enabling display unit 1010). The second user interface includes a portion of the results, and at least a portion of the content continues to be displayed while the second user interface is being displayed, and the display area of the second user interface on the display unit is smaller than the display area of at least a portion of the content on the display unit. In some embodiments, the second user interface is superimposed on the displayed content.
[0277] In some embodiments, the portion of the results includes one or more media items. The processing unit 1008 is further configured to receive, via the second user interface, a selection of a media item among the one or more media items (e.g., using the detection unit 1012). The processing unit 1008 is further configured to display, on the display unit, media content associated with the selected media item (e.g., using the enabling display unit 1010).
[0278] In some embodiments, the processing unit 1008 is further configured to detect a second user input (e.g., using the detection unit 1012) while the second user interface is being displayed. The processing unit 1008 is further configured to abort displaying the second user interface (e.g., using the enabling display unit 1010) in response to detecting the second user input.
[0279] In some embodiments, the second user input is detected on a remote control unit of an electronic device. The second user input includes a first predetermined motion pattern on the touch-sensitive surface of the remote control unit.
[0280] In some embodiments, the processing unit 1008 is further configured to detect a third user input (e.g., using a detection unit 1012) while displaying a second user interface. In response to detecting a third user input, the processing unit 1008 is further configured to replace the display of the second user interface with a display of the third user interface on the display unit (e.g., using a display-enabled unit 1010). The third user interface includes at least a portion of the result and occupies at least a majority of the display area of the display unit.
[0281] In some embodiments, the third user input is detected on a remote control unit of an electronic device, and the third user input includes a second predetermined motion pattern on the touch-sensitive surface of the remote control unit.
[0282] In some embodiments, the processing unit 1008 is further configured to acquire a second result different from the first result (for example, using an acquisition unit 1022) in response to detecting a third user input. The second result satisfies the user request at least partially, and the third user interface includes at least a portion of the second result.
[0283] In some embodiments, the second result is based on a user request received before the detection of user input. In some embodiments, the focus of the second user interface is on an item in the result portion while the third user input is detected, and the second result is contextually related to the item.
[0284] In some embodiments, the displayed content includes media content. The processing unit 1008 is further configured to pause playback of the media content on the electronic device (for example, using the execution unit 1020) in response to detecting a third user input.
[0285] In some embodiments, at least a portion of the results includes one or more media items. The processing unit 1008 is further configured to receive a selection of media items from among one or more media items via a third user interface (for example, using a detection unit 1012). The processing unit 1008 is further configured to display the media content associated with the media items on a display unit (for example, using a display-enabled unit 1010).
[0286] In some embodiments, the processing unit 1008 is further configured to detect a fourth user input associated with a direction on the display unit (e.g., using a detection unit 1012) while the third user interface is being displayed. In response to detecting the fourth user input, the processing unit 1008 is further configured to switch the focus of the third user interface from the first item to a second item on the third user interface (e.g., using a switching unit 1024). The second item is positioned in the direction described above relative to the first item.
[0287] In some embodiments, the processing unit 1008 is further configured to detect a fifth user input (e.g., using a detection unit 1012) while displaying a third user interface. The processing unit 1008 is further configured to display a search field (e.g., using a display-enabling unit 1010) in response to the detection of the fifth user input. The processing unit 1008 is further configured to display a virtual keyboard interface on the display unit (e.g., using a display-enabling unit 1010), and input received via the virtual keyboard interface results in text being entered into the search field.
[0288] In some embodiments, the processing unit 1008 is further configured to detect a sixth user input (e.g., using a detection unit 1012) while displaying a third user interface. The processing unit 1008 is further configured to sample second audio data (e.g., using a sampling unit 1016 and an audio input unit 1004) in response to detecting a sixth user input. The second audio data comprises a second user request. The processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the second user request is a request to narrow down the results of the user request. The processing unit 1008 is further configured to display a subset of the results (e.g., using a displayable unit 1010) via the third user interface in accordance with the determination that the second user request is a request to narrow down the results of the user request.
[0289] In some embodiments, a subset of the results is displayed at the top of the third user interface. The processing unit 1008 is further configured to obtain (e.g., using the acquisition unit 1018) a third result that at least partially satisfies the second user request, in accordance with the determination that the second user request is not a request to narrow down the results of the user request. The processing unit 1008 is further configured to display (e.g., using the displayable unit 101) a portion of the third result via the third user interface. In some embodiments, a portion of the third result is displayed at the top of the third user interface.
[0290] In some embodiments, the processing unit 1008 is further configured to acquire a fourth result (for example, using the acquisition unit 1022) that at least partially satisfies a user request or a second user request. The processing unit 1008 is further configured to display a portion of the fourth result (for example, using the displayable unit 1010) via a third user interface.
[0291] In some embodiments, the fourth result is displayed in the row after the top row of the third user interface.
[0292] In some embodiments, the focus of the third user interface is on one or more items of the third user interface while the sixth user input is detected, and the fourth result is contextually related to one or more items.
[0293] In some embodiments, the processing unit 1008 is further configured to detect a seventh user input (e.g., using a detection unit 1012) while displaying the third user interface. The processing unit 1008 is further configured to stop displaying the third user interface (e.g., using a display-enabled unit 1010) in response to detecting the seventh user input.
[0294] In some embodiments, the displayed content is media content, and playback of the media content on the electronic device is paused in response to the detection of a third user input. The processing unit 1008 is further configured to resume playback of the media content on the electronic device (for example, using the execution unit 1020) in response to the detection of a seventh user input. In some embodiments, the seventh user input includes pressing a menu button on the remote control device of the electronic device.
[0295] According to some embodiments, the processing unit 1008 is further configured to display content on the display unit (for example, using a display-enabling unit 1010). The processing unit 1008 is further configured to detect user input (for example, using a detection unit 1012) while displaying content. In response to detecting user input, the processing unit 1008 is further configured to display a user interface on the display unit (for example, using a display-enabling unit 1010). The user interface includes a plurality of exemplary natural language requests that are contextually relevant to the displayed content, and receiving a user utterance corresponding to one of the plurality of exemplary natural language requests causes the digital assistant to perform the respective action.
[0296] In some embodiments, the displayed content includes media content. In some embodiments, multiple exemplary natural language requests include natural language requests to change one or more settings associated with the media content. In some embodiments, the media content continues to play while the user interface is displayed.
[0297] In some embodiments, the processing unit 1008 is further configured to output audio associated with media content (for example, using the output unit 1018). The amplitude of the audio is not reduced in response to detected user input. In some embodiments, the displayed content includes a main menu user interface.
[0298] In some embodiments, multiple exemplary natural language requests include exemplary natural language requests relating to each of several core capabilities of the digital assistant. In some embodiments, the displayed content includes a second user interface having results associated with previous user requests. In some embodiments, multiple exemplary natural language requests include natural language requests to refine the results. In some embodiments, the user interface includes textual instructions for calling the digital assistant and interacting with it. In some embodiments, the user interface includes a visual indicator indicating that the digital assistant is not receiving voice input. In some embodiments, the user interface is overlaid on the displayed content.
[0299] In some embodiments, the processing unit 1008 is further configured to reduce the brightness of the displayed content (for example, using the display-enabled unit 1010) in order to make the user interface more prominent, in response to detection of user input.
[0300] In some embodiments, user input is detected on a remote control device of an electronic device. In some embodiments, user input includes pressing a button on the remote control device and releasing the button within a predetermined period after pressing it. In some embodiments, the button is configured to invoke a digital assistant. In some embodiments, the user interface includes textual instructions for displaying a virtual keyboard interface.
[0301] In some embodiments, the processing unit 1008 is further configured to detect a second user input (for example, using a detection unit 1012) after displaying a user interface. In response to detecting the second user input, the processing unit 1008 is further configured to display a virtual keyboard interface on the display unit (for example, using a display unit 1012).
[0302] In some embodiments, the processing unit 1008 is further configured to change the focus of the user interface to a search field on the user interface (for example, using a displayable unit 1010). In some embodiments, the search field is configured to receive text search queries via a virtual keyboard interface. In some embodiments, the virtual keyboard interface cannot be used to interact with a digital assistant. In some embodiments, the second user input includes a predetermined motion pattern on the touch-sensitive surface of a remote control device for an electronic device.
[0303] In some embodiments, a plurality of exemplary natural language requests are displays at a predetermined time after detecting user input. In some embodiments, the processing unit 1008 is further configured to display each of the plurality of exemplary natural language requests one by one in a predetermined order (e.g., using a display-enabled unit 1010). In some embodiments, the processing unit 1008 is further configured to replace the display of a previously displayed exemplary natural language request among the plurality of exemplary natural language requests with a subsequent exemplary natural language request among the plurality of exemplary natural language requests (e.g., using a display-enabled unit 1010).
[0304] In some embodiments, the content includes a second user interface having one or more items. When user input is detected, the focus of the second user interface is on one of the one or more items. Multiple exemplary natural language requests are contextually related to one of the one or more items.
[0305] According to some embodiments, the processing unit 1008 is further configured to display content on a display unit (for example, using a display-enabling unit 1010). The processing unit 1008 is further configured to detect user input (for example, using a detection unit 1012). In response to detecting user input, the processing unit 1008 is further configured to display one or more suggested examples of natural language utterances (for example, using a display-enabling unit 1010). The one or more suggested examples are contextually related to the displayed content and, when uttered by the user, cause the digital assistant to perform a corresponding action.
[0306] In some embodiments, the processing unit 1008 is further configured to detect a second user input (for example, using a detection unit 1012). The processing unit 1008 is further configured to sample speech data (for example, using a sampling unit 1016) in response to detecting a second user input. The processing unit 1008 is further configured to determine (for example, using a determination unit 1014) whether the sampled speech data contains one of one or more proposed examples of natural language utterances. The processing unit 1008 is further configured to perform a corresponding action on the utterance (for example, using an execution unit 1020) in accordance with the determination that the sampled speech data contains one of one or more proposed examples of natural language utterances.
[0307] According to some embodiments, the processing unit 1008 is further configured to display content on a display unit (for example, using a display-enabled unit 1010). The processing unit 1008 is further configured to detect user input (for example, using a detection unit 1012) while displaying content. The processing unit 1008 is further configured to sample audio data (for example, using a sampling unit 1016) in response to detecting user input. The audio data includes user utterances that express a media search request. The processing unit 1008 is further configured to retrieve a plurality of media items that satisfy the media search request (for example, using a retrieval unit 1022). The processing unit 1008 is further configured to display at least a portion of the plurality of media items on a display unit via a user interface (for example, using a display-enabled unit 1010).
[0308] In some embodiments, content remains visible on the display unit while at least a portion of multiple media items are displayed. The display area occupied by the user interface is smaller than the display area occupied by the content.
[0309] In some embodiments, the processing unit 1008 is further configured to determine (for example, using the determination unit 1014) whether the number of media items in a plurality of media items is less than or equal to a predetermined number. In accordance with the determination that the number of media items in a plurality of media items is less than or equal to a predetermined number, at least a portion of the plurality of media items contains the plurality of media items.
[0310] In some embodiments, the number of media items in at least a portion of the multiple media items is equal to a predetermined number, based on the determination that the number of media items in multiple media items is greater than a predetermined number.
[0311] In some embodiments, each of the multiple media items is associated with a relevance score for a media search request, and the relevance score of at least a portion of the multiple media items is the highest among the multiple media items.
[0312] In some embodiments, each of at least a portion of a group of media items is associated with a popularity rating, and at least a portion of the group of media items are arranged within the user interface based on their popularity ratings.
[0313] In some embodiments, the processing unit 1008 is further configured to detect a second user input (e.g., using a detection unit 1012) while displaying at least a portion of a plurality of media items. In response to detecting the second user input, the processing unit 1008 is further configured to expand the user interface (e.g., using a display-enabled unit 1010) to occupy at least a majority of the display area of the display unit.
[0314] In some embodiments, the processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the number of media items in a plurality of media items is less than or equal to a predetermined number, in response to detecting a second user input. The processing unit 1008 is further configured to obtain a second plurality of media items that at least partially satisfy the media search request, in accordance with the determination that the number of media items in a plurality of media items is less than or equal to a predetermined number, and the second plurality of media items differ from at least a portion of the media items. The processing unit 1008 is further configured to display the second plurality of media items on a display unit (e.g., using a display-enabled unit 101) via an expanded user interface.
[0315] In some embodiments, the processing unit 1008 is further configured to determine (for example, using the determination unit 1014) whether a media search request contains more than one search parameter. Following the determination that the media search request contains more than one search parameter, a second set of media items are organized within the expanded user interface according to the more than one search parameter of the media search request.
[0316] In some embodiments, the processing unit 1008 is further configured to display at least a second portion of the media items via an expanded user interface (for example, using a displayable unit 1010) in accordance with a determination that the number of media items in the plurality of media items is greater than a predetermined number. The at least second portion of the plurality of media items is different from at least a portion of the plurality of media items.
[0317] In some embodiments, at least a second portion of a plurality of media items includes two or more media types, and at least a second portion of a plurality of media items is organized according to each of the two or more media types within an expanded user interface.
[0318] In some embodiments, the processing unit 1008 is further configured to detect a third user input (e.g., using a detection unit 1012). The processing unit 1008 is further configured to scroll the enlarged user interface (e.g., using a display-enabled unit 1010) in response to detecting a third user input. The processing unit 1008 is further configured to determine (e.g., using a determination unit 1014) whether the enlarged user interface has scrolled beyond a predetermined position on the enlarged user interface. The processing unit 1008 is further configured to display at least a third portion of a plurality of media items on the enlarged user interface (e.g., using a display-enabled unit 1010) in response to determining that the enlarged user interface has scrolled beyond a predetermined position on the enlarged user interface. The at least third portion of the plurality of media items is arranged on the enlarged user interface according to one or more media content providers associated with the third plurality of media items.
[0319] The operations described above with reference to Figures 5A to 5I can be optionally performed by the components shown in Figures 1 to 3 and Figures 4A to 4B. For example, display operations 502, 508-514, 520, 524, 530, 536, 546, 556, 560, 562, 576, 582, 588, 592, detection operations 504, 538, 542, 550, 558, 566, 570, determination operations 506, 516, 522, 526, 528, 574, 578, sampling operations 518, 572, execution operations 532, 584, acquisition operations 534, 544, 580, 586, 590, cancellation operations 540, 568, receiving unit 554, and switching operations 552, 564 may be performed by one or more of the operating system 252, GUI module 256, application module 262, digital assistant module 426, and processor(s) 204, 404. To those skilled in the art, it will be obvious how the other processes are carried out based on the components shown in Figures 1-3 and 4A-4B.
[0320] Based on several embodiments, Figure 11 shows a functional block diagram of an electronic device 1100 configured according to the principles of various embodiments described, for example, to provide voice control of media playback and real-time updates of a virtual assistant's knowledge. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software to perform the principles of the various embodiments described. Those skilled in the art will understand that the functional blocks described in Figure 11 may be combined or separated into subblocks to perform the principles of the various embodiments described. Accordingly, the description herein supports, at its discretion, any possible combination or division of the functional blocks described herein, or further definitions.
[0321] As shown in Figure 11, the electronic device 1100 may include an input unit 1103 (e.g., remote control unit 124, or similar) configured to receive user input such as tactile input or gesture input; an audio input unit 1104 (e.g., microphone 272, or similar) configured to receive audio data; a speaker unit 116 (e.g., speaker 268, or similar) configured to output audio; and a communication unit 1107 (e.g., communication subsystem 224, or similar) configured to transmit and receive information from external devices over a network. In some embodiments, the electronic device 1100 may optionally include a display unit 1102 (e.g., display unit 126, or similar) configured to display media, interfaces, and other content. The electronic device 1100 may further include an input unit 1103, an audio input unit 1104, a speaker unit 1106, a communication unit 1107, and a processing unit 1108 optionally coupled to a display unit 1102. In some embodiments, the processing unit 1108 may include a display-enabled unit 1110, a detection unit 1112, a determination unit 1114, a sampling unit 1116, an output unit 1118, an execution unit 1120, an acquisition unit 1122, a identification unit 1124, and a transmission unit 1126.
[0322] According to some embodiments, the processing unit 1108 is configured to display content on a display unit (e.g., display unit 1102 or a separate display unit) (e.g., using a display-enabled unit 1110). The processing unit 1108 is further configured to detect user input (e.g., using a detection unit 1112) while displaying content. The processing unit 1108 is further configured to sample audio data (e.g., using a sampling unit 1016 and an audio input unit 1104) in response to detecting user input. The audio data includes user utterances. The processing unit 1108 is further configured to acquire a determination of user intent corresponding to the user utterance (e.g., using an acquisition unit 1122). The processing unit 1108 is further configured to acquire a determination (e.g., using an acquisition unit 1122) of whether the user intent includes a request to adjust the state or settings of an application on an electronic device. The processing unit 1108 is further configured to adjust the state or settings of an application to satisfy the user intent (for example, using the task execution unit 1120) when it has determined that the user intent includes a request to adjust the state or settings of an application on an electronic device.
[0323] In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to play a specific media item. Adjusting the state or settings of an application to satisfy user intent includes playing a specific media item.
[0324] In some embodiments, the displayed content includes a user interface having a media item, and the user utterance does not explicitly limit to a specific media item to be played. The processing unit 1108 is further configured to determine (for example, using a determination unit 1114) whether the focus of the user interface is on a media item. In accordance with the determination that the focus of the user interface is on a media item, the processing unit 1108 is further configured to identify the media item as a specific media item to be played (for example, using a identification unit 1124).
[0325] In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to launch the application on the electronic device. In some embodiments, the displayed content includes media content being played on the electronic device, and the state or settings relate to the media content being played on the electronic device. In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to fast forward or rewind the media content being played on the electronic device. In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to jump forward or backward within the media content to play a specific portion of the media content. In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to pause playback of media content on the electronic device. In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to turn on or off subtitles for media content.
[0326] In some embodiments, the displayed content includes a user interface having a first media item and a second media item.
[0327] In some embodiments, a request to adjust the state or settings of an application on an electronic device includes a request to switch the focus of the user interface from a first media item to a second media item. Adjusting the state or settings of an application to satisfy user intent includes switching the focus of the user interface from a first media item to a second media item.
[0328] In some embodiments, the displayed content includes media content being played on a media device. User utterance is a natural language expression indicating that the user did not hear a portion of the audio associated with the media content. A request to adjust the state or settings of an application on an electronic device includes a request to replay the portion of the media content corresponding to the portion of audio the user did not hear. The processing unit 1108 is further configured to rewind the media content by a predetermined amount to an earlier portion of the media content (for example, using the task execution unit 1120) and restart playback of the media content from the earlier portion (for example, using the task execution unit 1120).
[0329] In some embodiments, the processing unit 1108 is further configured to turn on closed captions (for example, using the task execution unit 1120) before restarting playback of the media content from the previous portion.
[0330] In some embodiments, a request to adjust the state or settings of an application on an electronic device further includes a request to increase the volume of audio associated with media content. Adjusting the state or settings of an application further includes increasing the volume of audio associated with media content before restarting playback of the media content from a previous portion.
[0331] In some embodiments, speech utterances within audio associated with media content are converted to text. Adjusting the application state or settings further includes displaying portions of the text while playback of media content is restarted from a previous portion.
[0332] In some embodiments, the processing unit 1108 is further configured to obtain a determination of the user's emotion associated with the user's utterance (for example, using the acquisition unit 1122). The user's intent is determined based on the determined user emotion.
[0333] In some embodiments, the processing unit 1108 is further configured to obtain (e.g., using the acquisition unit 1122) a determination as to whether the user intent is one of a plurality of predetermined request types, in response to obtaining a determination that the user intent does not include a request to adjust the state or settings of an application on an electronic device. In response to obtaining a determination that the user intent is one of a plurality of predetermined request types, the processing unit 1108 is further configured to obtain (e.g., using the acquisition unit 1122) a result that at least partially satisfies the user intent and to display the result on the display unit in text format (e.g., using the display-enabled unit 1110).
[0334] In some embodiments, a plurality of predetermined request types include a request for the current time at a specific location. In some embodiments, a plurality of predetermined request types include a request for a joke. In some embodiments, a plurality of predetermined request types include a request for information about media content being played on an electronic device. In some embodiments, the text format result is overlaid on the displayed content. In some embodiments, the displayed content includes media content being played on an electronic device, and the media content continues to play while the text format result is displayed.
[0335] In some embodiments, the processing unit 1108 is further configured to obtain (e.g., using the acquisition unit 1122) a result that at least partially satisfies a second user intent, in response to obtaining a determination that the user intent is not one of a plurality of predetermined request types, and to determine (e.g., using the determination unit 1114) whether the displayed content includes media content being played on an electronic device. The processing unit 1108 is further configured to determine (e.g., using the determination unit 1114) whether the media content can be paused, in accordance with the determination that the displayed content includes media content. The processing unit 1108 is further configured to display (e.g., using the display-enabled unit 1110) a second user interface having part of the second result on the display unit, in accordance with the determination that the media content cannot be paused. The display area occupied by the second user interface on the display unit is smaller than the display area occupied by the media content on the display unit.
[0336] In some embodiments, the user intent includes a request for a weather forecast for a specific location. The user intent includes a request for information associated with a sports team or athlete. In some embodiments, the user intent is not a media search query, and the second result includes one or more media items having media content that at least partially satisfies the user intent. In some embodiments, the second result further includes non-media data that at least partially satisfies the user intent. In some embodiments, the user intent is a media search query, and the second result includes multiple media items corresponding to the media search query.
[0337] In some embodiments, the processing unit 1108 is further configured to display a third user interface having a portion of the second result on the display unit (for example, using a display-enabled unit 1110) upon determination that the displayed content does not include media content being played on an electronic device, the third user interface occupies more than half of the display area of the display unit.
[0338] In some embodiments, the displayed content includes a main menu user interface.
[0339] In some embodiments, the displayed content includes a third user interface having prior results related to a previous user request received before detecting user input. Upon determination that the displayed content does not include media content currently being played on the electronic device, the display of the prior results within the third user interface is replaced with a display of the second result.
[0340] In some embodiments, the processing unit 1108 is further configured to determine (for example, using the determination unit 1114) whether the displayed content includes a second user interface having a previous result from a previous user request, based on the determination that the displayed content includes media content being played on an electronic device. Based on the determination that the displayed content includes a second user interface having a previous result from a previous user request, the previous result is replaced with the second result.
[0341] In some embodiments, the processing unit 1108 is further configured to pause playback of media content on an electronic device (e.g., using a task execution unit 1120) upon determining that the media content can be paused, and to display a third user interface having a portion of the second result on a display unit (e.g., using a display-enabled unit 1110), the third user interface occupying more than half of the display area of the display unit.
[0342] In some embodiments, the processing unit 1108 is further configured to transmit audio data to a server (e.g., using the transmission unit 1126 and the communication unit 1107) to perform natural language processing, and to instruct the server (e.g., using the transmission unit 1126) that the audio data is associated with a media application. This instruction causes the natural language processing to be biased towards media-related user intent.
[0343] In some embodiments, the processing unit 1108 is further configured to transmit the audio data to a server (e.g., a transmission unit 1126) to perform speech-to-text processing.
[0344] In some embodiments, the processing unit 1108 is further configured to instruct the server (for example, using the transmission unit 1126) that the audio data is associated with a media application. This instruction biases the speech-to-text processing towards media-related text results.
[0345] In some embodiments, the processing unit 1108 is further configured to acquire a textual representation of the user's utterance (for example, using an acquisition unit 1122), the textual representation being based on previous user utterances received before sampling the audio data.
[0346] In some embodiments, the text representation is based on the time when the previous user utterance was received before the audio data was sampled.
[0347] In some embodiments, the processing unit 1108 is further configured to obtain (for example, using the acquisition unit 1122) a determination that the user intent does not correspond to one of several core capabilities associated with the electronic device. The processing unit 1108 is further configured to cause the second electronic device to perform a task (for example, using the task execution unit 1120) to assist in satisfying the user intent.
[0348] In some embodiments, the processing unit 1108 is further configured to obtain (for example, using the acquisition unit 1122) a determination of whether the user utterance contains ambiguous terms. Upon obtaining the determination that the user utterance contains ambiguous terms, the processing unit 1108 is further configured to obtain (for example, using the acquisition unit 1122) two or more candidate user intents based on the ambiguous terms and to display the two or more candidate user intents on the display unit (for example, using the display-enabled unit 1110).
[0349] In some embodiments, the processing unit 1108 is further configured to receive a user selection (for example, using the detection unit 1112) of one of the two or more user intent candidates while displaying two or more candidate user intents. The user intent is determined based on the user selection.
[0350] In some embodiments, the processing unit 1108 is further configured to detect a second user input (e.g., using a detection unit). In response to detecting the second user input, the processing unit 1108 is further configured to sample second audio data (e.g., using a sampling unit 1116). The second audio data includes a second user utterance representing a user selection.
[0351] In some embodiments, two or more interpretations are displayed without outputting utterances associated with two or more candidate user intentions.
[0352] According to some embodiments, the processing unit 1108 is further configured to display content on a display unit (e.g., display unit 1102 or a separate display unit) (e.g., using a display-enabled unit 1110). The processing unit 1108 is further configured to detect user input (e.g., using a detection unit 1112) while displaying content. The processing unit 1108 is further configured to display a virtual keyboard interface on the display unit (e.g., using a display-enabled unit 1110) in response to detecting user input. The processing unit 1108 is further configured to make selectable affordances appear on the display of the second electronic device (e.g., using a task execution unit 1120). The selection of affordances allows text input to be received by the electronic device (e.g., using a communication unit 1107) via the keyboard of the second electronic device.
[0353] In some embodiments, the processing unit 1108 is further configured to receive text input via a keyboard of a second electronic device (e.g., using a detection unit 1112), where the text input represents a user request. The processing unit 1108 retrieves results that at least partially satisfy the user request (e.g., using an acquisition unit 1122), and is further configured to display a user interface on a display unit (e.g., using a display-enabled unit 1110), where the user interface includes at least a portion of the results.
[0354] In some embodiments, the displayed content includes a second user interface having a plurality of exemplary natural language requests. In some embodiments, the displayed content includes media content. In some embodiments, the displayed content includes a third user interface having results from previous user requests, the third user interface occupying at least a majority of the display area of the display unit. In some embodiments, a virtual keyboard interface is superimposed on at least a portion of the third user interface. In some embodiments, user input is detected via a remote control device of an electronic device, the remote control device being a different device from the second electronic device. In some embodiments, user input includes a predetermined motion pattern on the touch-sensitive surface of the remote control device. In some embodiments, user input is detected via a second electronic device.
[0355] The operations described above with reference to Figures 7A to 7C and Figure 9 can be optionally performed by the components shown in Figures 1 to 3 and Figure 4A. The operations described above with reference to Figures 7A to 7C and Figure 9 can be optionally performed by the components shown in Figures 1 to 3 and Figures 4A to 4B. For example, display operations 702, 716, 732, 736, 738, 742, 746, 902, 906, 914, detection operations 704, 718, 904, 910, determination operations 708, 710, 712, 714, 720, 724, 728, 736, 740, sampling operation 706, execution operations 722, 726, 744, 908, acquisition operations 730, 734, 912, and switching operations 552, 564 may be performed by one or more of the following: operating systems 252, 352, GUI modules 256, 356, application modules 262, 362, digital assistant module 426, and processors (one or more) 204, 304, 404. To those skilled in the art, it will be obvious how the other processes are carried out based on the components shown in Figures 1-3 and 4A-4B.
[0356] According to some embodiments, a computer-readable storage medium (e.g., a non-temporary computer-readable storage medium) is provided which stores one or more programs executed by one or more processors of an electronic device, and which include instructions that perform any of the methods described herein.
[0357] According to some embodiments, an electronic device (e.g., a portable electronic device) is provided that includes means for carrying out any of the methods described herein.
[0358] According to some embodiments, an electronic device (e.g., a portable electronic device) is provided that includes a processing unit configured to perform any of the methods described herein.
[0359] According to some embodiments, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and a memory for storing one or more programs executed by the one or more processors, wherein the one or more programs include instructions for performing any of the methods described herein.
[0360] Exemplary methods, non-temporary computer-readable storage media, systems, and electronic devices are described in the following sections. 1. A method for operating a digital assistant in a media system, the method being: In an electronic device having one or more processors and memory, Displaying content on a display unit, Detecting user input and To determine whether user input corresponds to the first input format, According to the determination that the user input corresponds to the first input format, Displaying multiple exemplary natural language requests on a display unit, wherein the multiple exemplary natural language requests are contextually relevant to the displayed content, and receiving a user utterance corresponding to one of the multiple exemplary natural language requests causes the digital assistant to perform the respective action. A method that includes this. 2. The method according to item 1, wherein user input is detected on a remote control device of an electronic device. 3. The method according to item 2, wherein the first input method includes pressing a button on a remote control device and releasing the button within a predetermined period of time. 4. The method according to any one of items 1 to 3, wherein multiple exemplary natural language requests are displayed on a display unit via a first user interface, and the first user interface is overlaid on the displayed content. 5. The method described in any one of items 1 through 4, wherein the displayed content includes media content, and the media content continues to play while displaying multiple exemplary natural language requests. 6. The method of any one of items 1 to 5, further comprising displaying a visual indicator on the display unit that indicates the digital assistant is not processing voice input, in accordance with the determination that user input corresponds to a first input format. 7. The method according to any one of items 1 to 6, wherein, when it is determined that user input corresponds to a first input format, several exemplary natural language requests are displayed on the display unit after a predetermined time. 8. The method according to any one of items 1 to 7, wherein each of a group of exemplary natural language requests is presented separately at different times in a predetermined order. 9. Displaying multiple exemplary natural language requests A method according to any one of items 1 through 8, comprising displaying multiple lists of exemplary natural language requests, each list being displayed in rotation at different times. 10. Based on the determination that the user input does not correspond to the first input format, To determine whether user input corresponds to a second input format, In accordance with the determination that user input corresponds to the second input format, Sampling audio data and Determining whether the audio data includes the user request, Based on the determination that the audio data contains the user request, perform a task that satisfies the user request at least partially. The method described in any one of items 1 through 9, further including the method described in any one of items 1 through 9. 11. The method according to item 10, wherein the second input method includes pressing a button on a remote control device of an electronic device and holding the button down for a longer period than a predetermined time. 12. The method according to item 10 or 11, further comprising displaying a request for clarification of user intent on the display unit in accordance with the determination that the audio data does not contain the user request. 13. The method according to any one of items 10 to 12, wherein the displayed content includes media content, and the media content continues to play on the electronic device while the audio data is being sampled and while the task is being performed. 14. Outputting audio associated with media content, In accordance with the determination that the user input corresponds to the second input format, the amplitude of the sound is reduced, The method described in item 13, further including the method described in item 13. 15. A method of any one of items 10 to 14 wherein the task is performed without outputting any utterances related to the task from an electronic device. 16. A method according to any one of items 10 to 15, wherein audio data is sampled while user input is being detected. 17. A method according to any one of items 10 to 15, wherein audio data is sampled during a predetermined period after detecting user input. 18. Audio data is sampled via a first microphone on a remote control device of an electronic device, and the method is as follows: While sampling audio data, background audio data is sampled via a second microphone on the remote control device, Using background audio data to remove background noise from audio data, The method described in any one of items 10 through 17, further including the method described in item 10 through 17. 19. Audio associated with the displayed content is output from the electronic device via an audio signal, and the method is as follows: Removing background noise from audio data using an audio signal. The method described in any one of items 10 through 18, further including the method described in any one of items 10 through 18. 20. The method according to any one of items 10 to 19, further comprising displaying a visual cue on the display unit prompting the user to provide a verbal request in response to detecting user input. 21. The task to be executed is, To obtain results that at least partially satisfy the user requirements, Displaying a second user interface on a display unit, wherein the second user interface includes a portion of the result, at least a portion of the content remains displayed while the second user interface is displayed, and the display area of the second user interface on the display unit is smaller than the display area of at least a portion of the content on the display unit. A method of any one of items 10 through 20, including the method described above. 22. The method described in item 21, wherein a second user interface is overlaid on the displayed content. 23. The result section contains one or more media items, and the method is: Receiving a selection of one or more media items via a second user interface, The display unit will show the media content associated with the selected media item, The method described in item 21 or 22, further including the method described in item 21 or 22. twenty four. Detecting a second user input while displaying the second user interface, In response to detecting a second user input, the display of the second user interface is stopped, The method described in item 21 or 22, further including the method described in item 21 or 22. 25. The method according to item 24, wherein a second user input is detected on a remote control unit of an electronic device, and the second user input includes a first predetermined motion pattern on the touch-sensing surface of the remote control unit. 26. Detecting a third user input while displaying the second user interface, In response to detecting a third user input, the display of the second user interface is replaced with the display of the third user interface on the display unit, wherein the third user interface includes at least a portion of the result and occupies at least a majority of the display area of the display unit. The method described in item 21 or 22, further including the method described in item 21 or 22. 27. The method according to item 26, wherein a third user input is detected on a remote control unit of an electronic device, and the third user input includes a second predetermined motion pattern on the touch-sensing surface of the remote control unit. 28. In response to the detection of a third user input, The method according to item 26 or 27, further comprising obtaining a second result different from the first result, wherein the second result at least partially satisfies the user requirements, and a third user interface includes at least a portion of the second result. 29. The method described in item 28, wherein the second result is based on a user request received before detecting user input. 30. The method described in item 28 or 29, wherein the focus of the second user interface is on an item in the result portion while a third user input is detected, and the second result is contextually related to the item. 31. The method according to any one of items 26 to 30, wherein the displayed content includes media content, and playback of the media content on the electronic device is paused in response to the detection of a third user input. 32. At least a portion of the results includes one or more media items, and the method is Receiving a selection of one or more media items via a third user interface, Displaying media content associated with media items on the display unit, The method described in any one of items 26 to 31, further including the method described in any one of items 26 to 31. 33. While displaying the third user interface, a fourth user input associated with orientation on the display unit is detected, In response to detecting a fourth user input, The third user interface's focus is switched from the first item to the second item on the third user interface, with the second item positioned in the direction described above relative to the first item. The method described in any one of items 26 to 32, further including the method described in any one of items 26 to 32. 34. Detecting a fifth user input while displaying a third user interface, In response to detecting a fifth user input, Displaying the search field, The display unit displays a virtual keyboard interface, and input received via the virtual keyboard interface results in text being entered into the search field. The method described in any one of items 26 to 33, further including the method described in any one of items 26 to 33. 35. Detecting a sixth user input while displaying the third user interface, In response to detecting a sixth user input, The process involves sampling a second audio data set, wherein the second audio data set includes a second user request. The second user request is to determine whether it is a request to narrow down the results of the user request, Based on the determination that the second user request is a request to narrow down the results of the user request, Displaying a subset of results through a third user interface, The method described in any one of items 26 to 34, further including the method described in any one of items 26 to 34. 36. The method described in item 35, wherein a subset of the results is displayed at the top of the third user interface. 37. The second user request is determined not to be a request to narrow down the results of the user request, To obtain a third result that at least partially satisfies the second user requirement, Displaying a portion of the third result through a third user interface, The method described in item 35 or 36, further including the method described in item 35 or 36. 38. The method described in item 37, wherein the third result section is displayed at the top of the third user interface. 39. To obtain a fourth result that at least partially satisfies the user requirements or the second user requirements, Displaying a portion of the fourth result through a third user interface, The method described in any one of items 35 to 38, further including the method described in any one of items 35 to 38. 40. The method described in item 39, wherein the fourth result is displayed in the row after the top row of the third user interface. 41. The method according to item 39 or 40, wherein the focus of the third user interface is on one or more items of the third user interface while the sixth user input is detected, and the fourth result is contextually related to one or more items. 42. Detecting a seventh user input while displaying the third user interface, In response to detecting a seventh user input, the display of the third user interface is stopped, The method described in any one of items 26 to 41, further including the method described in any one of items 26 to 41. 43. The method described in item 42, wherein the displayed content is media content, playback of the media content on the electronic device is paused in response to the detection of a third user input, and playback of the media content on the electronic device is resumed in response to the detection of a seventh user input. 44. The method of item 42 or 43, wherein the seventh user input includes pressing a menu button on a remote control device of an electronic device. 45. A method for operating a digital assistant in a media system, the method being: In an electronic device having one or more processors and memory, Displaying content on a display unit, Detecting user input while displaying content, In response to detecting user input, The display unit displays a user interface, the user interface includes multiple exemplary natural language requests that are contextually relevant to the displayed content, and the digital assistant performs the corresponding action upon receiving a user utterance corresponding to one of the multiple exemplary natural language requests. A method that includes this. 46. The method described in item 45, where the displayed content includes media content. 47. The method described in item 46, wherein multiple exemplary natural language requests include a natural language request to change one or more settings associated with media content. 48. The method described in item 46 or 47, wherein media content continues to play while the user interface is displayed. 49. A method according to any one of items 46 to 48, further comprising outputting audio associated with media content, wherein the amplitude of the audio is not reduced in response to detection of user input. 50. The method described in item 45, wherein the displayed content includes the main menu user interface. 51. The method of item 50, wherein the multiple exemplary natural language requests include exemplary natural language requests relating to each of the multiple core capabilities of the digital assistant. 52. The method of item 45, wherein the displayed content includes a second user interface having results associated with a previous user request. 53. The method described in item 52, which includes a natural language request to narrow down the results, with multiple exemplary natural language requests. 54. A method of any one of items 45 to 53, wherein the user interface includes text instructions for calling a digital assistant and interacting with it. 55. The method described in any one of items 45 to 54, wherein the user interface includes a visual indicator that indicates that the digital assistant is not receiving voice input. 56. A method of any one of items 45 to 55, wherein a user interface is overlaid on displayed content. 57. A method according to any one of items 45 to 56, further comprising reducing the brightness of displayed content to make the user interface more prominent in response to detection of user input. 58. A method according to any one of items 45 to 57, wherein user input is detected on a remote control device of an electronic device. 59. The method of item 58, wherein user input includes pressing a button on a remote control device and releasing the button within a predetermined period of time after pressing the button. 60. The method described in item 59, wherein a button is configured to invoke a digital assistant. 61. A method according to any one of items 45 to 60, wherein the user interface includes textual instructions for displaying a virtual keyboard interface. 62. After displaying the user interface, detect a second user input, Up...
Claims
1. A method performed by computer, In an electronic device having one or more processors and memory, Receiving a first audio input, including a media request, from a secondary electronic device, In response to receiving the first audio input, a user interface is displayed on a first portion of the display unit, wherein the user interface includes a plurality of media items. Receiving a second audio input from the secondary electronic device, wherein the second audio input includes references to individual media items of the plurality of media items, Based on the determination that the reference to the individual media item contains an ambiguous reference, the individual media item is identified based on the focus of the user interface, In accordance with the determination that the reference to the individual media item does not contain an ambiguous reference, the user interface will not identify the individual media item based on its focus, A method comprising: displaying preview information for the individual media items of the plurality of media items in response to receiving the second audio input.
2. A method according to claim 1, wherein the focus of the user interface is on the individual media item.
3. A method according to claim 1 or 2, A method for identifying the individual media items based on the focus of the user interface and the second voice input.
4. A method according to any one of claims 1 to 3, wherein the preview information includes at least one of a synopsis, a cast list, a release date, and one or more user ratings.
5. A method according to any one of claims 1 to 4, wherein the reference of the plurality of media items to the individual media items is an ambiguous reference that does not include the name of the individual media item.
6. A method according to any one of claims 1 to 5, A method comprising displaying the preview information relating to the individual media items of the plurality of media items, in accordance with the determination that the second audio input includes a request to adjust the state of the application settings of the electronic device.
7. A method according to any one of claims 1 to 6, wherein the second voice input includes a request to view a specific user interface corresponding to information about the individual media items of the plurality of media items.
8. The method according to claim 7, wherein the request to view a specific user interface is: A request to view the synopsis, A request to view the cast list, A request to view the publication date, A request to view one or more user ratings, and a method including one of the following.
9. A method according to any one of claims 1 to 8, Receiving a third voice input from the secondary electronic device, wherein the third voice input includes a request to navigate to the main menu interface. A method comprising: displaying the main menu interface in response to receiving the third voice input.
10. A method according to any one of claims 1 to 9, Receiving a third audio input from the secondary electronic device, wherein the third audio input includes a reference to a second individual media item of the plurality of media items, A method comprising: changing the focus of the user interface to the second individual media item in response to receiving the third audio input.
11. A method according to any one of claims 1 to 10, A method comprising: in response to receiving the second audio input, initiating playback of media content associated with the individual media items of the plurality of media items.
12. A method according to any one of claims 1 to 11, Receiving a third audio input from the secondary electronic device, wherein the third audio input includes a request to navigate to an application for purchasing media items. A method comprising: displaying the application for purchasing media items in response to receiving the third voice input.
13. A method according to any one of claims 1 to 12, wherein the plurality of media items include a specific category of results, Receiving a third audio input from the secondary electronic device, wherein the third audio input includes a request to adjust the focus of the user interface to the resulting specific category, A method comprising: adjusting the focus of the user interface to the specific category of the result in response to receiving the third voice input.
14. A computer program that causes a computer to perform the method according to any one of claims 1 to 13.
15. A memory for storing the computer program described in claim 14, A computer system comprising one or more processors capable of executing the computer program stored in the memory, wherein the computer system is configured to communicate with a display generation component and one or more input devices.
16. A computer system configured to communicate with a display generation component and one or more input devices, A computer system comprising means for performing the method described in any one of claims 1 to 13.
Citation Information
Patent Citations
Content display, content display method and content display system constituted of the content display and retrieval server apparatus
JP2010130409A
Method and device for performing user function by using voice recognition
JP2013143151A
Information processing device, information processing method, information processing program and terminal device
JP2013198085A
Voice control of multimedia content
US20060041926A1