Intelligent automated assistant for TV user interaction
A virtual assistant system processes user speech to control media devices, enhancing user interaction by facilitating intuitive content selection and recommendation, thus improving the user experience.
Patent Information
- Application Number
- JP2025129486
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-09-26
- Filing Date
- 2025-08-01
- Publication Date
- 2025-12-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing media control devices, such as televisions and set-top boxes, lack intuitive control mechanisms and make it cumbersome for users to discover desired media content, degrading the user experience.
A virtual assistant system that processes user speech input to determine intent and control media interactions, displaying user interfaces to facilitate content selection and recommendation based on displayed content and history.
Provides an intuitive and efficient user experience by allowing natural language control of media devices, simplifying content discovery and access.
Smart Images

Figure 2025176016000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 019,312, entitled "INTELLIGENT AUTOMATED ASSISTANT FOR TV USER INTERACTIONS," filed June 30, 2014, and U.S. Patent Application No. 14 / 498,503, entitled "INTELLIGENT AUTOMATED ASSISTANT FOR TV USER INTERACTIONS," filed September 26, 2014, which applications are incorporated herein by reference in their entireties for all purposes.
[0002] This application is also related to co-pending U.S. patent application Ser. No. 62 / 019,292, filed Jun. 30, 2014, entitled "REAL-TIME DIGITAL ASSISTANT KNOWLEDGE UPDATES" (Attorney Docket No. 106843097900 (P22498USP1)), which is incorporated herein by reference in its entirety. [Technical field]
[0003] This application relates generally to controlling television user interactions, and more particularly to processing utterances to a virtual assistant to control television user interactions. [Background technology]
[0004] Intelligent automated assistants (or virtual assistants) provide an intuitive interface between users and electronic devices. These assistants can enable users to interact with devices or systems using natural language in spoken and / or textual form. For example, a user can access the services of an electronic device by providing spoken user input in natural language form to a virtual assistant associated with the electronic device. The virtual assistant can perform natural language processing on the spoken user input to infer the user's intent and manipulate the user's intent into a task. The task can then be performed by executing one or more functions of the electronic device, and in some embodiments, associated output can be returned to the user in natural language form.
[0005] While mobile phones (e.g., smartphones), tablet computers, and the like benefit from virtual assistant control, many other user devices lack such convenient control mechanisms. For example, learning user interactions with media control devices (e.g., televisions, television set-top boxes, cable boxes, gaming devices, streaming media devices, digital video recorders, etc.) can be complex and difficult. Furthermore, with the proliferation of sources available through such devices (e.g., over-the-air TV, subscription TV services, streaming video services, cable-on-demand video services, web-based video services, etc.), discovering desired media content to consume can be cumbersome or even tedious for some users. As a result, many media control devices can degrade the user experience and frustrate many users. Summary of the Invention
[0006] A system and process for controlling television interactions using a virtual assistant are disclosed. In one embodiment, speech input from a user can be received. Media content can be determined based on the speech input. A first user interface having a first size can be displayed, and the first user interface can include selectable links to the media content. A selection of one of the selectable links can be received. In response to the selection, a second user interface having a second size larger than the first size can be displayed, and the second user interface includes media content associated with the selection.
[0007] In another example, speech input from a user can be received at a first device having a first display. The user's intent for the speech input can be determined based on content displayed on the first display. Media content can be determined based on the user intent. The media content can be played on a second device associated with a second display.
[0008] In another example, speech input can be received from a user, the speech input including a query associated with content displayed on a television display. User intent for the query can be determined based on one or more of the content displayed on the television display and a media content viewing history. Results of the query can be displayed based on the determined user intent.
[0009] In another embodiment, media content can be displayed on the display. Input from a user can be received. A virtual assistant query can be determined based on media content and / or media content browsing history. A recommended virtual assistant query can be displayed on the display. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 illustrates an exemplary system for controlling television user interactions using a virtual assistant.
[0011] [Figure 2] FIG. 2 is a block diagram of an exemplary user device, according to various embodiments.
[0012] [Figure 3] FIG. 1 is a block diagram of an exemplary media control device in a system for controlling television user interaction.
[0013] [Figure 4A] FIG. 1 illustrates an exemplary speech input interface on video content. [Figure 4B] FIG. 1 illustrates an exemplary speech input interface on video content. [Figure 4C] FIG. 1 illustrates an exemplary speech input interface on video content. [Figure 4D] FIG. 1 illustrates an exemplary speech input interface on video content. [Figure 4E] FIG. 1 illustrates an exemplary speech input interface on video content.
[0014] [Figure 5] 1 illustrates an exemplary media content interface on video content.
[0015] [Figure 6A] FIG. 1 illustrates an exemplary media details interface on video content. [Figure 6B] FIG. 1 illustrates an exemplary media details interface on video content.
[0016] [Figure 7A] FIG. 1 illustrates an exemplary media transition interface. [Figure 7B]FIG. 1 illustrates an exemplary media transition interface.
[0017] [Figure 8A] FIG. 10 illustrates an exemplary speech input interface on menu content. [Figure 8B] FIG. 10 illustrates an exemplary speech input interface on menu content.
[0018] [Figure 9] FIG. 10 illustrates an example virtual assistant result interface on menu content.
[0019] [Figure 10] FIG. 1 illustrates an exemplary process for controlling television interactions using a virtual assistant and displaying associated information using different interfaces.
[0020] [Figure 11] FIG. 1 illustrates exemplary television media content on a mobile user device.
[0021] [Figure 12] FIG. 1 illustrates an exemplary television control using a virtual assistant.
[0022] [Figure 13] FIG. 1 illustrates exemplary photo and video content on a mobile user device.
[0023] [Figure 14] FIG. 1 illustrates an exemplary media display control using a virtual assistant.
[0024] [Figure 15] FIG. 1 illustrates an example virtual assistant interaction with results on a mobile user device and a media display device.
[0025] [Figure 16] FIG. 1 illustrates an exemplary virtual assistant interaction with media results on a media display device and a mobile user device.
[0026] [Figure 17] FIG. 1 illustrates exemplary proximity-based media device control.
[0027] [Figure 18] FIG. 1 illustrates an exemplary process for controlling television interactions using a virtual assistant and multiple user devices.
[0028] [Figure 19] FIG. 1 illustrates an exemplary speech input interface with a virtual assistant query for video background content.
[0029] [Figure 20] FIG. 10 illustrates an exemplary information virtual assistant response on video content.
[0030] [Figure 21] FIG. 1 illustrates an exemplary speech input interface with a virtual assistant query for media content associated with animated background content.
[0031] [Figure 22] FIG. 1 illustrates an exemplary virtual assistant responsive interface with selectable media content.
[0032] [Figure 23A] FIG. 1 illustrates an exemplary page of a program menu. [Figure 23B] FIG. 1 illustrates an exemplary page of a program menu.
[0033] [Figure 24] FIG. 1 illustrates an exemplary media menu divided into categories.
[0034] [Figure 25] FIG. 1 illustrates an exemplary process for controlling television interaction using presented media content on a display and media content viewing history.
[0035] [Figure 26] FIG. 1 illustrates an example interface with virtual assistant query recommendations based on animated background content.
[0036] [Figure 27] FIG. 10 illustrates an exemplary interface for confirming a selection of recommended queries.
[0037] [Figure 28A] FIG. 1 illustrates an example virtual assistant answer interface based on a selected query. [Figure 28B] FIG. 1 illustrates an example virtual assistant answer interface based on a selected query.
[0038] [Figure 29] FIG. 1 illustrates an exemplary interface with media content notifications and virtual assistant query recommendations based on the notifications.
[0039] [Figure 30] FIG. 1 illustrates a mobile user device with exemplary photo and video content playable on a media control device.
[0040] [Figure 31] FIG. 1 illustrates an exemplary mobile user device interface with virtual assistant query recommendations based on playable user device content and based on video content displayed on a separate display.
[0041] [Figure 32]FIG. 10 illustrates an exemplary interface with virtual assistant query recommendations based on playable content from a separate user device.
[0042] [Figure 33] FIG. 1 illustrates an exemplary process for recommending virtual assistant interactions for controlling media content.
[0043] [Figure 34] FIG. 1 illustrates a functional block diagram of an electronic device configured to control television interactions using a virtual assistant and display related information using different interfaces, according to various embodiments.
[0044] [Figure 35] FIG. 1 illustrates a functional block diagram of an electronic device configured to control television interactions using a virtual assistant and multiple user devices, according to various embodiments.
[0045] [Figure 36] FIG. 1 illustrates a functional block diagram of an electronic device configured to control television interactions using media content displayed on a display and a media content viewing history, according to various embodiments.
[0046] [Figure 37] FIG. 1 illustrates a functional block diagram of an electronic device configured to recommend virtual assistant interactions for controlling media content, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0047] In the following description of the embodiments, reference is made to the accompanying drawings, which show, by way of illustration, specific embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the various embodiments.
[0048] This relates to a system and process for controlling television user interaction using a virtual assistant. In one embodiment, the virtual assistant can be used to interact with a media control device such as a television set-top box that controls content displayed on a television display. A mobile user device or remote control equipped with a microphone can be used to receive speech input for the virtual assistant. The user's intention can be determined from the speech input, and the virtual assistant can perform tasks according to the user's intention, including playing media on a connected television and controlling any other functions of a television set-top box or similar device (e.g., managing video recordings, searching for media content, navigating menus, etc.).
[0049] The virtual assistant interaction can be displayed on a connected television or other display. In one embodiment, media content can be determined based on speech input received from a user. A first, smaller-sized, first user interface can be displayed, including a selectable link to the determined media content. After receiving a media link selection, a second, larger-sized, second user interface can be displayed, including the media content associated with the selection. In another embodiment, the interface used to communicate the virtual assistant interaction can be enlarged or reduced to minimize the amount of space occupied while conveying the desired information.
[0050] In some embodiments, multiple devices associated with multiple displays can be used to communicate information to a user in various ways, as well as to determine user intent from speech input. For example, speech input from a user can be received at a first device having a first display. User intent can be determined from the speech input based on content displayed on the first display. Media content can be determined based on the user intent, and the media content can be played on a second device associated with a second display.
[0051] Television display content can also be used as context input to determine user intent from speech input. For example, speech input can be received from a user, including a query associated with content displayed on the television display. User intent for a query can be determined based on the content displayed on the television display, as well as a browsing history of media content on the television display (e.g., to disambiguate a query based on a character in a TV program currently playing). Results for the query can then be displayed based on the determined user intent.
[0052] In some embodiments, virtual assistant query recommendations can be provided to the user (e.g., informing the user of available commands, recommending interesting content, etc.). For example, media content can be displayed on the display, and an input requesting a virtual assistant query recommendation can be received from the user. Based on the media content displayed on the display and the browsing history of the media content displayed on the display, a virtual assistant query recommendation can be determined (e.g., recommending a query related to a TV program currently being played). The recommended virtual assistant query can then be displayed on the display.
[0053] Using a virtual assistant to control television user interactions in accordance with various embodiments discussed herein can provide an efficient and enjoyable user experience. Using a virtual assistant capable of receiving natural language queries or commands can make user interactions with a media control device intuitive and simple. If desired, available features can be recommended to the user, including meaningful query recommendations based on playing content, which can help the user learn control capabilities. Furthermore, intuitive verbal commands can provide easy access to available media. However, it should be understood that many other advantages can be achieved according to various embodiments discussed herein.
[0054] FIG. 1 illustrates an exemplary system 100 for controlling television user interaction using a virtual assistant. Controlling television user interaction as discussed herein is merely an example of controlling media based on one type of display technology and is used for reference purposes; it should be understood that the concepts discussed herein can be used to generally control any media content interaction, such as on any of a variety of devices and associated displays (e.g., monitors, laptop displays, desktop computer displays, mobile user device displays, projector displays, etc.). Thus, the term "television" can refer to any type of display associated with any of a variety of devices. Furthermore, the terms "virtual assistant," "digital assistant," "intelligent automated assistant," or "automated digital assistant" can refer to any information processing system that interprets spoken and / or textual natural language input to infer user intent and perform actions based on the inferred user intent. For example, to perform an action based on the inferred user intent, the system can perform one or more of the following: That is, identifying a task flow having steps and parameters designed to fulfill the inferred user intent, inputting specific requests from the inferred user intent into the task flow, executing the task flow by invoking programs, methods, services, APIs, etc., and generating an output response to the user in an audible (e.g., verbal) and / or visual form.
[0055] A virtual assistant can accept user requests, at least in part, in the form of natural language commands, requests, statements, narratives, and / or queries. Typically, a user request requires either an informational answer or task performance by the virtual assistant (e.g., displaying a particular medium). A satisfactory response to a user request can include providing the requested informational answer, performing the requested task, or a combination of the two. For example, a user can ask a virtual assistant a question such as, "Where am I now?" Based on the user's current location, the virtual assistant can reply, "You're in Central Park." A user can also request task performance, for example, "Please remind me to call my mother at 4 p.m. today." In response, the virtual assistant can confirm the request and then create an appropriate reminder item in the user's electronic schedule. During the performance of a requested task, the virtual assistant can interact with the user in a continuous dialog, sometimes exchanging information multiple times over an extended period of time. There are many other ways to interact with a virtual assistant to request information or the performance of various tasks. In addition to providing verbal responses and taking programmed actions, the virtual assistant can also provide other visual or audio responses (e.g., as text, alerts, music, videos, animations, etc.). Further, as described herein, an exemplary virtual assistant can control the playback of media content (e.g., play videos on a television) and display information on a display.
[0056] One example of a virtual assistant is described in applicant's U.S. Utility Patent Application No. 12 / 987,982 for "Intelligent Automated Assistant," filed January 10, 2011, the entire disclosure of which is incorporated herein by reference.
[0057] As shown in FIG. 1, in some embodiments, a virtual assistant can be implemented according to a client-server model. The virtual assistant can include a client-side portion running on the user device 102 and a server-side portion running on the server system 110. The client-side portion can also run on the television set-top box 104 in conjunction with the remote control 106. The user device 102 can include any electronic device such as a mobile phone (e.g., a smartphone), a tablet computer, a portable media player, a desktop computer, a laptop computer, a PDA, or a wearable electronic device (e.g., digital glasses, a wristband, a watch, a brooch, an armband, etc.). The television set-top box 104 can include any media control device such as a cable box, a satellite box, a video player, a video streaming device, a digital video recorder, a game system, a DVD player, a Blu-ray Disc player, a combination of such devices, etc. The television set-top box 104 can be connected to a display 112 and a speaker 111 via a wired or wireless connection. Display 112 (with or without speakers 111) can be any type of display, such as a television display, a monitor, a projector, etc. In some embodiments, television set-top box 104 can be connected to an audio system (e.g., an audio receiver), and speakers 111 can be separate from display 112. In other embodiments, display 112, speakers 111, and television set-top box 104 can be combined into a single device, such as a smart television, with advanced processing and network connectivity capabilities. In such embodiments, the functionality of television set-top box 104 can be implemented as an application on the combined device.
[0058] In some embodiments, the television set-top box 104 can function as a media control center for multiple types and sources of media content. For example, the television set-top box 104 can enable user access to live television (e.g., over-the-air television, satellite television, or cable television). Accordingly, the television set-top box 104 can include a cable tuner, a satellite tuner, etc. In some embodiments, the television set-top box 104 can also record television programs for later time-shifted viewing. In other embodiments, the television set-top box 104 can provide access to one or more streaming media services, such as cable-delivered, on-demand television programs, videos, and music (e.g., from various free, paid, and subscription-based streaming services) and Internet-delivered television programs, videos, and music. In still other embodiments, the television set-top box 104 can enable playback or display of media content from any other source, such as displaying photos from a mobile user device, playing videos from a coupled storage device, playing music from a coupled music player, etc. Additionally, the television set-top box 104 may also include various other combinations of the media control features discussed herein, as desired.
[0059] The user device 102 and the television set-top box 104 can communicate with the server system 110 via one or more networks 108, which may include the Internet, an intranet, or any other wired or wireless public or private network. Additionally, the user device 102 can communicate with the television set-top box 104 directly via the network 108 or by any other wired or wireless communication mechanism (e.g., Bluetooth, Wi-Fi, radio frequency, infrared transmission, etc.). As illustrated, the remote control 106 can communicate with the television set-top box 104 using any type of communication, such as a wired connection or any type of wireless communication (e.g., Bluetooth, Wi-Fi, radio frequency, infrared transmission, etc.), including via the network 108. In some embodiments, a user can interact with the television set-top box 104 via interface elements (e.g., buttons, microphones, cameras, joysticks, etc.) integrated into the user device 102, the remote control 106, or the television set-top box 104. For example, speech input including media-related queries or commands for the virtual assistant can be received at the user device 102 and / or the remote control 106, and the speech input can be used to perform media-related tasks on the television set-top box 104. Similarly, haptic commands for controlling media on the television set-top box 104 can be received at the user device 102 and / or the remote control 106 (as well as from other devices not shown). Thus, various functions of the television set-top box 104 can be controlled in a variety of ways, providing users with multiple options for controlling media content from multiple devices.
[0060] The client-side portion of the exemplary virtual assistant running on the user device 102 and / or television set-top box 104 using the remote control 106 can provide client-side functionality, such as user-responsive input and output processing and communication with the server system 110. The server system 110 can provide server-side functionality to any number of clients residing on each user device 102 or each television set-top box 104.
[0061] The server system 110 can include one or more virtual assistant servers 114, which can include a client-facing I / O interface 122, one or more processing modules 118, data and model storage 120, and an I / O interface 116 to external services. The client-facing I / O interface 122 can enable client-facing input and output processing for the virtual assistant server 114. The one or more processing modules 118 can utilize the data and model storage 120 to determine the user's intent based on natural language input and can perform task execution based on the estimated user intent. In some embodiments, the virtual assistant server 114 can communicate with external services 124, such as telephone services, calendar services, information services, messaging services, navigation services, television program services, streaming media services, etc., via the network(s) 108 for task completion or information gathering. The I / O interface 116 to external services can enable such communication.
[0062] Server system 110 may be implemented on one or more standalone data processing devices or a distributed network of computers, and in some embodiments, server system 110 may employ various virtual devices and / or services of third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing and / or infrastructure resources of server system 110.
[0063] Although the functionality of the virtual assistant is shown in FIG. 1 as including both a client-side portion and a server-side portion, in some embodiments, the assistant's functionality (or, generally, speech recognition and media control) can be implemented as a standalone application installed on a user device, television set-top box, smart TV, etc. Furthermore, in different embodiments, the distribution of functionality between the client and server portions of the virtual assistant can vary. For example, in some embodiments, the client running on the user device 102 or television set-top box 104 can be a thin client that provides only user-responsive input and output processing functions and delegates all other functions of the virtual assistant to a back-end server.
[0064] 2 shows a block diagram of an exemplary user device 102, according to various embodiments. The user device 102 may include a memory interface 202, one or more processors 204, and a peripherals interface 206. One or more communication buses or signal lines may couple various components within the user device 102 together. The user device 102 may further include various sensors, subsystems, and peripheral devices coupled to the peripherals interface 206. The sensors, subsystems, and peripheral devices may collect information and / or enable various functions of the user device 102.
[0065] For example, the user device 102 may include a motion sensor 210, a light sensor 212, and a proximity sensor 214 to enable orientation, light, and proximity sensing functions, which are coupled to the peripherals interface 206. One or more other sensors 216, such as a positioning system (e.g., a GPS receiver), a temperature sensor, a biometric sensor, a gyroscope, a compass, an accelerometer, and the like, may also be connected to the peripherals interface 206 to enable related functions.
[0066] In some embodiments, a camera subsystem 220 and an optical sensor 222 may be utilized to enable camera functions such as taking pictures and recording video clips. Various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters may be included to enable communication functions via one or more wired and / or wireless communication subsystems 224. An audio subsystem 226 may be coupled to a speaker 228 and a microphone 230 to enable voice-enabled functions such as voice recognition, voice duplication, digital recording, and telephone functions.
[0067] In some embodiments, user device 102 may further include an I / O subsystem 240 coupled to peripherals interface 206. I / O subsystem 240 may include a touchscreen controller 242 and / or other input controller(s) 244. Touchscreen controller 242 may be coupled to a touchscreen 246. Touchscreen 246 and touchscreen controller 242 may detect contact and its movement or cessation using any of a number of touch-sensing technologies, such as, for example, capacitive, resistive, infrared, surface acoustic wave technology, proximity sensor arrays, etc. Other input controller(s) 244 may be coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses.
[0068] In some embodiments, user device 102 may further include a memory interface 202 coupled to memory 250. Memory 250 may include any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, a portable computer diskette (magnetic), a random access memory (RAM) (magnetic), a read-only memory (ROM) (magnetic), an erasable programmable read-only memory (EPROM) (magnetic), a portable optical disk such as a CD, CD-R, CD-RW, DVD, DVD-R, or DVD-RW, or a flash memory such as a compact flash card, a secure digital card, a USB memory device, a memory stick, or the like. In some embodiments, the non-transitory computer-readable storage medium of memory 250 may be used to store instructions (e.g., to perform some or all of the various processes described herein) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system capable of fetching instructions from and executing those instructions from an instruction execution system, apparatus, or device. In other embodiments, instructions (e.g., for performing some or all of the various processes described herein) may be stored on a non-transitory computer-readable storage medium of server system 110, or may be split between a non-transitory computer-readable storage medium of memory 250 and a non-transitory computer-readable storage medium of server system 110. In the context of this document, a "non-transitory computer-readable storage medium" may be any medium that contains or is capable of storing a program for use by or in connection with an instruction execution system, apparatus, or device.
[0069] In some embodiments, memory 250 can store an operating system 252, a communications module 254, a graphical user interface module 256, a sensor processing module 258, a telephony module 260, and applications 262. Operating system 252 can include instructions for handling basic system services and instructions for performing hardware-dependent tasks. Communications module 254 can enable communication with one or more additional devices, one or more computers, and / or one or more servers. Graphical user interface module 256 can enable graphic user interface processing. Sensor processing module 258 can enable sensor-related processes and functions. Telephony module 260 can enable telephony-related processes and functions. Application module 262 can enable various functionality of user applications, such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions.
[0070] As described herein, memory 250 also stores client-side virtual assistant instructions (e.g., in virtual assistant client module 264), as well as various user data 266 (e.g., user-specific vocabulary data, setting data, and / or the user's electronic address book, to-do list, shopping list, television program preferences, etc.), for example, to provide client-side functionality of the virtual assistant. User data 266 can also be used to support the virtual assistant or to perform speech recognition for any other application.
[0071] In various embodiments, the virtual assistant client module 264 can accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces (e.g., I / O subsystem 240, audio subsystem 226, etc.) of the user device 102. The virtual assistant client module 264 can also provide output in audio (e.g., speech output), visual, and / or tactile form. For example, the output can be provided as voice, sound, an alert, a text message, a menu, a graphic, a video, an animation, a vibration, and / or a combination of two or more of the above. In operation, the virtual assistant client module 264 can communicate with the virtual assistant server using the communication subsystem 224.
[0072] In some embodiments, the virtual assistant client module 264 can collect additional information from the surrounding environment of the user device 102 using various sensors, subsystems, and peripheral devices to establish a context associated with the user, the current user interaction, and / or the current user input. Such context can also include information from other devices, such as information from the television set-top box 104. In some embodiments, the virtual assistant client module 264 can provide context information or a subset thereof along with the user input to the virtual assistant server to help infer the user's intent. The virtual assistant can also use the context information to determine how to prepare and deliver output to the user. Furthermore, the context information can be used by the user device 102 or the server system 110 to support accurate speech recognition.
[0073] In some embodiments, context information associated with a user input can include sensor information such as lighting, environmental noise, ambient temperature, images or videos of the surrounding environment, and distances to other objects. Context information can also include information associated with the physical state of the user device 102 (e.g., device orientation, device location, device temperature, power level, speed, acceleration, motion patterns, cellular signal strength, etc.) or the software state of the user device 102 (e.g., running processes, installed programs, past and present network activity, background services, error logs, resource usage, etc.). Context information can also include information associated with the state of connected devices or other devices associated with the user (e.g., media content displayed by the television set-top box 104, media content available to the television set-top box 104, etc.). Any of these types of context information can be provided to the virtual assistant server 114 as context information associated with the user input (or can be used by the user device 102 itself).
[0074] In some embodiments, the virtual assistant client module 264 can selectively provide information stored on the user device 102 (e.g., user data 266) in response to a request from the virtual assistant server 114 (or can be used by the user device 102 itself when performing speech recognition and / or virtual assistant functions). The virtual assistant client module 264 can also elicit additional input from the user via a natural language dialog or other user interface upon request by the virtual assistant server 114. The virtual assistant client module 264 can pass additional input to the virtual assistant server 114 to assist the virtual assistant server 114 in inferring intent and / or achieving the user's intent expressed in the user request.
[0075] In various embodiments, memory 250 may include additional or fewer instructions. Furthermore, various functions of user device 102 may be implemented in hardware and / or firmware, including in the form of one or more signal processing circuits and / or application specific integrated circuits.
[0076] FIG. 3 is a block diagram of an exemplary television set-top box 104 in a system 300 for controlling television user interaction. System 300 can include a subset of elements of system 100. In some embodiments, system 300 can perform only certain functions and can work with other elements of system 100 to perform other functions. For example, elements of system 300 can handle certain media control functions (e.g., playback of locally stored media, recording functions, channel tuning, etc.) without interacting with server system 110, and system 300 can handle other media control functions (e.g., playback of remotely stored media, downloading media content, processing certain virtual assistant queries, etc.) in cooperation with server system 110 and other elements of system 100. In other embodiments, elements of system 300 can perform functions of a larger system 100, including accessing external services 124 over a network. It should be understood that functionality may be distributed between local devices and remote server devices in various other ways.
[0077] 3, in one embodiment, television set-top box 104 may include a memory interface 302, one or more processors 304, and a peripheral interface 306. One or more communication buses or signal lines may couple the various components within television set-top box 104 together. Television set-top box 104 may further include various sensors, subsystems, and peripheral devices coupled to peripheral interface 306. The subsystems and peripheral devices may collect information and / or enable various functions of television set-top box 104.
[0078] For example, the television set-top box 104 may include a communications subsystem 324. Various communications ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters may be included to enable communications capabilities via one or more wired and / or wireless communications subsystems 324.
[0079] In some embodiments, the television set-top box 104 may further include an I / O subsystem 340 coupled to the peripherals interface 306. The I / O subsystem 340 may include an audio / video output controller 370. The audio / video output controller 370 may be coupled to the display 112 and the speakers 111, or in some cases may provide audio and video output (e.g., via an audio / video port, wireless transmission, etc.). The I / O subsystem 340 may further include a remote controller 342. The remote controller 342 may be communicatively coupled to the remote control 106 (e.g., via a wired connection, Bluetooth, Wi-Fi, etc.). The remote control 106 may include a microphone 372 for capturing audio input (e.g., speech input from a user), button(s) 374 for capturing tactile input, and a transceiver 376 for enabling communication with the television set-top box 104 via the remote controller 342. The remote control 106 may also include other input mechanisms such as a keyboard, joystick, touchpad, etc. The remote control 106 may further include output mechanisms such as lights, a display, speakers, etc. Input received at the remote control 106 (e.g., user utterances, button presses, etc.) may be communicated to the television set-top box 104 via a remote controller 342. The I / O subsystem 340 may further include other input controller(s) 344. The other input controller(s) 344 may be coupled to other input / control devices 348, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses.
[0080] In some embodiments, television set-top box 104 may further include a memory interface 302 coupled to memory 350. Memory 350 may include any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, a portable computer diskette (magnetic), a random access memory (RAM) (magnetic), a read-only memory (ROM) (magnetic), an erasable programmable read-only memory (EPROM) (magnetic), a portable optical disk such as a CD, CD-R, CD-RW, DVD, DVD-R, or DVD-RW, or a flash memory such as a compact flash card, a secure digital card, a USB memory device, a memory stick, or the like. In some embodiments, the non-transitory computer-readable storage medium of memory 350 may be used to store instructions (e.g., to perform some or all of the various processes described herein) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system capable of fetching instructions from and executing those instructions from an instruction execution system, apparatus, or device. In other embodiments, instructions (e.g., for performing some or all of the various processes described herein) may be stored on a non-transitory computer-readable storage medium of server system 110, or may be split between a non-transitory computer-readable storage medium of memory 350 and a non-transitory computer-readable storage medium of server system 110. In the context of this document, a "non-transitory computer-readable storage medium" may be any medium that contains or is capable of storing a program for use by or in connection with an instruction execution system, apparatus, or device.
[0081] In some embodiments, memory 350 can store an operating system 352, a communications module 354, a graphical user interface module 356, an on-device media module 358, an off-device media module 360, and applications 362. Operating system 352 can include instructions for handling basic system services and for performing hardware-dependent tasks. Communications module 354 can enable communication with one or more additional devices, one or more computers, and / or one or more servers. Graphical user interface module 356 can enable graphic user interface processing. On-device media module 358 can enable storage and playback of media content stored locally on television set-top box 104 and other locally available media content (e.g., tuning cable channels). Off-device media module 360 can enable streaming playback or download of media content stored remotely (e.g., on a remote server, on user device 102, etc.). Application module 362 can enable various functionalities of user applications, such as electronic messaging, web browsing, media processing, games, and / or other processes and functions.
[0082] As described herein, memory 350 also stores client-side virtual assistant instructions (e.g., in virtual assistant client module 364), as well as various user data 366 (e.g., user-specific vocabulary data, setting data, and / or the user's electronic address book, to-do list, shopping list, television program preferences, etc.), for example, to provide client-side functionality of the virtual assistant. User data 366 can also be used to support the virtual assistant or to perform speech recognition for any other application.
[0083] In various embodiments, the virtual assistant client module 364 can accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces (e.g., I / O subsystem 340, etc.) of the television set-top box 104. The virtual assistant client module 364 can also provide output in audio form (e.g., speech output), visual form, and / or tactile form. For example, the output can be provided as voice, sound, an alert, a text message, a menu, a graphic, a video, an animation, a vibration, and / or a combination of two or more of the above. In operation, the virtual assistant client module 364 can communicate with the virtual assistant server using the communication subsystem 324.
[0084] In some embodiments, the virtual assistant client module 364 can collect additional information from the surrounding environment of the television set-top box 104 using various sensors, subsystems, and peripheral devices to establish a context associated with the user, the current user interaction, and / or the current user input. Such context can also include information from other devices, such as information from the user device 102. In some embodiments, the virtual assistant client module 364 can provide context information or a subset thereof along with the user input to the virtual assistant server to help infer the user's intent. The virtual assistant can also use the context information to determine how to prepare and deliver output to the user. Furthermore, the context information can be used by the television set-top box 104 or the server system 110 to support accurate speech recognition.
[0085] In some embodiments, the context information associated with the user input can include sensor information such as lighting, environmental noise, ambient temperature, distance to other objects, etc. Context information can further include information associated with the physical state of the television set-top box 104 (e.g., device location, device temperature, power level, etc.) or the software state of the television set-top box 104 (e.g., running processes, installed programs, past and present network activity, background services, error logs, resource usage, etc.). Context information can also include information associated with the state of connected devices or other devices associated with the user (e.g., content displayed by the user device 102, playable content on the user device 102, etc.). Any of these types of context information can be provided to the virtual assistant server 114 as context information associated with the user input (or can be used by the television set-top box 104 itself).
[0086] In some embodiments, the virtual assistant client module 364 can selectively provide information stored in the television set-top box 104 (e.g., user data 366) in response to a request from the virtual assistant server 114 (or can be used by the television set-top box 104 itself when performing speech recognition and / or virtual assistant functions). The virtual assistant client module 364 can also elicit additional input from the user via a natural language dialog or other user interface upon request by the virtual assistant server 114. The virtual assistant client module 364 can pass additional input to the virtual assistant server 114 to assist the virtual assistant server 114 in inferring intent and / or achieving the user's intent expressed in the user request.
[0087] In various embodiments, memory 350 may include additional or fewer instructions. Furthermore, various functions of television set-top box 104 may be implemented in hardware and / or firmware, including in the form of one or more signal processing circuits and / or application specific integrated circuits.
[0088] It should be understood that system 100 and system 300 are not limited to the components and configurations shown in Figures 1 and 3, and similarly, user device 102, television set-top box 104, and remote control 106 are not limited to the components and configurations shown in Figures 2 and 3. System 100, system 300, user device 102, television set-top box 104, and remote control 106 can all include fewer or other components in multiple configurations according to various embodiments.
[0089] Throughout this disclosure, references to a "system" may include system 100, system 300, or one or more elements of either system 100 or system 300. For example, a typical system referred to herein may include at least a television set-top box 104 that receives user input from a remote control 106 and / or a user device 102.
[0090] 4A-4E illustrate an exemplary speech input interface 484 that may be displayed on a display (e.g., display 112) to convey speech input information to a user. In one embodiment, speech input interface 484 may be displayed over video 480, which may include any moving image or paused video. For example, video 480 may include live television, playback video, streaming movies, playback of recorded programs, etc. Speech input interface 484 may be configured to occupy a minimal amount of space so as not to significantly interfere with the user's viewing of video 480.
[0091] In one embodiment, the virtual assistant can be triggered to listen for speech input containing a command or query (or to start recording the speech input for subsequent processing or start real-time processing of the speech input). For example, a user can press a physical button on the remote control 106, a user can press a physical button on the user device 102, a user can press a virtual button on the user device 102, a user can utter a trigger phrase recognizable by a constantly listening device (e.g., uttering "Hey Assistant" to start listening for commands), or a user can make a gesture detectable by a sensor (e.g., making a gesture in front of the camera). Listening can be triggered in various ways, including instructions such as: In another embodiment, a user can press and hold a physical button on the remote control 106 or the user device 102 to start listening. In yet another embodiment, a user can press and hold a physical button on the remote control 106 or the user device 102 while uttering a query or command, and can release the button when finished. Similarly, various other instructions can be received to start receiving speech input from the user.
[0092] In response to receiving an instruction to listen for speech input, a speech input interface 484 may be displayed. FIG. 4A shows a notification area 482 extending upward from a lower portion of the display 112. Upon receiving an instruction to listen for speech input, the speech input interface 484 may be displayed in the notification area 482 and may be animated to slide upward from the lower edge of the viewing area of the display 112, as shown. FIG. 4B shows the speech input interface 484 after it has slid upward and appeared. The speech input interface 484 may be configured to occupy a minimal amount of space at the bottom of the display 112 to avoid interfering with the video 480. In response to receiving an instruction to listen for speech input, a readiness confirmation 486 may be displayed. The readiness confirmation 486 may include a microphone symbol, as shown, or may include any other image, icon, animation, or symbol that communicates that the system (e.g., one or more elements of the system 100) is ready to capture speech input from a user.
[0093] When the user begins speaking, listen confirmation 487, shown in FIG. 4C , may be displayed to confirm that the system is capturing speech input. In some examples, listen confirmation 487 may be displayed in response to receiving speech input (e.g., capturing speech). In other examples, ready confirmation 486 may be displayed for a predetermined period of time (e.g., 500 milliseconds, 1 second, 3 seconds, etc.), after which listen confirmation 487 may be displayed. Listen confirmation 487 may include a waveform symbol as shown, or may include an active waveform animation that moves (e.g., changes frequency) in response to user speech. In other examples, listen confirmation 487 may include any other image, icon, animation, or symbol that communicates that the system is capturing speech input from the user.
[0094] Upon detecting that the user has finished speaking (e.g., based on a pause, speech interpretation indicating the end of the query, or any other endpoint detection method), a processing confirmation 488, shown in FIG. 4D , may be displayed to confirm that the system has completed capturing the speech input and is processing the speech input (e.g., interpreting the speech input, determining the user's intent, and / or performing an associated task). The processing confirmation 488 may include an hourglass symbol as shown, or any other image, icon, animation, or symbol that communicates that the system is processing the captured speech input. In another example, the processing confirmation 488 may include a spinning circle or an animation of colored / glowing dots moving around a circle.
[0095] After interpreting the captured speech input as text (or in response to successfully converting the speech input to text), a command receipt confirmation 490 and / or a phonetic transcription 492, shown in FIG. 4E , may be displayed to confirm that the system received and interpreted the speech input. The phonetic transcription 492 may include a phonetic transcription of the received speech input (e.g., "What sporting events are currently airing?"). In some embodiments, the phonetic transcription 492 may be animated to slide up from the bottom of the display 112 and may be displayed in the position shown in FIG. 4E for a period of time (e.g., a few seconds), and then the phonetic transcription may slide up to the top of the speech input interface 484 and disappear from view (e.g., as if text were scrolling up and eventually disappearing from view). In other embodiments, the phonetic transcription may not be displayed, and the user's command or query may be processed and the associated task may be performed without displaying the phonetic transcription (e.g., a simple channel change may be performed immediately without displaying the phonetic transcription of the user's utterance).
[0096] In other embodiments, phonetic transcription of speech can be performed in real time as the user speaks. As the words are being phonetically transcribed, they can be displayed in speech input interface 484. For example, they can be displayed next to listen confirmation 487. After the user has finished speaking, command receipt confirmation 490 can be briefly displayed, after which a task associated with the user's command can be performed.
[0097] Additionally, in other embodiments, command acknowledgment 490 may convey information about the command that was received and understood. For example, for a simple request to change to a different channel, a logo or number associated with that channel may be displayed briefly (e.g., for a few seconds) as the channel is changed as command acknowledgment 490. In another embodiment, for a request to pause a video (e.g., video 480), a pause symbol (e.g., two vertical parallel bars) may be displayed as command acknowledgment 490. The pause symbol may remain on the display, for example, until the user performs another action (e.g., a play command to resume playback). Similarly, for any other command, a symbol, logo, animation, etc. (e.g., symbols for rewind, fast forward, stop, play, etc.) may be displayed. Thus, command acknowledgment 490 may be used to convey command-specific information.
[0098] In some embodiments, the speech input interface 484 may be hidden after receiving a user query or command. For example, the speech input interface 484 may be animated to slide downward until it disappears from the bottom of the display 112. The speech input interface 484 may be hidden when no further information needs to be displayed to the user. For example, for common or simple commands (e.g., change to channel 10, change to a sports channel, play, pause, fast forward, rewind, etc.), the speech input interface 484 may be hidden immediately after acknowledging receipt of the command, and the associated task(s) may be performed immediately. While various embodiments herein illustrate and describe interfaces at the bottom or top edge of the display, it should be understood that any of a variety of interfaces may be located in other locations around the display. For example, the speech input interface 484 may emerge from a side edge of the display 112, in the center of the display 112, at a corner of the display 112, etc. Similarly, various other interface embodiments described herein may be arranged in a variety of different locations on the display and in a variety of different orientations. Additionally, although the various interfaces described herein are shown as being opaque, any of the various interfaces may be transparent, or in some cases, allow an image (a blurred image or the entire image) to be viewed through the interface (e.g., overlaying interface content over media content without completely obscuring the underlying media content).
[0099] In other embodiments, the results of the query can be displayed in the speech input interface 484 or in a different interface. FIG. 5 shows an exemplary media content interface 510 on the video 480, displaying exemplary results of the phonetic query of FIG. 4E. In some embodiments, the results of the virtual assistant query can include media content instead of or in addition to text content. For example, the results of the virtual assistant query can include television programs, videos, music, etc. Some results can include media that is immediately available for playback, while other results can include media that may be available for purchase, etc.
[0100] As shown, the media content interface 510 can be larger than the speech input interface 484. In one embodiment, the speech input interface 484 can be a smaller first size to accommodate speech input information, and the media content interface 510 can be a larger second size to accommodate query results, where the media content interface 510 can include text, still images, and moving images. In this way, the size of the interface for conveying virtual assistant information can be scaled according to the content to be conveyed, thereby limiting intrusion into the screen area (e.g., minimizing occlusion of other content such as video 480).
[0101] As illustrated, the media content interface 510 may include a selectable video link 512 (as a result of a virtual assistant query), a selectable text link 514, and an additional content link 513. In some embodiments, a link can be selected by using a remote control (e.g., remote control 106) to navigate a focus, cursor, or the like to a particular element and selecting it. In other embodiments, a link can be selected using a voice command to the virtual assistant (e.g., watch the soccer game, view details about the basketball game, etc.). The selectable video link 512 may include a still image or a moving image and may be selectable to play an associated video. In one embodiment, the selectable video link 512 may include a playback video of the associated video content. In another embodiment, the selectable video link 512 may include a live feed of a television channel. For example, the selectable video link 512 may include a live feed of a soccer game on a sports channel as a result of a virtual assistant query about a sporting event currently being broadcast on television. The selectable video links 512 may also include any other videos, animations, images, etc. (e.g., a triangular play symbol.) Additionally, the links 512 may link to any type of media content, such as movies, television shows, sporting events, music, etc.
[0102] The selectable text link 514 can include text content associated with the selectable video link 512 or can include a text representation of the results of the virtual assistant query. In one embodiment, the selectable text link 514 can include a description of the media resulting from the virtual assistant query. For example, the selectable text link 514 can include the name of a television program, a movie title, a description of a sporting event, the name or number of a television channel, etc. In one embodiment, selection of the text link 514 can play the associated media content. In another example, selection of the text link 514 can provide additional details about the media content or other virtual assistant query results. The additional content link 513 can link to and display additional results of the virtual assistant query.
[0103] While an example of a particular media content is shown in FIG. 5, it should be understood that the results of a virtual assistant query for media content may include any type of media content. For example, media content that may be returned as a result of a virtual assistant may include videos, television programs, music, television channels, etc. Furthermore, in some embodiments, a category filter may be provided in any of the interfaces described herein to enable a user to filter search or query results or displayed media options. For example, a selectable filter may be provided to filter results by type (e.g., movies, music albums, books, television programs, etc.). In other embodiments, a selectable filter may include a genre descriptor or content descriptor (e.g., comedy, interviews, specific programs, etc.). In yet other embodiments, a selectable filter may include time (e.g., this week, last week, last year, etc.). It should be understood that a filter may be provided in any of the various interfaces described herein to enable a user to filter results based on categories associated with the displayed content (e.g., filtering by type when the media results have various types, filtering by genre when the media results have various genres, filtering by time when the media results have various times, etc.).
[0104] In other embodiments, media content interface 510 may include a paraphrase of the query in addition to the media content results. For example, a paraphrase of the user's query may be displayed above the media content results (above selectable video links 512 and selectable text links 514). In the embodiment of FIG. 5, such a paraphrase of the user's query may include, "Several sporting events are currently airing." Similarly, other text introducing the media content results may be displayed.
[0105] In some examples, after displaying any interface, including interface 510, a user can begin capturing additional speech input with a new query (which may or may not be related to a previous query). A user query may include a command to act on an interface element, such as a command to select video link 512. In another example, a user's speech may include a query associated with displayed content, such as displayed menu information, a playing video (e.g., video 480), etc. A response to such a query may be determined based on the displayed information (e.g., displayed text) and / or metadata associated with the displayed content (e.g., metadata associated with the playing video). For example, a user may ask a question regarding a media result displayed in an interface (e.g., interface 510), and the metadata associated with that media may be searched to provide an answer or result. Such an answer or result may then be provided in another interface or within the same interface (e.g., any of the interfaces discussed herein).
[0106] As described above, in one embodiment, additional details about the media content can be displayed in response to selection of text link 514. FIGS. 6A and 6B show an exemplary media details interface 618 on video 480 after selection of text link 514. In one embodiment, when providing additional detailed information, media content interface 510 can expand to media details interface 618, as illustrated by interface expansion transition 616 in FIG. 6A. Specifically, as shown in FIG. 6A, the size of the selected content can be expanded, and additional text information can be provided by expanding the interface upward on display 112 to occupy more of the screen real estate. The interface can expand to accommodate the additional detailed information desired by the user. In this manner, the size of the interface can scale with the amount of content desired by the user, thereby minimizing intrusion into screen real estate while still conveying the desired content.
[0107] 6B shows the details interface 618 after it has been fully expanded. As shown, the details interface 618 can be larger than either the media content interface 510 or the speech input interface 484 to accommodate the desired detailed information. The details interface 618 can include detailed media information 622, including various detailed information associated with the media content or another result of the virtual assistant query. The detailed media information 622 can include the program title, program description, program broadcast time, channel, episode synopsis, movie description, actor names, character names, sporting event participants, producer names, director names, or any other detailed information associated with the results of the virtual assistant query.
[0108] In one embodiment, the details interface 618 may include a selectable video link 620 (or another link for playing media content), where the selectable video link 620 may include a larger version of the corresponding selectable video link 512. Thus, the selectable video link 620 may include a still image or a moving image and may be selectable to play an associated video. The selectable video link 620 may include a playback video of the associated video content, a live feed of a television channel (e.g., a live feed of a soccer game on a sports channel), etc. The selectable video link 620 may also include any other video, animation, image, etc. (e.g., a triangular play symbol).
[0109] As described above, a video may be played in response to selection of a video link, such as video link 620 or video link 512. FIGS. 7A and 7B illustrate exemplary media transition interfaces that may be displayed in response to selection of a video link (or other command to play video content). As illustrated, video 480 may be replaced with video 726. In one embodiment, video 726 may be expanded to overlay or cover video 480, as illustrated by interface expansion transition 724 in FIG. 7A. The result of the transition may include expanded media interface 728 in FIG. 7B. As with other interfaces, the size of expanded media interface 728 may be sufficient to provide the user with desired information, which here includes expanding to fill display 112. Thus, expanded media interface 728 may be larger than any other interface, as the desired information may include playing media content across the entire display. Although not shown, in some embodiments, descriptive information may be temporarily overlaid on video 726 (e.g., along the bottom of the screen). Such descriptive information may include the name of an associated program, video, channel, etc. The descriptive information can then be hidden from view (eg, after a few seconds).
[0110] 8A-8B illustrate an exemplary speech input interface 836 that may be displayed on display 112 to convey speech input information to a user. In one embodiment, speech input interface 836 may be displayed on menu 830. Menu 830 may include various media options 832, and speech input interface 836 may similarly be displayed on any other type of menu (e.g., a content menu, a category menu, a control menu, a setup menu, a program menu, etc.). In one embodiment, speech input interface 836 may be configured to occupy a relatively large amount of screen real estate of display 112. For example, speech input interface 836 may be larger than speech input interface 484 discussed above. In one embodiment, the size of the speech input interface to be used (e.g., either smaller interface 484 or larger interface 836) may be determined based on background content. For example, a smaller speech input interface (e.g., interface 484) may be displayed when the background content includes moving images. On the other hand, when the background content includes a still image (e.g., paused video) or a menu, for example, a larger speech input interface (e.g., interface 836) may be displayed. In this manner, when the user is watching video content, a smaller speech input interface may be displayed, minimizing intrusion into screen real estate, but when the user is navigating a menu or viewing paused video or other still images, a larger speech input interface may be displayed, occupying additional real estate and thereby conveying more information or having a more significant effect. Similarly, other interfaces discussed herein may be sized differently based on the background content.
[0111] As discussed above, the virtual assistant can be triggered to listen for speech input containing a command or query (or to start recording the speech input for subsequent processing or start real-time processing of the speech input). For example, a user can press a physical button on the remote control 106, a user can press a physical button on the user device 102, a user can press a virtual button on the user device 102, a user can utter a trigger phrase recognizable by a constantly listening device (e.g., uttering "Hey Assistant" to start listening for commands), or a user can make a gesture detectable by a sensor (e.g., making a gesture in front of the camera). Listening can be triggered in various ways, including instructions such as: In another embodiment, the user can press and hold a physical button on the remote control 106 or the user device 102 to start listening. In yet another embodiment, the user can press and hold a physical button on the remote control 106 or the user device 102 while uttering a query or command, and can release the button when finished. Similarly, various other instructions can be received to start receiving speech input from the user.
[0112] In response to receiving an instruction to listen for speech input, a speech input interface 836 can be displayed above the menu 830. FIG. 8A shows a large notification area 834 extending upward from the bottom portion of the display 112. Upon receiving an instruction to listen for speech input, the speech input interface 836 can be displayed in the large notification area 834 and animated to slide upward from the bottom edge of the viewing area of the display 112, as shown. In some examples, as the overlying interface is displayed (e.g., in response to receiving an instruction to listen for speech input), a background menu, paused video, still image, or other background content can be collapsed in the z-direction (as if moving further into the display 112) and / or moved backward. The background interface collapse transition 831 and associated inward arrow illustrate how background content (e.g., the menu 830) can be collapsed (making the displayed menu, image, text, etc. smaller). This can provide the visual effect of background content appearing to move away from the user, out of the way of the new foreground interface (e.g., interface 836). Figure 8B shows a minimized background interface 833 that includes a minimized (smaller) version of menu 830. As shown, the minimized background interface 833 (which may include a border) can appear farther away from the user, while still ceding focus to foreground interface 836. Background content (including background video content) in any of the other embodiments discussed herein can similarly shrink and / or move backward in the z-direction as overlapping interfaces are displayed.
[0113] 8B shows speech input interface 836 after it has been slid upward to reveal it. As discussed above, various confirmations may be displayed while receiving speech input. While not shown here, speech input interface 836 may similarly display larger versions of ready confirmation 486, listen confirmation 487, and / or process confirmation 488, similar to speech input interface 484 discussed above with reference to FIGS. 4B, 4C, and 4D, respectively.
[0114] As shown in FIG. 8B , a command acknowledgment 838 may be displayed (as with the smaller command acknowledgment 490 discussed above) to confirm that the system received and interpreted the speech input. Also, a phonetic transcription 840 may be displayed, which may include a phonetic transcription of the received speech input (e.g., "What's the weather in New York?"). In some examples, the phonetic transcription 840 may be animated to slide up from the bottom of the display 112 and may be displayed in the position shown in FIG. 8B for a period of time (e.g., a few seconds), and then may slide up to the top of the speech input interface 836 and disappear from view (e.g., as if text were scrolling up and eventually disappearing from view). In other examples, the phonetic transcription may not be displayed, and the user's command or query may be processed and the associated task may be performed without displaying the phonetic transcription.
[0115] In other embodiments, phonetic transcription of utterances can be performed in real time as the user speaks. Words can be displayed in speech input interface 836 as they are being phonetically transcribed. For example, words can be displayed next to a larger version of listen confirmation 487 discussed above. After the user has finished speaking, command receipt confirmation 838 can be briefly displayed, after which a task associated with the user's command can be performed.
[0116] Additionally, in other embodiments, command acknowledgment 838 may convey information about the command that was received and understood. For example, in the case of a simple request to tune to a particular channel, the command acknowledgment 838 may briefly display (e.g., for a few seconds) a logo or number associated with that channel when the channel is tuned. In another embodiment, in the case of a request to select a displayed menu item (e.g., one of media options 832), the command acknowledgment 838 may display an image associated with the selected menu item. Thus, command acknowledgment 838 may be used to convey command-specific information.
[0117] In some embodiments, the speech input interface 836 may be hidden after receiving a user query or command. For example, the speech input interface 836 may be animated to slide downward until it disappears from the bottom of the display 112. The speech input interface 836 may be hidden when no further information needs to be displayed to the user. For example, for common or simple commands (e.g., change to channel 10, change to a sports channel, play that movie, etc.), the speech input interface 836 may be hidden immediately after acknowledging receipt of the command, and the associated task(s) may be performed immediately.
[0118] In other examples, the results of the query can be displayed within the speech input interface 836 or in a different interface. FIG. 9 shows a virtual assistant result interface 942 on an example menu 830 (specifically, on a reduced background interface 833) with example results of the phonetized query of FIG. 8B. In some examples, the results of the virtual assistant query can include a text answer, such as text answer 944. The results of the virtual assistant query can also include media content that addresses the user's query, such as content associated with a selectable video link 946 and a purchase link 948. In particular, in this example, a user can ask for weather information about a specific location in New York. The virtual assistant can provide a text answer 944 that directly answers the user's query (e.g., indicates that the weather looks good and provides temperature information). Instead of or in addition to the text answer 944, the virtual assistant can provide a selectable video link 946 along with a purchase link 948 and associated text. Additionally, the media associated with links 946 and 948 can provide a response to the user's query. Here, the media associated with links 946 and 948 may include a 10-minute clip of weather information for a particular location (specifically, a five-day forecast for New York from a television channel called the Weather Channel).
[0119] In one embodiment, the clip addressing the user's query may include a time cue portion of previously broadcast content (which may be available from a recording or streaming service). In one embodiment, the virtual assistant can identify such content by searching for detailed information about available media content (e.g., including metadata about the recorded broadcast along with detailed timing information or details about the streaming content) based on the user intent associated with the speech input. In some embodiments, the user may not have access to certain content or may not have a subscription to certain content. In such cases, the user may be encouraged to purchase the content, such as via a purchase link 948. Selecting the purchase link 948 or video link 946 may automatically collect or charge the cost of the content from the user's account.
[0120] FIG. 10 shows an example process 1000 for controlling television interactions using a virtual assistant and displaying associated information using different interfaces. At block 1002, speech input from a user can be received. For example, the speech input can be received at a user device 102 or a remote control 106 of the system 100. In some embodiments, the speech input (or a data representation of some or all of the speech input) can be transmitted to and received by the server system 110 and / or the television set-top box 104. In response to the user starting to receive the speech input, various notifications can be displayed on a display (such as the display 112). For example, as discussed above with reference to FIGS. 4A-4E, a readiness confirmation, a listening confirmation, a processing confirmation, and / or a command receipt confirmation can be displayed. Additionally, the received user speech input can be phonetically transcribed, and the phonetically transcribed can be displayed.
[0121] Referring again to process 1000 of FIG. 10, at block 1004, media content can be determined based on speech input. For example, media content that addresses a user query directed at the virtual assistant (e.g., by searching available media content) can be determined. For example, media content related to phonetic transcription 492 of FIG. 4E ("What sporting events are currently being broadcast?") can be determined. Such media content can include live sporting events displayed on one or more television channels available for viewing by the user.
[0122] At block 1006, a first user interface of a first size comprising selectable media links may be displayed. For example, as shown in Figure 5, a media content interface 510 comprising selectable video links 512 and selectable text links 514 may be displayed on display 112. As discussed above, media content interface 510 may be a smaller size to avoid interfering with background video content.
[0123] At block 1008, a selection of one of the links may be received. For example, a selection of one of links 512 and / or 514 may be received. At block 1010, a second, larger, sized second user interface may be displayed that includes media content associated with the selection. As shown in FIG. 6B, for example, a details interface 618 that includes a selectable video link 620 and detailed media information 622 may be displayed on the display 112. As discussed above, the details interface 618 may be larger in size to convey desired additional detailed media information. Similarly, as shown in FIG. 7B, selection of the video link 620 may display an expanded media interface 728 that includes a video 726. As discussed above, the expanded media interface 728 may be larger in size to still provide the user with desired media content. In this manner, the various interfaces discussed herein may be sized to accommodate desired content (including expanding to a larger sized interface or shrinking to a smaller sized interface), potentially while occupying limited screen real estate. Thus, the process 1000 can be used to control television interactions using a virtual assistant and display associated information using different interfaces.
[0124] In another example, a larger size interface can be displayed over a control menu rather than over background video content. For example, speech input interface 836 can be displayed over menu 830 as shown in FIG. 8B, and assistant results interface 942 can be displayed over menu 830 as shown in FIG. 9, while smaller media content interface 510 can be displayed over video 480 as shown in FIG. 5. In this manner, the size of the interface (e.g., the amount of screen real estate it occupies) can be determined at least in part by the type of background content.
[0125] FIG. 11 illustrates exemplary television media content on a user device 102, which may include a mobile phone, tablet computer, remote control, etc., with a touchscreen 246 (or another display). FIG. 11 illustrates an interface 1150 including a TV listing with a plurality of television programs 1152. The interface 1150 may correspond to a particular application on the user device 102, such as a television control application, a television content listing application, an Internet application, etc. In some examples, content displayed on the user device 102 (e.g., on the touchscreen 246) can be used to determine user intent from speech input related to the content, and the user intent can be used to play or display the content on another device and display (e.g., on the television set-top box 104 and the display 112 and / or speaker 111). For example, content displayed on the interface 1150 on the user device 102 can be used to disambiguate a user request and determine user intent from speech input, and the determined user intent can then be used to play or display media via the television set-top box 104.
[0126] FIG. 12 shows an exemplary television control using a virtual assistant. FIG. 12 shows an interface 1254, which can include a virtual assistant interface formatted as a conversational dialog between the assistant and the user. For example, the interface 1254 can include an assistant greeting 1256 that prompts the user to make a request. Subsequent received user utterances can then be phonetically transcribed, such as phonetically transcribed user utterance 1258, and the conversational exchange can be displayed. In some embodiments, the interface 1254 can appear on the user device 102 in response to a trigger (such as a button press, a key phrase, or the like) that initiates the receipt of speech input.
[0127] In one example, a user request to play content via television set-top box 104 (e.g., on display 112 and speaker 111) may include an ambiguous reference to something displayed on user device 102. For example, phonetically transcribed user utterance 1258 includes a reference to “the” soccer game (“Turn on the soccer game.”). The particular soccer game desired may be unclear from the speech input alone. However, in some examples, content displayed on user device 102 can be used to disambiguate the user request and determine user intent. In one example, content displayed on user device 102 before the user makes the request (e.g., before interface 1254 appears on touchscreen 246) can be used to determine user intent (as can content appearing in interface 1254, such as previous queries and results). In the illustrated example, content displayed in interface 1150 of FIG. 11 can be used to determine user intent from the command to turn on “the” soccer game. The TV listing for television programs 1152 includes a variety of different programs, one of which is titled "Soccer" broadcast on channel 5. The appearance of the soccer listing can be used to determine the user's intention from the mention of "that" soccer game. In particular, the user's reference to "that" soccer game can be interpreted as the soccer program appearing in the TV listing of interface 1150. Thus, the virtual assistant can play the particular soccer game desired by the user (for example, by tuning the television set-top box 104 to the appropriate channel and displaying the game).
[0128] In other examples, a user may reference television programs (e.g., Channel 8 programs, news, drama programs, advertisements, first program, etc.) displayed in interface 1150 in various other ways, and user intent may similarly be determined based on the displayed content. It should be appreciated that metadata associated with the displayed content (e.g., descriptions of the TV programs), fuzzy matching techniques, synonym matching, etc., may further be used in conjunction with the displayed content to determine user intent. For example, the term “advertising” may be matched to the description “television shopping” (e.g., using synonym and / or fuzzy matching techniques) to determine user intent from a request to view “advertising.” Similarly, descriptions of specific TV programs may be analyzed in determining user intent. For example, the term “legal” may be identified in a detailed description of a courtroom drama, and user intent may be determined from a user request to watch a “legal” program based on the detailed description associated with the content displayed in interface 1150. Thus, the displayed content and data associated therewith may be used to disambiguate user requests and determine user intent.
[0129] FIG. 13 illustrates exemplary photo and video content on a user device 102, which may include a mobile phone, tablet computer, remote control, etc., equipped with a touchscreen 246 (or another display). FIG. 13 illustrates an interface 1360 including a list of photos and videos. Interface 1360 may correspond to a particular application on the user device 102, such as a media content application, a file navigation application, a storage application, a remote storage management application, a camera application, etc. As shown, interface 1360 may include videos 1362, a photo album 1364 (e.g., a group of multiple photos), and photos 1366. As discussed above with reference to FIGS. 11 and 12 , content displayed on the user device 102 can be used to determine user intent from speech input related to the content. The user intent can then be used to play or display the content on another device and display (e.g., on the television set-top box 104 and display 112 and / or speaker 111). For example, content displayed on interface 1360 on user device 102 can be used to disambiguate user requests and determine user intent from speech input, and the determined user intent can then be used to play or display media via television set-top box 104.
[0130] FIG. 14 shows an exemplary media display control using a virtual assistant. FIG. 14 shows an interface 1254, which can include a virtual assistant interface formatted as a conversational dialog between the assistant and the user. As shown, the interface 1254 can include an assistant greeting 1256 that prompts the user to make a request. User utterances can then be phonetically transcribed within the dialog as shown by the example of FIG. 14. In some embodiments, the interface 1254 can appear on the user device 102 in response to a trigger (such as a button press, a key phrase, or the like) that initiates the receipt of speech input.
[0131] In one example, a user request to play media content or display media via television set-top box 104 (e.g., on display 112 and speaker 111) may include an ambiguous reference to something displayed on user device 102. For example, phonetically transcribed user utterance 1468 includes a reference to “the” video (“Display the video.”). The specific video referenced may be unclear from speech input alone. However, in some examples, content displayed on user device 102 may be used to disambiguate the user request and determine user intent. In one example, content displayed on user device 120 before the user makes the request (e.g., before interface 1254 appears on touchscreen 246) may be used to determine user intent (as may content appearing in interface 1254, such as previous queries and results). In the example of user utterance 1468, content displayed in interface 1360 of FIG. 13 may be used to determine user intent from the command to display “the” video. The list of photos and videos in interface 1360 includes a variety of different photos and videos, including video 1362, photo album 1354, and photo 1366. Because only one video appears in interface 1360 (e.g., video 1362), the appearance of video 1362 in interface 1360 can be used to determine the user's intent from the utterance of "that" video. In particular, the user's reference to "that" video can be interpreted as video 1362 (titled "Graduation Video") appearing in interface 1360. Thus, the virtual assistant can play video 1362 (for example, by having video 1362 transmitted from user device 102 or remote storage to television set-top box 104 and starting playback).
[0132] In another example, the phonetically transcribed user utterance 1470 includes a reference to “the” album (“Play a slideshow of that album.”). The specific album referenced may be unclear from the speech input alone. The content displayed on the user device 102 can be used again to disambiguate the user request. Specifically, the content displayed in interface 1360 of FIG. 13 can be used to determine user intent from the command to play a slideshow of “the” album. The list of photos and videos in interface 1360 includes photo album 1354. The appearance of photo album 1364 in interface 1360 can be used to determine the user intent from the utterance of “the” album. Specifically, the user's reference to “the” album can be interpreted as photo album 1364 (titled “Graduation Album”) appearing in interface 1360. Thus, in response to user utterance 1470, the virtual assistant can display a slideshow containing photos from the photo album 1364 (for example, by sending photos from the photo album 1364 from the user device 102 or remote storage to the television set-top box 104 and starting a slideshow of photos).
[0133] In yet another example, the phonetically transcribed user utterance 1472 includes a reference to a “latest” photo (“Show the latest photo on the kitchen TV.”). The specific photo being referenced may be unclear from the speech input alone. The content displayed on the user device 102 can again be used to disambiguate the user request. In particular, the content displayed in interface 1360 of FIG. 13 can be used to determine user intent from the command to display the “latest” photo. The list of photos and videos in interface 1360 includes two individual photos 1366. The appearance of the photos 1366 in interface 1360 (particularly the order in which the photos 1366 appear within the interface) can be used to determine the user intent from the utterance of the “latest” photo. In particular, the user's reference to the “latest” photo can be interpreted as the photo 1366 (dated June 21, 2014) appearing at the bottom of interface 1360. Thus, in response to user utterance 1472, the virtual assistant can display the latest photo 1366 on the interface 1360 (for example, by transmitting the latest photo 1366 from the user device 102 or remote storage to the television set-top box 104 and displaying the photo).
[0134] In other examples, a user may view the media content displayed in interface 1360 in a variety of other ways (e.g., most recent two photos, all video news, all photos, graduation album, graduation video, photos since June 21, etc.), and user intent may similarly be determined based on the displayed content. It should be appreciated that metadata associated with the displayed content (e.g., timestamp, location, information, title, description, etc.), fuzzy matching techniques, synonym matching, etc., may further be used in conjunction with the displayed content to determine user intent. Thus, the displayed content and its associated data may be used to disambiguate the user request and determine user intent.
[0135] It should be appreciated that any type of displayed content in any application interface of any application can be used in determining user intent. For example, speech input can reference images displayed on a web page in an internet browser application, and the displayed web page content can be analyzed to identify the desired image. Similarly, speech input by title, genre, artist, band name, etc. can reference music tracks in a music listing in a music application, and the displayed content (and in some embodiments, associated metadata) in the music application can be used to determine user intent from the speech input. The determined user intent can then be used to display or play media via another device, such as television set-top box 104, as discussed above.
[0136] In some embodiments, user identification, user authentication, and / or device authentication may be employed to determine whether media control can be granted, determine media content available for display, determine access permissions, etc. For example, it may be determined whether a particular user device (e.g., user device 102) is authorized to control media, for example, on the television set-top box 104. The user device may be authenticated based on registration, pairing, trust determination, passcode, secret question, system settings, etc. In response to determining that a particular user device is authorized, an attempt to control the television set-top box 104 may be permitted (e.g., media content may be played in response to determining that the requesting device is authorized to control media). In contrast, media control commands or requests from unauthorized devices may be ignored and / or users of such devices may be prompted to register their devices for use in controlling the particular television set-top box 104.
[0137] In another example, a particular user can be identified, and personal information associated with the user can be used to determine the user intent of the request. For example, a user can be identified based on speech input, such as by voice recognition using the user's voiceprint. In some examples, a user can utter a particular phrase, which can be analyzed for speech recognition. In other examples, a speech input request directed to a virtual assistant can be analyzed using speech recognition to identify the speaker. A user can also be identified based on the source of the speech input sample (e.g., on the user's personal device 102). A user can also be identified based on a password, passcode, menu selection, etc. Speech input received from the user can then be interpreted based on the identified user's personal information. For example, the user intent of a speech input can be determined based on previous requests from the user, media content owned by the user, media content stored on the user's device, user preferences, user settings, user demographics (e.g., languages spoken, etc.), user profile information, user payment method, or various other personal information associated with a particular identified user. For example, based on the personal information, ambiguity can be avoided for speech input that references a favorites list, and the user's personal favorites list can be identified. Based on the user identification, speech inputs that refer to "my" photos, "my" videos, "my" programs, etc. can similarly be disambiguated to accurately identify photos, videos, and programs associated with the user (e.g., photos stored on a personal user device, etc.) Similarly, speech inputs that request the purchase of content can be disambiguated to determine that the identified user's payment method (versus another user's payment method) should be changed for the purchase.
[0138] In some examples, user authentication can be used to determine whether a user is allowed to access media content, whether they are allowed to purchase media content, etc. For example, voice recognition can be used to verify a particular user's identity (e.g., using their voiceprint) to allow the user to make purchases using their payment method. Similarly, a password or the like can be used to authenticate a user and enable purchases. In another example, voice recognition can be used to verify a particular user's identity to determine whether the user is allowed to view a particular program (e.g., a program with a particular parental guideline rating, a movie with a particular age rating, etc.). For example, a child's request for a particular program can be denied based on voice recognition indicating that the requester is not an authorized user (e.g., a parent) who is allowed to view such content. In other examples, voice recognition can be used to determine whether a user has access to particular subscription content (e.g., restricting access to content on premium channels based on voice recognition). In some examples, a user can utter a particular phrase, which can be analyzed for voice recognition. In other examples, a spoken input request directed to a virtual assistant can be analyzed using voice recognition to identify the speaker. Thus, certain media content may be played in response to initially determining that a user is authenticated in any of a variety of ways.
[0139] FIG. 15 shows an exemplary virtual assistant interaction with results on a mobile user device and a media display device. In some embodiments, the virtual assistant can provide information and control on two or more devices, such as the user device 102 and the television set-top box 104. Further, in some embodiments, the same virtual assistant interface used for control and information on the user device 102 can be used to issue a request to control media on the television set-top box 104. Thus, the virtual assistant system can determine whether the results should be displayed on the user device 102 or the television set-top box 104. In some embodiments, when employing the user device 102 to control the television set-top box 104, the intrusion of the virtual assistant interface on the display (e.g., display 112) associated with the television set-top box 104 can be minimized by displaying information on the user device 102 (e.g., on the touch screen 246). In other embodiments, the virtual assistant information can be displayed only on the display 112, or the virtual assistant information can be displayed on both the user device 102 and the display 112.
[0140] In some embodiments, a determination can be made as to whether the results of the virtual assistant query should be displayed directly on the user device 102 or on the display 112 associated with the television set-top box 104. In one embodiment, in response to determining that the user intent of the query includes a request for information, an information response can be displayed on the user device 102. In another example, in response to determining that the user intent of the query includes a request to play media content, media content in response to the query can be played via the television set-top box 104.
[0141] FIG. 15 shows a virtual assistant interface 1254 illustrating an example of a conversational dialog between a virtual assistant and a user. An assistant greeting 1256 can prompt the user to make a request. In the first query, the phonetized user utterance 1574 (which can also be typed or entered in other ways) includes a request for an information answer associated with the displayed media content. In particular, the phonetized user utterance 1574 inquires, for example, about who is playing in a soccer game that may be displayed on an interface on the user device 102 (e.g., listed in interface 1150 of FIG. 11) or on the display 112 (e.g., listed in interface 510 of FIG. 5 or played as video 726 on the display 112 of FIG. 7B). Based on the displayed media content, the user intent of the phonetized user utterance 1574 can be determined. For example, the particular soccer game in question can be identified based on the content displayed on the user device 102 or the display 112. The user intent of the phonetically transcribed user utterance 1574 may include obtaining an informational answer detailing the teams playing in the soccer match identified based on the displayed content. In response to determining that the user intent includes a request for an informational answer, the system may determine to display the response within interface 1254 of FIG. 15 (as opposed to on display 112). In some examples, the response to the query may be determined based on metadata associated with the displayed content (e.g., based on a description of the soccer match in a television listing). Thus, as shown, in interface 1254, an assistant response 1576 identifying teams Alpha and Theta as playing in the match may be displayed on touchscreen 246 of user device 102. Thus, in some examples, the informational response may be displayed within interface 1254 on user device 102 based on determining that the query includes an information request.
[0142] However, the second query in interface 1254 includes a media request. Specifically, the phonetized user utterance 1578 requests that the displayed media content be changed to “games.” The user intent of the phonetized user utterance 1578 can be determined based on the displayed content, such as the games listed in interface 510 of FIG. 5 (e.g., to identify which game the user desires), the games listed in interface 1150 of FIG. 11 , or games referenced in a previous query (e.g., in the phonetized user utterance 1574). Thus, the user intent of the phonetized user utterance 1578 can include changing the displayed content to a particular game (here, a soccer match between Teams Alpha and Theta). In one example, the game can be displayed on the user device 102. However, in other examples, the game can be displayed via the television set-top box 104 based on a query including a request to play media content. In particular, in response to determining that the user intent includes a request to play media content, the system determines to display media content results on the display 112 via the television set-top box 104 (as opposed to within the interface 1254 of FIG. 15). In some embodiments, the virtual assistant may display a response or paraphrase (e.g., "Change to a soccer game") in the interface 1254 or on the display 112 confirming the intended action.
[0143] FIG. 16 shows an exemplary virtual assistant interaction with media results on a media display device and a mobile user device. In some embodiments, the virtual assistant can provide access to media on both the user device 102 and the television set-top box 104. Furthermore, in some embodiments, the same virtual assistant interface used for media on the user device 102 can be used to issue a request for media on the television set-top box 104. Thus, the virtual assistant system can determine whether the results should be displayed on the user device 102 via the television set-top box 104 or on the display 112.
[0144] In some embodiments, a determination may be made as to whether to display media on device 102 or display 112 based on the media result format, user preferences, default settings, explicit commands in the request itself, etc. For example, the format of the media results for a query may be used to determine on which device to display the media results by default (e.g., without specific instructions). Television programs may be better suited to display on a television, large-format videos may be better suited to display on a television, thumbnail photos may be better suited to display on a user device, small-format web videos may be better suited to display on a user device, and various other media formats may be better suited to display on either a relatively large television screen or a relatively small user device display. Thus, in response to a determination (e.g., based on the media format) that media content should be displayed on a particular display, the media content may be displayed on that particular display by default.
[0145] FIG. 16 shows the virtual assistant interface 1254 with examples of queries related to playing or displaying media content. The assistant greeting 1256 can prompt the user to make a request. In the first query, the phonetized user utterance 1680 includes a request to display a soccer game. Similar to the example discussed above, the user intent of the phonetized user utterance 1680 can be determined based on the displayed content, such as the games listed in the interface 510 of FIG. 5, the games listed in the interface 1150 of FIG. 11, or the games referenced in a previous query (e.g., to identify which game the user wants). Thus, the user intent of the phonetized user utterance 1680 can include, for example, displaying a particular soccer game that may be broadcast on television. In response to determining that the user intent includes a request to display media formatted for television (e.g., a soccer game that will be broadcast on television), the system can automatically determine to display the desired media on the display 112 via the television set-top box 104 (as opposed to on the user device 102 itself). The virtual assistant system can then tune the television set-top box 104 to the soccer game and display it on the display 112 (for example, by performing the necessary tasks and / or sending appropriate commands).
[0146] However, in the second query, the phonetically transcribed user utterance 1682 includes a request to display photos of the team's players (e.g., photos of "Team Alpha"). Similar to the example described above, a user intent of the phonetically transcribed user utterance 1682 can be determined. The user intent of the phonetically transcribed user utterance 1682 can include performing a search (e.g., a web search) for photos associated with "Team Alpha" and displaying the resulting photos. In response to determining that the user intent includes a request to display media that can be presented in thumbnail format, or media associated with a web search or other unspecified media without a specific format, the system can automatically determine to display the desired media results on the touch screen 246 in the interface 1254 of the user device 102 (as opposed to displaying the resulting photos on the display 112 via the television set-top box 104). For example, as shown, thumbnail photos 1684 can be displayed in the interface 1254 on the user device 102 in response to the user's query. Thus, the virtual assistant system can, by default, display media in a particular format, or media that can be presented in a particular format (e.g., in a group of thumbnails), on the user device 102.
[0147] It should be appreciated that in some embodiments, the soccer match referenced in user utterance 1680 may be displayed on user device 102, and photos 1684 may be displayed on display 112 via television set-top box 104. However, a default device for display may be automatically determined based on the media format, thereby simplifying media commands for the user. In other embodiments, a default device for displaying requested media content may be determined based on user preferences, default settings, the device most recently used to display content, voice recognition identifying the user, devices associated with the user, etc. For example, a user may set preferences or a default configuration to display certain types of content (e.g., videos, slideshows, television programs, etc.) on display 112 via television set-top box 104, and other types of content (e.g., thumbnails, photos, web videos, etc.) on touchscreen 246 of user device 102. Similarly, preferences or default configurations may be set to respond to certain queries by displaying content on one device or the other. In another embodiment, all content may be displayed on the user device 102 unless the user commands otherwise.
[0148] In yet other embodiments, a user query can include a command to display content on a particular display. For example, user utterance 1472 in FIG. 14 includes a command to display a photo on a kitchen TV. As a result, the system can display the photo on a TV display associated with the user's kitchen, as opposed to displaying the photo on user device 102. In other embodiments, the user can instruct which display device to use in various other ways (e.g., on a TV, a large screen, a living room, a bedroom, their tablet, their phone, etc.). Thus, the display device to use to display media content results of a virtual assistant query can be determined in various different ways.
[0149] FIG. 17 illustrates exemplary proximity-based media device control. In some embodiments, a user may have multiple televisions and television set-top boxes in the same home or on the same network. For example, a home may have a television and set-top box set in the living room, another set in the bedroom, and another set in the kitchen. In other embodiments, multiple set-top boxes may be connected to the same network, such as a shared network in an apartment or office building. While a user may pair, connect, or possibly authenticate a remote control 106 and user device 102 for a particular set-top box to prevent unauthorized access, in other embodiments, a remote control and / or user device may be used to control more than one set-top box. A user may, for example, use a single user device 102 to control set-top boxes in the bedroom, living room, and kitchen. A user may also, for example, use a single user device 102 to control not only their own set-top box in their own apartment, but also neighboring set-top boxes in neighboring apartments (e.g., to share content from the user device 102 with neighbors, such as displaying a slideshow of photos stored on the user device 102 on a neighbor's TV). Because a user may control multiple different set-top boxes using a single user device 102, the system may determine which of the multiple set-top boxes to send a command to. Similarly, because a home may be equipped with multiple remote controls 106 capable of operating multiple set-top boxes, the system may similarly determine which of the multiple set-top boxes to send a command to.
[0150] In one embodiment, device proximity can be used to determine which of multiple set-top boxes a command should be sent to (or on which display the requested media content should be displayed) on a nearby TV. Proximity can be determined between the user device 102 or remote control 106 and each of multiple set-top boxes. The issued command can then be sent to the closest set-top box (or the requested media content can be displayed on the closest display). Proximity can be determined (or at least approximated) in any of a variety of ways, such as time-of-flight measurements (e.g., using radio frequencies), Bluetooth LE, electronic ping signals, proximity sensors, sound travel measurements, etc. The measured or estimated distances can then be compared, and a command can be issued to the closest device (e.g., the closest set-top box).
[0151] FIG. 17 illustrates a multi-device system 1790 including a first set-top box 1792 with a first display 1786 and a second set-top box 1794 with a second display 1788. In one example, a user can issue a command from a user device 102 to display media content (e.g., without necessarily specifying where or on which device). A distance 1795 to the first set-top box 1792 and a distance 1796 to the second set-top box 1794 can then be determined (or estimated). As shown, distance 1796 can be greater than distance 1795. Based on proximity, a command from the user device 102 can be issued to the first set-top box 1792, which is the closest device and most likely to match the user's intent. In some examples, a single remote control 106 can be used to control more than one set-top box. Based on proximity, a desired device to control at a given time can be determined. A distance 1797 to the second set top box 1794 and a distance 1798 to the first set top box 1792 can then be determined (or estimated). As shown, distance 1798 can be greater than distance 1797. Based on the proximity, a command from the remote control 106 can be issued to the second set top box 1794, which is the closest device and most likely to match the user's intent. For example, the distance measurements can be refreshed periodically or for each command to accommodate for a user moving to a different room and wanting to control a different device.
[0152] It should be understood that a user can specify a different device for a command, overriding proximity in some cases. For example, a list of available display devices can be displayed on the user device 102 (e.g., first display 1786 and second display 1788 listed by setup name, designated room, etc., or first set-top box 1792 and second set-top box 1794 listed by setup name, designated room, etc.). The user can select one of the devices from the list. A command can then be sent to the selected device. A request for media content issued at the user device 102 can then be fulfilled by displaying the desired media on the selected device. In other examples, the user can speak the desired device as part of a verbal command (e.g., show the game on the kitchen TV, change to the cartoon channel in the living room, etc.).
[0153] In yet other examples, a default device for displaying the requested media content may be determined based on status information associated with a particular device. For example, it may be determined whether headphones (or a headset) are attached to the user device 102. In response to determining that headphones are attached to the user device 102 when a request to display the media content is received, the requested content may be displayed on the user device 102 by default (e.g., assuming the user consumes content on the user device 102 rather than on a television). In response to determining that headphones are not attached to the user device 102 when a request to display the media content is received, the requested content may be displayed on either the user device 102 or the television according to any of the various determination methods discussed herein. Similarly, other device status information, such as ambient light around the user device 102 or set-top box 104, the proximity of other devices to the user device 102 or set-top box 104, the orientation of the user device 102 (e.g., landscape orientation may be more likely to show the desired view on the user device 102), the display state of the set-top box 104 (e.g., in sleep mode), the time since the last interaction on a particular device, or any of various other status indicators for the user device 102 and / or set-top box 104, can be used to determine whether the requested media content should be displayed on the user device 102 or the set-top box 104.
[0154] 18 shows an example process 1800 for controlling television interactions using a virtual assistant and multiple user devices. At block 1802, speech input from a user can be received at a first device having a first display. For example, speech input from a user can be received at a user device 102 or a remote control 106 of the system 100. In some embodiments, the first display can include a touch screen 246 of the user device 102 or a display associated with the remote control 106.
[0155] At block 1804, a user's intent may be determined from the speech input based on the content displayed on the first display. For example, content such as television programs 1152 in interface 1150 of FIG. 11 or photos and videos in interface 1360 of FIG. 13 may be analyzed and used to determine the user's intent for the speech input. In some examples, a user may ambiguously refer to content displayed on the first display, and the content shown on the first display may be analyzed to interpret the reference (e.g., to determine the user's intent for "the" video, "the" album, "the" game, etc.), as discussed above with reference to FIGS. 12 and 14 .
[0156] Referring again to process 1800 of Figure 18, at block 1806, media content may be determined based on user intent. For example, specific videos, photographs, photo albums, television programs, sporting events, music tracks, etc. may be identified based on user intent. In the examples of Figures 11 and 12 discussed above, for example, a specific soccer game displayed on channel 5 may be identified based on user intent referencing "the" soccer game displayed in interface 1150 of Figure 11. In the examples of Figures 13 and 14 discussed above, a specific video 1362 titled "Graduation Video," a specific photo album 1364 titled "Graduation Album," or a specific photo 1366 may be identified based on user intent determined from the example speech input of Figure 14.
[0157] 18 , at block 1808, the media content may be displayed on a second device associated with a second display. For example, the determined media content may be played on a display 112 with speakers 111 via a television set-top box 104. Playing the media content may include tuning to a particular television channel, playing a particular video, displaying a photo slideshow, displaying particular photos, playing a particular audio track, etc. on the television set-top box 104 or another device.
[0158] In some embodiments, a determination can be made as to whether a response to a speech input directed to a virtual assistant should be displayed on a first display associated with a first device (e.g., user device 102) or a second display associated with a second device (e.g., television set-top box 104). For example, as discussed above with reference to FIGS. 15 and 16, an information answer or media content suitable for display on a smaller screen can be displayed on the user device 102, while a media response or media content suitable for display on a larger screen can be displayed on the display associated with the set-top box 104. As discussed above with reference to FIG. 17, in some embodiments, the distance between the user device 102 and multiple set-top boxes can be used to determine which set-top box to play media content on or which set-top box to issue a command to. Similarly, various other determinations can be made to provide a convenient and user-friendly experience in which multiple devices can interact.
[0159] In some examples, just as content displayed on user device 102 can be used to inform the interpretation of speech input, as discussed above, content displayed on display 112 can likewise be used to inform the interpretation of speech input. In particular, content displayed on a display associated with television set-top box 104, along with metadata associated with that content, can be used to determine user intent from speech input, disambiguate user queries, respond to queries related to the content, etc.
[0160] FIG. 19 shows an exemplary speech input interface 484 (described above), with a virtual assistant query about the video 480 displayed in the background. In some embodiments, the user query may include a question about the media content displayed on the display 112. For example, the phonetic transcription 1916 includes a query requesting the identification of actresses ("Who are those actresses?"). The content displayed on the display 112 (along with metadata or other descriptive information about the content) can be used to determine the user's intent from the speech input related to the content, as well as to determine a response to the query (including informational responses and media responses that provide the user with media selections). For example, the video 480, a description of the video 480, a list of characters and actors in the video 480, rating information for the video 480, genre information for the video 480, and various other descriptive information associated with the video 480 can be used to disambiguate the user's request and determine a response to the user query. Associated metadata can include, for example, identification information for the characters 1910, 1912, and 1914 (e.g., character names along with the names of the actresses playing the characters). Similarly, metadata for any other content may include a title, description, a list of characters, a list of actors, a list of players, a genre, a producer name, a director name, or a display schedule associated with the content displayed on the display or a viewing history of media content on the display (e.g., recently viewed media).
[0161] In one embodiment, a user query directed to the virtual assistant may include an ambiguous reference to something displayed on the display 112. The phonetic transcription 1916 may include, for example, a reference to "those" actresses ("Who are those actresses?"). The specific actress the user is asking about may be unclear from speech input alone. However, in some embodiments, the content displayed on the display 112 and associated metadata may be used to disambiguate the user request and determine the user's intent. In the illustrated embodiment, the content displayed on the display 112 may be used to determine the user's intent from the reference to "those" actresses. In one embodiment, the television set-top box 104 may identify content to play along with details associated with the content. In this case, the television set-top box 104 may identify the title of the video 480 along with various descriptive content. In other embodiments, television programs, sporting events, or other content may be used in conjunction with associated metadata to determine the user's intent. Furthermore, in any of the various embodiments discussed herein, speech recognition results and intent determination may weight terms associated with the displayed content higher than alternatives. For example, while an actor of an on-screen character appears on-screen (or while a program in which they appear is playing), their actor name can be weighted more highly, thereby enabling accurate speech recognition and intent determination of likely user requests associated with the displayed content.
[0162] In one embodiment, a list of characters and / or actors associated with video 480 can be used to identify all or the most prominent actresses appearing in video 480, which may include actresses 1910, 1912, and 1914. The identified actresses can be returned as possible results (fewer or additional actresses may be included if the metadata resolution is coarser). In another embodiment, metadata associated with video 480 can include identities of the actors and actresses appearing on screen at a given time, and from that metadata, the actresses appearing at the time of the query can be determined (e.g., specifically, actresses 1910, 1912, and 1914 are identified). In yet another embodiment, a facial recognition application can be used to identify actresses 1910, 1912, and 1914 from images displayed on display 112. In still other embodiments, various other metadata associated with video 480, and various other recognition techniques, can be used to identify the user's likely intent in referring to "these" actresses.
[0163] In some embodiments, the content displayed on the display 112 may change during the making of a query and the determination of a response. Thus, the viewing history of the media content may be used to determine the user's intent and the response to the query. For example, if the video 480 moves to a different view (e.g., with another character) before the response to the query is generated, the result of the query may be determined based on the user's view at the time the query was uttered (e.g., the character displayed on the screen when the user initiated the query). In some instances, a user may pause media playback to issue a query, and the content displayed during the pause, along with associated metadata, may be used to determine the user's intent and the response to the query.
[0164] Given the determined user intent, query results can be provided to the user. FIG. 20 shows an example assistant response interface 2018 that includes an assistant response 2020, which can include a response determined from the query of phonetic transcription 1916 in FIG. 19 . Assistant response 2020 can include a list of the name of each actress in video 480 and her associated character, as shown ("Actress Jennifer Jones plays the character Blanche, actress Elizabeth Arnold plays the character Julia, and actress Whitney Davidson plays the character Melissa"). The listed actresses and characters in response 2020 can correspond to characters 1910, 1912, and 1914 appearing on display 112. As described above, in some embodiments, content displayed on display 112 can change during the making of the query and the determination of the response. Thus, response 2020 can include information about content or characters that are no longer appearing on display 112.
[0165] As with other interfaces displayed on display 112, Assistant response interface 2018 can occupy a minimal amount of screen real estate while providing sufficient space to convey the desired information. In some examples, as with other text displayed in an interface on display 112, Assistant response 2020 can scroll up from the bottom of display 112 to the position shown in FIG. 20 , be displayed for a certain amount of time (e.g., a delay based on the length of the response), and then scroll up and out of view. In other examples, after a delay, interface 2018 can be slid downward to disappear from view.
[0166] 21 and 22 show another example of determining user intent and responding to a query based on content displayed on display 112. FIG. 21 shows an exemplary speech input interface 484 illustrating a virtual assistant query regarding media content associated with video 480. In some embodiments, the user query can include a request for media content associated with the media displayed on display 112. For example, a user can request other movies, television programs, sporting events, etc. associated with a particular media, for example, based on characters, actors, genres, etc. For example, phonetic transcription 2122 includes a query requesting other media associated with the actress in video 480, referencing the name of the actress's character in video 480 ("What else has Blanche appeared in?"). Similarly, the content displayed on display 112 (together with metadata or other descriptive information about that content) can be used to determine the user intent from speech input related to that content, as well as to determine the response to the query (either an information response or a resulting response in media selection).
[0167] In some embodiments, a user query directed to the virtual assistant may include ambiguous references using character names, actor names, program names, player names, etc. Without the context of the content displayed on display 112 and its associated metadata, it may be difficult to accurately interpret such references. Phonetic transcription 2122 includes, for example, a reference to a character named "Blanche" in video 480. The specific actress or other person the user is asking about may be unclear from speech input alone. However, in some embodiments, the content displayed on display 112 and associated metadata can be used to disambiguate the user request and determine the user intent. In the illustrated embodiment, the content displayed on display 112 and associated metadata can be used to determine the user intent from the character name "Blanche." In this case, the character list associated with video 480 can be used to determine that "Blanche" may refer to the character "Blanche" in video 480. In another example, detailed metadata and / or facial recognition can be used to determine that a character named "Blanche" appears on the screen (or was visible on the screen at the start of a user's query), and the actress associated with that character can be determined to be the most likely intent of the user's query. For example, characters 1910, 1912, and 1914 can be determined to appear on display 112 (or were visible on display 112 at the start of a user's query), and their associated character names can then be referenced to determine the user intent of a query referencing the character Blanche. The actor list can then be used to identify the actress who plays Blanche, and a search can be conducted to identify other media in which the identified actress appears.
[0168] Given the determined user intent (e.g., interpretation of the character reference "Blanche") and the determined query results (e.g., other media associated with the actress who plays "Blanche"), a response can be provided to the user. FIG. 22 shows an exemplary assistant response interface 2224 including an assistant text response 2226 and a selectable video link 2228, which can respond to the query of phonetic transcription 2122 of FIG. 21. Assistant text response 2226 can include a paraphrase of the user request that introduces selectable video link 2228, as shown. Assistant text response 2226 can also include instructions for disambiguating the user's query (e.g., identifying actress Jennifer Jones, who plays the character Blanche in video 480). Such a paraphrase can confirm to the user that the virtual assistant correctly interpreted the user's query and provided the desired results.
[0169] Assistant response interface 2224 may also include selectable video links 2228. In some embodiments, various types of media content, including movies (e.g., Movie A and Movie B in interface 2224), may be provided as results for a virtual assistant query. The media content displayed as a result of a query may include media that may be available for user consumption (free, for purchase, or as part of a subscription). The user can select the displayed media to view or consume the resulting content. For example, a user may select one of the selectable video links 2228 (e.g., using a remote control, voice command, etc.) to watch one of the other movies starring actress Jennifer Jones. In response to selecting one of the selectable video links 2228, the video associated with the selection may be played, replacing video 480 on display 112. Thus, the displayed media content and associated metadata can be used to determine user intent from speech input, and in some embodiments, playable media may be provided as a result.
[0170] When forming a query, a user may refer to actors, players, characters, locations, teams, sporting event details, movie themes, or various other information associated with the displayed content, and the virtual assistant system may similarly disambiguate such requests and determine user intent based on the displayed content and associated metadata. Similarly, in some embodiments, results may include media recommendations associated with the query, such as movies, television programs, or sporting events associated with the person who is the subject of the query (whether or not the user specifically requests such media content).
[0171] Further, in some embodiments, a user query can include a request for information associated with the media content itself, such as a query about a character, an episode, a movie plot, a previous scene, etc. As in the embodiments discussed above, the displayed content and associated metadata can be used to determine the user intent from such a query and determine a response. For example, a user may request a character description (e.g., "What is Blanche doing in this movie?"). The virtual assistant system can then identify the requested information about the character, such as a character description or cast, from the metadata associated with the displayed content (e.g., "Blanche is one of a group of lawyers and is known as a troublemaker in Hartford."). Similarly, a user may request an episode summary (e.g., "What happened in the last episode?"), and the virtual assistant system can search for and provide a description of the episode.
[0172] In some embodiments, content displayed on display 112 can include menu content, which can likewise be used to determine the user intent of speech input and responses to user queries. Figures 23A-23B are diagrams illustrating example pages of program menu 830. Figure 23A shows a first page of media options 832, and Figure 23B shows a second page of media options 832 (which can include successive next pages of a list of content that spans more than one page).
[0173] In one embodiment, a user request to play content may include ambiguous references in a menu 830 to what is displayed on the display 112. For example, the menu 830 viewed by the user may request to watch "the" football game, "the" basketball game, a vacuum cleaner advertisement, a legal program, etc. The particular program desired may be unclear from speech input alone. However, in some embodiments, content displayed on the device 112 may be used to disambiguate the user request and determine user intent. In the illustrated embodiment, media options in the menu 830 (in some embodiments, together with metadata associated with the media options) may be used to determine user intent from commands that include ambiguous references. For example, "the" football game may be interpreted as a football game on a sports channel. "The" basketball game may be interpreted as a basketball game on a college sports channel. An advertisement for a vacuum cleaner may be interpreted as a television shopping program (e.g., based on metadata associated with a program describing vacuum cleaners). A legal program may be interpreted as a courtroom drama based on metadata associated with the program and / or synonym, fuzzy, or other matching techniques. Thus, the appearance of various media options 832 in menu 830 on display 112 may be used to avoid ambiguity in a user request.
[0174] In some embodiments, the displayed menus can be navigated with a cursor, joystick, arrows, buttons, gestures, etc. In such cases, focus can be indicated for a selected item. For example, the selected item can be indicated with bold, underlined, bordered, larger in size than other menu items, shadowed, reflective, illuminated, and / or any other feature that emphasizes which menu item is selected and has focus. For example, selected media option 2330 in FIG. 23A can have focus as the currently selected media option, indicated by large, underlined type and a border.
[0175] In some embodiments, a request to play or select content or a menu item may include an ambiguous reference to the menu item having focus. For example, the menu 830 being viewed by the user in FIG. 23A may request that “the” program be played (e.g., “Play that program.”). Similarly, the user may request various other commands associated with the menu item having focus, such as play, delete, hide, watch reminder, record, etc. The particular menu item or program desired may be ambiguous from speech input alone. However, content displayed on the device 112 may be used to disambiguate the user request and determine user intent. In particular, the fact that the selected media option 2330 has focus in the menu 830 may be used to identify the desired media subject, either a command referencing “the” program, a subjectless command (e.g., play, delete, hide, etc.), or any other ambiguous command referencing the media content having focus. Thus, the menu item having focus may be used in determining user intent from speech input.
[0176] Similar to the browsing history of media content that may be displayed at the time the user initiated the request but has since elapsed, previously displayed menu or search result content may similarly be used to disambiguate subsequent user requests after navigating to subsequent menu or search result content. For example, FIG. 23B illustrates a second page of menu 830 with additional media options 832. A user may proceed to the second page illustrated in FIG. 23B but may again refer to the content displayed on the first page illustrated in FIG. 23A (e.g., media options 832 shown in FIG. 23A). For example, despite navigating to the second page of menu 830, the user may request to watch "the" football game, "the" basketball game, or a legal program, all of which are media options 832 that were recently displayed on previous pages of menu 830. Although such references may be ambiguous, the recently displayed menu content of the first page of menu 830 may be used to determine user intent. In particular, the recently viewed media options 832 in FIG. 23A may be analyzed to identify the specific soccer game, basketball game, or courtroom drama referenced in the exemplary ambiguous request. In some embodiments, results may be biased based on how recently the content was viewed (e.g., weighting recently viewed pages of results over previously viewed results). In this manner, the browsing history of what was recently viewed on the display 112 may be used to determine user intent. It should be appreciated that any recently viewed content may be used, such as previously viewed search results, previously viewed programs, previously viewed menus, etc. This allows a user to discover a particular view they viewed and revisit something they previously viewed without having to navigate to it.
[0177] In still other examples, various display cues displayed in a menu or results list on device 112 can be used to disambiguate a user request and determine user intent. FIG. 24 shows an exemplary media menu divided into categories, one of which (Movies) has focus. FIG. 24 shows category interface 2440, which can include a carousel-style interface of categorized media options, including TV option 2442, Movies option 2444, and Music option 2446. As shown, the Music category is only partially displayed, and the carousel interface can be shifted to the right (e.g., as indicated by the arrow) to reveal additional content, much like rotating media in a carousel. In the illustrated example, the Movies category has focus, indicated by an underlined title and border, although focus can be indicated in any of a variety of other ways (e.g., by making the category larger, adding a light so that it appears closer to the user than other categories, etc.).
[0178] In some embodiments, a request to play or select content or a menu item may include an ambiguous reference to a menu item in a group of items (such as a category). For example, the category interface 2440 that a user is viewing may request to play a soccer program (“Play a soccer program.”). The particular menu item or program desired may be unclear from the speech input alone. Furthermore, a query may be interpreted as more than one program displayed on the display 112. For example, a request for a soccer program may refer to either a soccer match listed in the TV Programs category or a soccer movie listed in the Movies category. Content displayed on the device 112 (including display cues) may be used to disambiguate the user request and determine user intent. In particular, the fact that the Movies category has focus in the category interface 2440 may be used to identify the particular soccer program desired, which is a soccer movie given focus on the Movies category. Thus, the category of media (or any other group of media) having focus as displayed on the display 112 may be used when determining user intent from the speech input. The user may also make various other requests associated with categories, such as requesting the display of certain category content (eg, show comedy movies, show horror movies, etc.).
[0179] In other embodiments, the user may browse menus or media items displayed on the display 112 in a variety of other ways. Similarly, user intent may be determined based on the displayed content. It should be appreciated that metadata associated with the displayed content (e.g., TV program descriptions, movie descriptions, etc.), fuzzy matching techniques, synonym matching, etc., may further be used in conjunction with the displayed content to determine user intent from speech input. Thus, various forms of user requests, including natural language requests, may be accommodated and user intent may be determined in accordance with various embodiments discussed herein.
[0180] It should be understood that the content displayed on the display 112 may be used alone or in conjunction with the content displayed on the user device 102 or on a display associated with the remote control 106 when determining user intent. Similarly, it should be understood that a virtual assistant query can be received by any of a variety of devices communicatively coupled to the television set-top box 104, and that the content displayed on the display 112 can be used to determine user intent, regardless of which device receives the query. The results of the query can also be displayed on the display 112 or on another display (e.g., on the user device 102).
[0181] Further, in any of the various embodiments discussed herein, the virtual assistant system can navigate menus and select menu options without requiring the user to specifically open the menu and navigate to the menu item. For example, after selecting media content or a menu button, such as selecting the movie option 2444 in FIG. 24, a menu of options may appear. Menu options may include alternatives to simply playing the media, such as setting a reminder to watch the media later, setting a media recording, adding the media to a favorites list, hiding the media from further view, and the like. While a user is viewing content or content with submenu options on a menu, the user can issue virtual assistant commands that may require navigating to a menu or submenu to select. For example, the category interface 2440 viewed by the user in FIG. 24 can issue any menu command associated with the movie option 2444 without manually opening the associated menu. For example, a user may request to add a soccer movie to a favorites list, record the evening news, or set a reminder to watch movie B without constantly navigating the menus or submenus associated with those media options where such commands may be available. Thus, the virtual assistant system can navigate menus and submenus to execute commands on behalf of the user, regardless of whether the menu options for the menus and submenus appear on the display 112. This can simplify user requests and reduce the number of clicks or selections the user must make to achieve the desired menu function.
[0182] FIG. 25 shows an example process 2500 for controlling television interaction using presented media content on a display and a media content viewing history. At block 2502, speech input from a user may be received, including a query associated with content displayed on the television display. For example, the speech input may include a query regarding a character, actor, movie, television program, sporting event, athlete, etc. appearing on display 112 of system 100 (represented by television set-top box 104). For example, phonetic transcription 1916 of FIG. 19 includes a query associated with an actress displayed in video 480 on display 112. Similarly, phonetic transcription 2122 of FIG. 21 includes, for example, a query associated with a character in video 480 displayed on display 112. The speech input may also include a query associated with menu or search content appearing on display 112, such as a query to select a particular menu item or to obtain information about a particular search result. For example, the displayed menu content may include media options 832 of menu 830 in FIGS. 23A and 23B. The displayed menu content may also include a TV option 2442, a movie option 2444, and / or a music option 2446, which appear in the category interface 2440 of FIG.
[0183] Referring back to process 2500 of FIG. 25, at block 2504, user intent for the query may be determined based on the displayed content and media content viewing history. For example, user intent may be determined based on displayed or recently displayed scenes of a television program, sporting event, movie, etc. User intent may also be determined based on displayed or recently displayed menu or search content. The displayed content, along with metadata associated with the content, may also be analyzed to determine user intent. For example, the content illustrated and described with reference to FIGS. 19, 21, 23A, 23B, and 24 may be used alone or in conjunction with metadata associated with the displayed content to determine user intent.
[0184] In block 2506, results of the query can be displayed based on the determined user intent. For example, results similar to the assistant response 2020 in the assistant response interface 2018 of FIG. 20 can be displayed on the display 112. In another example, results can be provided as text and selectable media, such as the assistant text response 2226 and selectable video link 2228 in the assistant response interface 2224 shown in FIG. 22. In yet another example, displaying the results of the query can include displaying or playing selected media content (e.g., playing a selected video on the display 112 via the television set-top box 104). Thus, user intent can be determined from speech input in a variety of ways using the displayed content and associated metadata as context.
[0185] In some embodiments, virtual assistant query recommendations can be provided to the user, for example, to inform the user of available queries, recommend content the user may enjoy, teach the user how to use the system, encourage the user to find additional media content for consumption, etc. In some embodiments, query recommendations can include comprehensive recommendations of possible commands (e.g., find comedy, view a TV guide, search for action movies, turn on closed captions, etc.). In other embodiments, query recommendations can include targeted recommendations related to the displayed content (e.g., add this program to your watchlist, share this program via social media, share the soundtrack of this movie, share books this guest is selling, share movie trailers that the guest is plugged in, etc.), user preferences (e.g., use of closed captions, etc.), content owned by the user, content recorded on the user's device, notifications, alerts, media content viewing history (e.g., recently viewed menu items, recently viewed scenes of a program, recent appearances of actors, etc.), etc. Recommendations can be displayed on any device, including on the display 112 via the television set-top box 104, on the user device 102, or on a display associated with the remote control 106. Additionally, recommendations can be determined based on nearby devices and / or devices in communication with the television set-top box 104 at a particular time (e.g., recommending content from devices of users in a room watching TV at a particular time). In other examples, recommendations can be determined based on a variety of other contextual information, including the time of day, crowd-sourced information (e.g., popular programs being viewed at a given time), live programming (e.g., live sporting events), media content viewing history (e.g., the last few programs viewed, a set of recently viewed search results, a group of recently viewed media options, etc.), or any of a variety of other contextual information.
[0186] FIG. 26 shows an exemplary recommendation interface 2650 including content-based virtual assistant query recommendation 2652. In one embodiment, query recommendations can be provided to an interface such as interface 2650 in response to input received from a user requesting a recommendation. For example, an input requesting a query recommendation can be received from user device 102 or remote control 106. In some embodiments, the input can include a button press, a button double-click, a menu selection, a voice command (e.g., display some recommendations, what can be done, what options are available, etc.), or one received at user device 102 or remote control 106. For example, a user can double-click a physical button on the remote control 106 to request a query recommendation, or can double-click a physical or virtual button on user device 102 when viewing an interface associated with television set-top box 104 to request a query recommendation.
[0187] Recommendation interface 2650 may be displayed over moving images such as video 480, or over any other background content (e.g., a menu, a still image, paused video, etc.). As with the other interfaces discussed herein, recommendation interface 2650 may be animated to slide up from the bottom of display 112, minimizing the amount of space it takes up while still sufficiently conveying the desired information, so as to limit interference with background video 480. In other implementations, the recommendation interface may be larger when the background content is still (e.g., paused video, menu, image, etc.).
[0188] In some embodiments, virtual assistant query recommendations can be determined based on the displayed media content or the viewing history of the media content (e.g., movies, television programs, sporting events, recently viewed programs, recently viewed menus, recently viewed movie scenes, recent scenes from currently airing television episodes, etc.). For example, FIG. 26 shows a content-based recommendation 2652 that can be determined based on the displayed video 480, where the displayed video 480 is displayed in the background and characters 1910, 1912, and 1914 appear on the display 112. Query recommendations can also be determined using metadata associated with the displayed content (e.g., descriptive details of the media content). The metadata can include various information associated with the displayed content, including the program title, character list, actor list, episode description, team roster, team rankings, program synopsis, movie details, plot description, director name, producer name, actor appearance time, sports standings, sports scores, genre, season episode list, related media content, or various other related information. For example, metadata associated with video 480 may include the names of characters 1910, 1912, and 1914, along with the actresses playing those characters. The metadata may also include a description of the plot of video 480, a description of previous or next episodes (if video 480 is a television episode of a series), etc.
[0189] FIG. 26 illustrates various content-based recommendations 2652 that may be presented in recommendation interface 2650 based on video 480 and metadata associated with video 480. For example, a character 1910 in video 480 may be named "Blanche," and the character name may be used to formulate a query recommendation for information about the character Blanche or the actress playing the character (e.g., "Who is the actress playing Blanche?"). The character 1910 may be identified from metadata associated with video 480 (e.g., character listings, actor listings, times associated with actor appearances, etc.). In other embodiments, facial recognition may be used to identify actresses and / or characters appearing on display 112 at a given time. Various other query recommendations may be provided associated with characters in the media itself, such as queries regarding a character's casting, profile, relationships to other characters, etc.
[0190] In another example, an actor or actress appearing on display 112 can be identified (e.g., based on metadata and / or facial recognition) and a query recommendation associated with the actor or actress can be provided. Such a query recommendation can include the role(s) played, film awards, age, other media appearances, career history, relatives, associates, or any of a variety of other details about the actor or actress. For example, character 1914 may be played by an actress named Whitney Davidson, and the actress' name Whitney Davidson can be used to formulate a query recommendation to identify other films, television programs, or other media in which actress Whitney Davidson appears (e.g., "What else has Whitney Davidson been in?").
[0191] In other embodiments, details about the program can be used to formulate query recommendations. Query recommendations can be formulated using episode summaries, plot summaries, episode lists, episode titles, series titles, and the like. For example, a recommendation to explain what happened in the last episode of a television program (e.g., "What happened in the last episode?") can be provided, and the virtual assistant system can respond with an episode summary (and its associated metadata) from a previous episode identified based on the episode currently displayed on display 112. In another embodiment, a recommendation to set up a recording of the next episode can be provided, which is achieved by the system identifying the next episode based on the currently airing episode displayed on display 112. In yet another embodiment, a recommendation to obtain information about the current episode or program appearing on display 112 can be provided, and the program title obtained from the metadata can be used to formulate query recommendations (e.g., "What is this episode of 'Their Show' about?" or "What is 'Their Show' about?").
[0192] In another example, a query recommendation can be formulated using categories, genres, ratings, awards, descriptions, etc. associated with the displayed content. For example, the video 480 can correspond to a television program described as a comedy with a female protagonist. From this information, a query recommendation can be formulated to identify other programs with similar characteristics (e.g., "Find other comedies starring women."). In other examples, recommendations can be determined based on user subscriptions, content available for playback (e.g., content on the television set-top box 104, content on the user device 102, content available for streaming, etc.), etc. For example, potential query recommendations can be filtered based on whether information or media results are available. Query recommendations that may not result in playable media content or information answers can be filtered out, and / or query recommendations with readily available information answers or playable media content can be provided (or weighted more heavily in determining which recommendations to provide). Thus, the displayed content and associated metadata can be used in various ways to determine query recommendations.
[0193] 27 shows an exemplary selection interface 2754 for confirming selection of a recommended query. In some embodiments, a user can select displayed query recommendations by speaking a query, selecting them with a button, navigating to them with a cursor, etc. In response to a selection, a confirmation interface such as selection interface 2754 can temporarily display the selected recommendation. In one embodiment, the selected recommendation 2756 can be animated (e.g., as indicated by an arrow) to move from where the selected recommendation 2756 appears in recommendation interface 2650 to the position shown in FIG. 27 next to command receipt confirmation 490, and other unselected recommendations can be hidden from the display.
[0194] 28A-28B are diagrams illustrating an exemplary virtual assistant answer interface 2862 based on a selected query. In some embodiments, an answer interface, such as answer interface 2862, can display information answers to the selected query. When switching from either recommendation interface 2650 or selection interface 2754, transition interface 2858 can be displayed, as shown in FIG. 28A. In particular, as the next content scrolls upward from the bottom of display 112, previously displayed content in the interface scrolls upward and disappears from the interface. For example, the selected recommendation 2756 can be slid or scrolled upward until it disappears at the top edge of the virtual assistant interface, and assistant results 2860 can be slid or scrolled upward from the bottom of display 112 until it arrives at the position shown in FIG. 28B.
[0195] The answer interface 2862 may include information answers and / or media results in response to the selected query recommendation (or in response to any other query). For example, in response to the selected query recommendation 2756, the assistant result 2860 may be determined and provided. In particular, in response to a request for a previous episode summary, the previous episode may be identified based on the displayed content, and an associated description or summary may be identified and provided to the user. In the illustrated embodiment, the assistant result 2860 may describe a previous episode of the program corresponding to the video 480 on the display 112 (e.g., "In episode 203 of 'Their Show,' Blanche is invited to a college psychology class as a guest speaker. Julia and Melissa show up unannounced and cause a stir."). Information answers and media results may also be presented in any of the other manners discussed herein (e.g., selectable video links), or the results may be presented in a variety of other manners (e.g., by speaking the answer, immediately playing content, showing animations, displaying images, etc.).
[0196] In another embodiment, notifications or alerts can be used to determine virtual assistant query recommendations. FIG. 29 shows a media content notification 2964 (although any notification can be taken into account when determining the recommendation) and a recommendation interface 2650 (which can include some of the same concepts discussed above with reference to FIG. 26) that includes both a notification-based recommendation 2966 and a content-based recommendation 2652. In some embodiments, the content of the notification can be analyzed to identify the name, title, subject, action, etc. related to the associated media. In the illustrated embodiment, notification 2964 includes an alert that notifies the user about alternative media content available for display; in particular, a sporting event is being broadcast live and the content of the game may be of interest to the user (e.g., "Team Theta and Team Alpha are tied with 5 minutes left in the game."). In some embodiments, the notification can be momentarily displayed at the top of display 112. The notification can slide down from the top of display 112 (as indicated by the arrow) to the position shown in FIG. 29, be displayed for a certain period of time, and then slide back up to disappear again at the top of display 112.
[0197] Notifications or alerts can inform users of various information, such as available alternative media content (e.g., alternatives to what may currently be displayed on the display 112), available live television programs, newly downloaded media content, recently added subscription content, recommendations received from friends, receipt of media sent from another device, etc. Notifications can also be personalized based on the media a household or identified user is watching (e.g., identified based on user authentication using account selection, voice recognition, password, etc.). In one embodiment, the system can interrupt programming and display notifications based on likely desired content, such as display notifications 2964 for users who may desire notification content based on their user profile, favorite team(s), favorite sport(ies), browsing history, etc. For example, sporting event scores, game status, time remaining, etc. can be obtained from sports data feeds, news outlets, social media discussions, etc. and used to identify possible alternative media content for notifying the user.
[0198] In other examples, popular media content (e.g., to many users) may be provided via alerts or notifications to recommend alternatives to the currently viewed content (e.g., notifying the user that a popular program or a program in the user's favorite genre has just started or is potentially available for viewing). In the illustrated example, the user may follow one or both of Team Theta and Team Alpha (or may follow soccer or a particular sport, league, etc.). The system may determine that available live content matches the user's preferences (e.g., a game on another channel matches the user's preferences, there is little time left in the game, the score is close). The system may then determine to alert the user via notification 2964 of potentially desired content. In some examples, the user may select notification 2964 (or a link within notification 2964) to switch to the recommended content (e.g., using a remote control button, cursor, verbal request, etc.).
[0199] A virtual assistant query recommendation can be determined based on the notification by analyzing the notification content to identify related media, related terms, names, titles, subjects, actions, etc. The identified information can then be used to formulate an appropriate virtual assistant query recommendation, such as a notification-based recommendation 2966, based on the notification 2964. For example, a notification about the exciting end of a live sporting event can be displayed. Then, when the user requests a query recommendation, a recommendation interface 2650 including query recommendations for viewing sporting events, inquiring about team performance, or discovering content related to the notification (e.g., change to the Theta / Alpha match, what is Team Theta's status, what other soccer games are being broadcast) can be displayed. Various other query recommendations can be similarly determined and provided to the user based on specific terms of interest identified in the notification.
[0200] Also, from the content on the user device, virtual assistant query recommendations related to media content (e.g., for consumption via television set-top box 104) can be determined, and the recommendations can be provided on the user device. In some embodiments, playable device content can be identified on a user device connected to or in communication with television set-top box 104. FIG. 30 shows a user device 102 with exemplary photo and video content in interface 1360. A determination can be made as to what content is available for playback on the user device or what content may be desired for playback. For example, playable media 3068 (e.g., photo and video applications) can be identified based on the active application, or playable media 3068 can be identified for stored content regardless of whether it is displayed on interface 1360 (e.g., in some embodiments, content can be identified from the active application, or in other embodiments, without being displayed at a given time). Playable media 3068 may include, for example, videos 1362, photo albums 1364, and photos 1366, each of which may comprise personal user content that may be transmitted to television set-top box 104 for display or playback. In other embodiments, any photos, videos, music, game interfaces, application interfaces, or other media content stored or displayed on user device 102 may be identified and used to determine query recommendations.
[0201] The identified playable media 3068 can be used to determine and provide a virtual assistant query recommendation to the user. FIG. 31 shows an exemplary TV assistant interface 3170 on the user device 102, comprising a virtual assistant query recommendation based on playable user device content and a virtual assistant query recommendation based on video content displayed on a separate display (e.g., a display 112 associated with the television set-top box 104). The TV assistant interface 3170 can include a virtual assistant interface for interacting with media content and / or the television set-top box 104, among other things. When viewing the interface 3170, the user can request a query recommendation on the user device 102, for example, by double-clicking a physical button. Similarly, other inputs can be used to indicate a request for a query recommendation. As shown, the assistant greeting 3172 can introduce the provided query recommendation (e.g., "Here are some recommendations for controlling your TV experience").
[0202] The virtual assistant query recommendations provided on the user device 102 can include recommendations based on various source devices as well as general recommendations. For example, device-based recommendations 3174 can include query recommendations based on content stored on the user device 102 (including content displayed on the user device 102). Content-based recommendations 2652 can be based on content displayed on a display 112 associated with the television set-top box 104. General recommendations 3176 can include general recommendations associated with particular media content or a particular device comprising the media content.
[0203] For example, device-based recommendations 3174 may be determined based on playable content (e.g., videos, music, photos, game interfaces, application interfaces, etc.) identified on user device 102. In the illustrated example, device-based recommendations 3174 may be determined based on playable media 3068 shown in FIG. 30 . For example, assuming photo album 1364 is identified as playable media 3068, details of photo album 1364 may be used to formulate a query. The system may identify the content as an album of multiple photos that can be displayed in a slideshow, and then (in some instances) use the album's title to formulate a query recommendation to display a slideshow of a particular album of photos (e.g., "View a slideshow of 'Graduation Album' from your photos."). In some examples, the recommendation may include an indication of the source of the content (e.g., "From your photos," "From Jennifer's phone," "From Daniel's tablet," etc.). The recommendation may also use other details to reference specific content, such as a recommendation to view photos from a specific date (e.g., display photos from June 21). In another example, the video 1362 can be identified as playable media 3068, and the title (or other identifying information) of the video can be used to formulate a query recommendation to play the video (e.g., "Show 'Graduation Video' from your videos.").
[0204] In other embodiments, content available on other connected devices can be identified and used to formulate virtual assistant query recommendations. For example, content from each of two user devices 102 connected to a common television set-top box 104 can be identified and used to formulate virtual assistant query recommendations. In some embodiments, the user can select which content to make visible to the system for sharing, and can hide other content from the system so that it is not included in the query recommendation or, in some cases, is not available for playback.
[0205] The content-based recommendation 2652 displayed in the interface 3170 of FIG. 31 can be determined, for example, based on content displayed on the display 112 associated with the television set-top box 104. In some embodiments, the content-based recommendation 2652 can be determined in the same manner as described above with reference to FIG. 26. In the illustrated embodiment, the content-based recommendation 2652 shown in FIG. 31 can be based on the video 480 displayed on the display 112 (for example, as in FIG. 26). In this manner, virtual assistant query recommendations can be derived based on content displayed on or available on any number of connected devices. In addition to targeted recommendations, general recommendations 3176 (for example, display a guide, what sports are being broadcast, what is being broadcast on channel 3, etc.) can be pre-determined and provided.
[0206] FIG. 32 shows an exemplary recommendation interface 2650 comprising a connected device-based recommendation 3275 along with a content-based recommendation 2652 displayed on a display 112 associated with a television set-top box 104. In some embodiments, the content-based recommendation 2652 can be determined in the same manner as described above with reference to FIG. 26. As described above, a virtual assistant query recommendation can be formulated based on content on any number of connected devices, and the recommendation can be provided on any number of connected devices. FIG. 32 shows a connected device-based recommendation 3275 that can be derived from content on the user device 102. For example, playable content can be identified on the user device 102, such as the photo and video content displayed in the interface 1360 as playable media 3068 in FIG. 30. The identified playable content on the user device 102 can then be used to formulate a recommendation that can be displayed on a display 112 associated with the television set-top box 104. In some embodiments, the connected device-based recommendation 3275 can be determined in the same manner as the device-based recommendation 3174 described above with reference to FIG. 31. Further, as described above, in some embodiments, the recommendation may include identifying source information, such as "from Jake's phone," as shown in the connected device-based recommendation 3275. Thus, a virtual assistant query recommendation provided on one device can be derived based on content from another device (e.g., displayed content, stored content, etc.). It should be understood that the connected device may include a television set-top box 104 and / or a remote storage device accessible to the user device 102 (e.g., accessing media content stored in the cloud to formulate recommendations).
[0207] It should be understood that any combination of virtual assistant query recommendations from various sources can be provided in response to a recommendation request. For example, recommendations from various sources can be randomly combined, or recommendations can be presented from various sources based on popularity, user preferences, selection history, etc. Furthermore, queries can be determined in various other ways and presented based on various other factors, such as query history, user preferences, query popularity, etc. Furthermore, in some embodiments, query recommendations can be automatically cycled by replacing the displayed recommendation with a new alternative recommendation after a delay. Furthermore, it should be understood that a user can select a displayed recommendation on any interface, for example, by tapping on a touchscreen, uttering a query, selecting a query using navigation keys, selecting a query using a button, selecting a query using a cursor, etc., and then providing an associated response (e.g., information and / or a media response).
[0208] Also, in any of the various embodiments, virtual assistant query recommendations can be filtered based on available content. For example, potential query recommendations that result in unavailable media content (e.g., no cable subscription) or that may have associated information answers may not qualify as recommendations and may be hidden without display. On the other hand, potential query recommendations that result in immediately playable media content to which the user has access may be weighted more than other potential recommendations, or in some cases, may be biased for display. In this way, the availability of media content for the user to view can be used when determining virtual assistant query recommendations for display.
[0209] Further, in any of various embodiments, pre-loaded query answers may be provided instead of or in addition to recommendations (e.g., in recommendation interface 2650). Such pre-loaded query answers may be selected and provided based on personal use and / or the current context. For example, a user watching a particular program may tap a button, double-click a button, etc. to receive recommendations. Instead of or in addition to query recommendations, context-based information may be automatically provided, such as identifying the song or soundtrack being played (e.g., "This song is a Performance Piece"), identifying the cast of the currently playing episode (e.g., "Actress Janet Quinn plays Genevieve"), identifying similar media (e.g., "Program Q is similar to this program"), or providing results of any of the other queries discussed herein.
[0210] Additionally, affordances can be provided in any of a variety of interfaces that allow a user to rate media content and inform the virtual assistant of the user's preferences (e.g., selectable rating scales). In other examples, a user can speak rating information as a natural language command (e.g., "I love this," "I hate this," "I don't like this show," etc.). In still other examples, various other functional and informational elements can be provided in any of the various interfaces illustrated and described herein. For example, an interface can further include links to important functions and locations, such as search links, purchase links, media links, etc. In another example, an interface can further include recommendations for what else to watch next based on currently playing content (e.g., selecting similar content). In yet another example, an interface can further include recommendations for what else to watch next based on personalized preferences and / or recent activity (e.g., selecting content based on user ratings, user-entered preferences, recently viewed programs, etc.). In still other examples, the interface can further include user interaction instructions (e.g., "press and hold to talk to your virtual assistant," "tap once to get recommendations," etc.). In some examples, providing preloaded answers, recommendations, etc. can make the user experience enjoyable while making content readily accessible to a wide variety of users (e.g., users of various skill levels, regardless of language or other control barriers).
[0211] FIG. 33 shows an example process 3300 for recommending a virtual assistant interaction (e.g., a virtual assistant query) for controlling media content. In block 3302, media content can be displayed on a display. For example, as shown in FIG. 26, a video 480 can be displayed on the display 112 via a television set-top box 104, or as shown in FIG. 30, an interface 1360 can be displayed on the touch screen 246 of the user device 102. In block 3304, input from a user can be received. The input can include a request for a virtual assistant query recommendation. The input can include a button press, a button double-click, a menu selection, a verbal query for a recommendation, etc.
[0212] In block 3306, a virtual assistant query can be determined based on media content and / or a browsing history of media content. For example, a virtual assistant query can be determined based on a displayed program, menu, application, list of media content, notification, etc. In one embodiment, a content-based recommendation 2652 can be determined based on a video 480 and associated metadata as described with reference to FIG. 26. In another embodiment, a notification-based recommendation 2966 can be determined based on a notification 2964 as described with reference to FIG. 29. In yet another embodiment, a device-based recommendation 3174 can be determined based on playable media 3068 on the user device 102 as described with reference to FIGS. 30 and 31. In yet another embodiment, a connected device-based recommendation 3275 can be determined based on playable media 3068 on the user device 102 as described with reference to FIG. 32.
[0213] Referring again to process 3300 of FIG. 33, in block 3308, the virtual assistant query can be displayed on the display. For example, the determined query recommendation can be displayed as shown in and described with reference to FIGS. 26, 27, 29, 31, and 32. As discussed above, query recommendations can be determined and displayed based on various other information. Furthermore, virtual assistant query recommendations provided on one display can be derived based on content from another device with another display. In this way, targeted virtual assistant query recommendations can be provided to the user, thereby assisting the user in learning potential queries and providing desirable content recommendations, among other benefits.
[0214] Additionally, in any of the various embodiments discussed herein, various aspects may be personalized for a particular user. User data, including contacts, preferences, location, favorite media, etc., may be used to interpret voice commands and enable user interaction with the various devices discussed herein. Various processes discussed herein may also be modified in various other ways according to user preferences, contacts, text, usage history, profile data, statistics, etc. Furthermore, such preferences and settings may be updated over time based on user interactions (e.g., frequently issued commands, frequently selected applications, etc.). The collection and use of user data available from various sources may be used to improve the delivery to users of invitation-only content or any other content that may be of interest to them. The present disclosure contemplates that, in some cases, this collected data may include personal information data that uniquely identifies or can be used to contact or locate a particular person. Such personal information data may include demographic data, location-based data, phone numbers, email addresses, home addresses, or any other identifying information.
[0215] This disclosure recognizes that the use of such personal information data in current technology can be used to benefit the user. For example, personal information data can be used to deliver targeted content that is of greater interest to the user. Thus, the use of such personal information data allows for computational control of the delivered content. Furthermore, other uses of personal information data to benefit the user are also contemplated by this disclosure.
[0216] This disclosure further contemplates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities must implement and consistently use generally recognized privacy policies and practices that meet or exceed industry or government requirements for maintaining personal information data confidentially and securely. For example, personal information from users should be collected for the entity's lawful and legitimate use and should not be shared or sold except for those lawful uses. Furthermore, such collection should occur only after receiving the user's informed consent. Furthermore, such entities will take all necessary measures to protect and secure access to such personal information and to ensure that others with access to that personal information comply with their own privacy policies and procedures. Furthermore, such entities may submit themselves to third-party assessment to attest to their compliance with widely accepted privacy policies and practices.
[0217] Notwithstanding the foregoing, the present disclosure also contemplates embodiments in which a user selectively prevents use of or access to personal information data. That is, the present disclosure contemplates providing hardware and / or software elements that prevent or block access to such personal information data. For example, in the case of an ad delivery service, the technology may be configured to allow a user to "opt in" or "opt out" of participating in the collection of personal information data during registration for the service. In another embodiment, a user may choose not to provide location information to a targeted content delivery service. In yet another embodiment, a user may choose not to provide precise location information but to allow the transfer of location zone information.
[0218] Thus, while this disclosure broadly encompasses the use of personal information data to practice one or more various disclosed embodiments, this disclosure also contemplates that various examples thereof may be implemented without requiring access to such personal information data. That is, various examples of the present technology are not rendered inoperable due to the absence of all or a portion of such personal information data. For example, content may be selected and delivered to a user by inferring preferences based on non-personal information or a minimal amount of personal information, such as content requested by devices associated with the user, other non-personal information available to content delivery services, or publicly available information.
[0219] According to some embodiments, FIG. 34 illustrates a functional block diagram of an electronic device 3400 configured to, for example, control television interactions using a virtual assistant and display related information using different interfaces in accordance with the principles of various described embodiments. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to implement the principles of various described embodiments. Those skilled in the art will understand that the functional blocks described in FIG. 34 can be combined or separated into sub-blocks to implement the principles of various described embodiments. Therefore, the description herein optionally supports any possible combination or division or further definition of the functional blocks described herein.
[0220] 34 , the electronic device 3400 may include a display unit 3402 (e.g., display 112, touch screen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3400 may further include an input unit 3404 (e.g., a microphone, a receiver, a touch screen, buttons, etc.) configured to receive information such as speech input, tactile input, gesture input, etc. The electronic device 3400 may further include a processing unit 3406 coupled to the display unit 3402 and the input unit 3404. In some embodiments, the processing unit 3406 may include a speech input receiving unit 3408, a media content determining unit 3410, a first user interface displaying unit 3412, a selection receiving unit 3414, and a second user interface displaying unit 3416.
[0221] The processing unit 3406 may be configured to receive speech input from a user (e.g., via the input unit 3404). The processing unit 3406 may be further configured to determine media content based on the speech input (e.g., using the media content determination unit 3410). The processing unit 3406 may be further configured to display a first user interface having a first size (e.g., on the display unit 3402 using the first user interface presentation unit 3412), the first user interface comprising one or more selectable links to media content. The processing unit 3406 may be further configured to receive a selection of one of the one or more selectable links (e.g., from the input unit 3404 using the selection receiving unit 3414). The processing unit 3406 may be further configured to, in response to the selection, display (e.g., on the display unit 3402 using the second user interface display unit 3416) a second user interface having a second size larger than the first size, the second user interface comprising a media content associated with the selection.
[0222] In some embodiments, the first user interface (e.g., of the first user interface presentation unit 3412) expands into a second user interface (e.g., of the second user interface presentation unit 3416) in response to a selection (e.g., of the selection receiving unit 3414). In other embodiments, the first user interface overlays the playing media content. In one embodiment, the second user interface overlays the playing media content. In another embodiment, the speech input (e.g., of the speech input receiving unit 3408 from the input unit 3404) comprises a query, and the media content (e.g., of the media content determination unit 3410) comprises results of the query. In yet another embodiment, the first user interface comprises links to the query results in addition to one or more selectable links to the media content. In another embodiment, the query comprises a weather query, and the first user interface comprises links to media content associated with the weather-related query. In another embodiment, the query comprises a location, and the link to the media content associated with the weather-related query comprises a link to a portion of the media content associated with the weather at the location.
[0223] In some embodiments, in response to the selection, processing unit 3406 may be configured to play media content associated with the selection. In one embodiment, the media content includes a movie. In another embodiment, the media content includes a television program. In another embodiment, the media content includes a sporting event. In some embodiments, the second user interface (e.g., of second user interface display unit 3416) includes a description of the media content associated with the selection. In other embodiments, the first user interface includes a link to purchase the media content.
[0224] Processing unit 3406 may be further configured to receive additional speech input from a user (e.g., via input unit 3404), the additional speech input including a query associated with the displayed content. Processing unit 3406 may be further configured to determine a response to the query associated with the displayed content based on metadata associated with the displayed content. In response to receiving the additional speech input, processing unit 3406 may be further configured to display a third user interface (e.g., on display unit 3402), the third user interface including the determined response to the query associated with the displayed content.
[0225] The processing unit 3406 may be further configured to receive an instruction to begin receiving speech input (e.g., via the input unit 3404). The processing unit 3406 may be further configured to display a readiness confirmation (e.g., on the display unit 3402) in response to receiving the instruction. The processing unit 3406 may be further configured to display a listen confirmation in response to receiving the speech input. The processing unit 3406 may be further configured to detect the end of the speech input and display a processing confirmation in response to detecting the end of the speech input. In some examples, the processing unit 3406 may be further configured to display a phonetic transcription of the speech input.
[0226] In some embodiments, electronic device 3400 includes a television. In some embodiments, electronic device 3400 includes a television set-top box. In some embodiments, electronic device 3400 includes a remote control. In some embodiments, electronic device 3400 includes a mobile phone.
[0227] In one embodiment, one or more selectable links in the first user interface (e.g., of first user interface display unit 3412) include moving images associated with the media content. In some embodiments, the moving images associated with the media content include a live feed of the media content. In another embodiment, one or more selectable links in the first user interface include still images associated with the media content.
[0228] In some examples, processing unit 3406 may be further configured to determine whether the currently displayed content includes moving images or a control menu, and in response to a determination that the currently displayed content includes moving images, select a small size as the first size for the first user interface (e.g., of first user interface display unit 3412), and in response to a determination that the currently shown content includes a control menu, select a large size, larger than the small size, as the first size for the first user interface (e.g., of first user interface display unit 3412). In other examples, processing unit 3406 may be further configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of a live sporting event, and display a notification including the determined alternative media content.
[0229] According to some embodiments, FIG. 35 shows a functional block diagram of an electronic device 3500 configured to control television interactions using, for example, a virtual assistant and multiple user devices in accordance with the principles of various described embodiments. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to implement the principles of various described embodiments. Those skilled in the art will understand that the functional blocks described in FIG. 35 can be combined or separated into sub-blocks to implement the principles of various described embodiments. Therefore, the description herein optionally supports any possible combination or division, or further definition, of the functional blocks described herein.
[0230] 35 , the electronic device 3500 may include a display unit 3502 (e.g., display 112, touchscreen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3500 may include an input unit 3504 (e.g., a microphone, a receiver, a touchscreen, buttons, etc.) further configured to receive information such as speech input, tactile input, gesture input, etc. The electronic device 3500 may further include a processing unit 3506 coupled to the display unit 3502 and the input unit 3504. In some embodiments, the processing unit 3506 may include a speech input receiving unit 3508, a user intention determination unit 3510, a media content determination unit 3512, and a media content playback unit 3514.
[0231] The processing unit 3506 may be configured to receive speech input from a user (e.g., from input unit 3504, using speech input receiving unit 3508) at a first device (e.g., device 3500) having a first display (e.g., display unit 3502, in some embodiments). The processing unit 3506 may be further configured to determine a user intent of the speech input based on content displayed on the first display (e.g., using user intent determination unit 3510). The processing unit 3506 may be further configured to determine media content based on the user intent (e.g., using media content determination unit 3512). The processing unit 3506 may be further configured to play media content (e.g., using media content playback unit 3514) on a second device (e.g., display unit 3502, in some embodiments) associated with a second display.
[0232] In one embodiment, the first device includes a remote control. In another embodiment, the first device includes a mobile phone. In another embodiment, the first device includes a tablet computer. In some embodiments, the second device includes a television set-top box. In another embodiment, the second device includes a television.
[0233] In some embodiments, the content displayed on the first display comprises an application interface. In one embodiment, the speech input (e.g., of the speech input receiving unit 3508 from the input unit 3504) comprises a request to display on media associated with the application interface. In one embodiment, the media content comprises media associated with the application interface. In another embodiment, the application interface comprises a photo album, and the media comprises one or more photos in the photo album. In yet another embodiment, the application interface comprises a list of one or more videos, and the media comprises one of the one or more videos. In yet another embodiment, the application interface comprises a television program listing, and the media comprises a television program in the television program listing.
[0234] In some examples, the processing unit 3506 may be further configured to determine whether the first device is authenticated, and in response to determining that the first device is authenticated, play media content on the second device. The processing unit 3506 may be further configured to identify a user based on the speech input and determine a user intent of the speech input based on data associated with the identified user (e.g., using the user intent determination unit 3510). The processing unit 3506 may be further configured to determine whether the user is authenticated based on the speech input, and in response to determining that the user is an authenticated user, play media content on the second device. In one example, determining whether the user is authenticated includes analyzing the speech input using speech recognition.
[0235] In other examples, processing unit 3506 may be further configured to, in response to determining that the user intent includes a request for information, display information associated with the media content on a first display of the first device. In response to determining that the user intent includes a request to play the media content, processing unit 3506 may be further configured to play information associated with the media content on a second device.
[0236] In some embodiments, the speech input includes a request to play content on the second device, and in response to the request to play content on the second device, the media content is played on the second device. Processing unit 3506 may be further configured to determine whether the determined media content should be displayed on the first display or the second display based on the media format, user preferences, or default settings. In some embodiments, in response to determining that the determined media content should be displayed on the second display, the media content is displayed on the second display. In other embodiments, in response to determining that the determined media content should be displayed on the first display, the media content is displayed on the first display.
[0237] In other examples, processing unit 3506 may be further configured to determine a proximity of each of two or more devices, including a second device and a third device. In some examples, playing media content on a second device associated with a second display based on the proximity of the second device relative to the proximity of the third device. In some examples, determining the proximity of each of the two or more devices includes determining the proximity based on Bluetooth LE.
[0238] In some examples, the processing unit 3506 may be further configured to display a list of display devices including the second device associated with the second display and receive a selection of the second device in the list of display devices. In one example, in response to receiving the selection of the second device, display media content on the second display. The processing unit 3506 may be further configured to determine whether headphones are attached to the first device. In response to determining that headphones are attached to the first device, the processing unit 3506 may be further configured to display media content on the first display. In response to determining that headphones are not attached to the first device, the processing unit 3506 may be further configured to display media content on the second display. In other examples, the processing unit 3506 may be further configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of a live sporting event, and to display a notification including the determined alternative media content.
[0239] According to some embodiments, FIG. 36 illustrates a functional block diagram of an electronic device 3600 configured to control television interactions, for example, using media content displayed on a display and a media content viewing history, in accordance with the principles of various described embodiments. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to carry out the principles of various described embodiments. Those skilled in the art will understand that the functional blocks described in FIG. 36 can be combined or separated into sub-blocks to implement the principles of various described embodiments. Thus, the description herein optionally supports any possible combination or division or further definition of the functional blocks described herein.
[0240] 36 , the electronic device 3600 may include a display unit 3602 (e.g., display 112, touch screen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3600 may further include an input unit 3604 (e.g., a microphone, a receiver, a touch screen, buttons, etc.) configured to receive information such as speech input, tactile input, gesture input, etc. The electronic device 3600 may further include a processing unit 3606 coupled to the display unit 3602 and the input unit 3604. In some embodiments, the processing unit 3606 may include a speech input receiving unit 3608, a user intent determining unit 3610, and a query result display unit 3612.
[0241] The processing unit 3606 may be configured to receive speech input from a user (e.g., from the input unit 3604 using the speech input receiving unit 3608), the speech input including a query associated with content displayed on a display (e.g., the display unit 3602 in some embodiments). The processing unit 3606 may be further configured to determine the user intent of the query based on one or more of the content displayed on the television display and a browsing history of the media content (e.g., using the user intent determination unit 3610). The processing unit 3606 may be further configured to display results of the query based on the determined user intent (e.g., using the query result display unit 3612).
[0242] In one embodiment, the speech input is received at a remote control. In another embodiment, the speech input is received at a mobile phone. In some embodiments, the query results are displayed on a television display. In another embodiment, the content displayed on the television display includes a movie. In yet another embodiment, the content displayed on the television display includes a television program. In yet another embodiment, the content displayed on the television display includes a sporting event.
[0243] In some embodiments, the query includes a request for information about a person associated with the content displayed on the television display, and the query results (e.g., in the query result display unit 3612) include information about the person. In one embodiment, the query results include media content associated with the person. In another embodiment, the media content includes one or more of a movie, television program, or sporting event associated with the person. In some embodiments, the query includes a request for information about a character associated with the content displayed on the television display, and the query results include information about the character or an actor playing the character. In one embodiment, the query results include media content associated with the actor playing the character. In another embodiment, the media content includes one or more of a movie, television program, or sporting event associated with the actor playing the character.
[0244] In some embodiments, the processing unit 3606 may be further configured to determine the results of the query based on metadata associated with the content displayed on the television display or a viewing history of the media content. In one embodiment, the metadata includes one or more of a title, a description, a list of characters, a list of actors, a list of players, a genre, or a viewing schedule associated with the content displayed on the television display or a viewing history of the media content. In another embodiment, the content displayed on the television display includes a list of media content, and the query includes a request to display one of the items in the list. In yet another embodiment, the content displayed on the television display further includes an item in the list of media content having focus, and determining the user intent of the query (e.g., using the user intent determination unit 3610) includes identifying the item having focus. In some embodiments, the processing unit 3606 may be further configured to determine the user intent of the query based on menus or search content recently displayed on the television display (e.g., using the user intent determination unit 3610). In one embodiment, the content displayed on the television display includes pages of enumerated media, and the recently displayed menus or search content includes previous pages of the enumerated media. In another example, the content displayed on the television display includes one or more categories of media, one of which has a focus. In one example, processing unit 3606 may be further configured to determine a user intent for the query based on one of the one or more categories of media that has a focus (e.g., using user intent determination unit 3610). In another example, the media categories include movies, television programs, and music.In other examples, the processing unit 3606 may be further configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of a live sporting event, and to display a notification including the determined alternative media content.
[0245] According to some embodiments, FIG. 37 shows a functional block diagram of an electronic device 3700, which is configured to recommend a virtual assistant interaction for controlling media content, for example, in accordance with the principles of various embodiments described. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to implement the principles of the various embodiments described. Those skilled in the art will understand that the functional blocks described in FIG. 37 can be combined or separated into sub-blocks to implement the principles of the various embodiments described. Therefore, the description herein optionally supports any possible combination or division, or further definition, of the functional blocks described herein.
[0246] 37 , the electronic device 3700 may include a display unit 3702 (e.g., display 112, touch screen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3700 may further include an input unit 3704 (e.g., a microphone, a receiver, a touch screen, buttons, etc.) configured to receive information such as speech input, tactile input, gesture input, etc. The electronic device 3700 may further include a processing unit 3706 coupled to the display unit 3702 and the input unit 3704. In some embodiments, the processing unit 3706 may include a media content display unit 3708, an input receiving unit 3710, a query determination unit 3712, and a query display unit 3714.
[0247] The processing unit 3706 can be configured to display media content on a display (e.g., display unit 3702) (e.g., using the media content display unit 3708). The processing unit 3706 can be further configured to receive input from a user (e.g., from the input unit 3704 using the input receiving unit 3710). The processing unit 3706 can be further configured to determine one or more virtual assistant queries based on the media content and one or more of the browsing history of the media content (e.g., using the query determination unit 3712). The processing unit 3706 can be further configured to display one or more virtual assistant queries on the display (e.g., using the query display unit 3714).
[0248] In one embodiment, input from a user is received on a remote control. In another embodiment, input from a user is received on a mobile phone. In some embodiments, one or more virtual assistant queries are overlaid on a moving image. In another embodiment, the input includes double-clicking a button. In one embodiment, the media content includes a movie. In another embodiment, the media content includes a television program. In yet another embodiment, the media content includes a sporting event.
[0249] In some embodiments, one or more virtual assistant queries include queries about people appearing in the media content. In other embodiments, one or more virtual assistant queries include queries about characters appearing in the media content. In another embodiment, one or more virtual assistant queries include queries about media content associated with people appearing in the media content. In some embodiments, the media content or the browsing history of the media content includes episodes of a television program, and one or more virtual assistant queries include queries about another episode of the television program. In another embodiment, the media content or the browsing history of the media content includes episodes of a television program, and one or more virtual assistant queries include a request to set a reminder to watch or record a subsequent episode of the media content. In yet another embodiment, one or more virtual assistant queries include a query for descriptive details of the media content. In one embodiment, the descriptive details include one or more of a program title, a character list, an actor list, an episode description, a team roster, a team ranking, or a program summary.
[0250] In some embodiments, the processing unit 3706 can be further configured to receive a selection of one of the one or more virtual assistant queries. The processing unit 3706 can be further configured to display results of the selected query of the one or more virtual assistant queries. In one embodiment, determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on one or more of query history, user preferences, or query popularity. In another embodiment, determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on media content available for viewing by the user. In yet another embodiment, determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on a received notification. In yet another embodiment, determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on an active application. In other embodiments, the processing unit 3706 can be further configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of a live sporting event, and display a notification including the determined alternative media content.
[0251] While the illustrative embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art (e.g., modifying any of the other systems or processes discussed herein in accordance with the concepts described with respect to any other systems or processes discussed herein), and such changes and modifications are to be understood as being included within the scope of the various illustrative embodiments as defined by the appended claims.
Claims
1. 1. A method for controlling television interactions using a virtual assistant, the method comprising: In electronic devices, receiving speech input from a user; determining media content based on the speech input; displaying a first user interface having a first size, the first user interface including one or more selectable links to the media content; receiving a selection of one of the one or more selectable links; and, in response to the selection, displaying a second user interface having a second size larger than the first size, the second user interface including the media content associated with the selection; A method comprising:
2. The method of claim 1 , wherein in response to the selection, the first user interface expands to the second user interface.
3. The method of claim 1 , wherein the first user interface overlays the media content being played.
4. The method of claim 1 , wherein the second user interface overlays the media content being played.
5. The method of claim 1 , wherein the speech input comprises a query and the media content comprises results of the query.
6. The method of claim 5 , wherein the first user interface includes, in addition to the one or more selectable links to the media content, a link to results of the query.
7. The method of claim 5 , wherein the query comprises a weather query, and the first user interface comprises a link to media content associated with the weather query.
8. The method of claim 7 , wherein the query includes a location, and the link to the media content associated with the query regarding the weather includes a link to a portion of media content associated with weather at the location.
9. The method of claim 1 , further comprising, in response to the selection, playing the media content associated with the selection.
10. The method of claim 1 , wherein the media content comprises a movie.
11. The method of claim 1 , wherein the media content comprises a television program.
12. The method of claim 1 , wherein the media content comprises a sporting event.
13. The method of claim 1 , wherein the second user interface includes a description of the media content associated with the selection.
14. The method of claim 1 , wherein the first user interface includes a link to purchase media content.
15. receiving additional speech input from the user, the additional speech input including a query associated with the displayed content; and determining a response to the query associated with the displayed content based on metadata associated with the displayed content; displaying a third user interface in response to receiving the additional speech input, the third user interface including the determined response to the query associated with the displayed content; and The method of claim 1 further comprising:
16. receiving an instruction to begin receiving speech input; In response to receiving the instruction, displaying a readiness confirmation; The method of claim 1 further comprising:
17. The method of claim 1 , further comprising displaying a listen confirmation in response to receiving the speech input.
18. Detecting an end of the speech input; displaying a processing confirmation in response to detecting the end of the speech input; The method of claim 1 further comprising:
19. The method of claim 1 , further comprising displaying a phonetic transcription of the speech input.
20. The method of claim 1 , wherein the electronic device comprises a television.
21. The method of claim 1 , wherein the electronic device comprises a television set-top box.
22. The method of claim 1 , wherein the electronic device comprises a remote control.
23. The method of claim 1 , wherein the electronic device comprises a mobile phone.
24. The method of claim 1 , wherein the one or more selectable links in the first user interface include a moving image associated with the media content.
25. 25. The method of claim 24, wherein the moving image associated with the media content comprises a live feed of the media content.
26. The method of claim 1 , wherein the one or more selectable links in the first user interface include a still image associated with the media content.
27. determining whether the currently displayed content includes a moving image or a control menu; In response to determining that the currently displayed content includes moving images, selecting a small size as the first size for the first user interface; In response to determining that the currently displayed content includes a control menu, selecting a large size as the first size for the first user interface, the large size being larger than the small size; The method of claim 1 further comprising:
28. determining alternative media content for display based on one or more of user preferences, program popularity, and the status of the live sporting event; displaying a notification including the determined alternative media content; The method of claim 1 further comprising:
29. 29. A non-transitory computer-readable storage medium comprising computer-executable instructions for performing the method of any one of claims 1 to 28.
30. 30. The non-transitory computer-readable storage medium of claim 29; a processor capable of executing the computer-executable instructions; and A system comprising:
31. A system for controlling television interactions using a virtual assistant, the system comprising: means for receiving speech input from a user; means for determining media content based on the speech input; means for displaying a first user interface having a first size, the first user interface including one or more selectable links to the media content; means for receiving a selection of one of the one or more selectable links; means for displaying, in response to the selection, a second user interface having a second size larger than the first size, the second user interface including the media content associated with the selection; and A system comprising:
32. 1. A method for controlling television interactions using a virtual assistant, the method comprising: In electronic devices, receiving speech input from a user at a first device having a first display; determining a user's intent for the speech input based on content displayed on the first display; determining media content based on the user intent; playing the media content on a second device associated with a second display; A method comprising:
33. 33. The method of claim 32, wherein the first device comprises a remote control.
34. 33. The method of claim 32, wherein the first device comprises a mobile phone.
35. 33. The method of claim 32, wherein the first device comprises a tablet computer.
36. 33. The method of claim 32, wherein the second device comprises a television set-top box.
37. 33. The method of claim 32, wherein the second display comprises a television.
38. 33. The method of claim 32, wherein the content displayed on the first display comprises an application interface.
39. 40. The method of claim 38, wherein the speech input comprises a request to display media associated with the application interface.
40. 40. The method of claim 39, wherein the media content includes the media associated with the application interface.
41. 41. The method of claim 40, wherein the application interface comprises a photo album and the media comprises one or more photographs in the photo album.
42. 41. The method of claim 40, wherein the application interface comprises a list of one or more videos, and the media includes one of the one or more videos.
43. 41. The method of claim 40, wherein the application interface comprises a television program listing and the media includes a television program in the television program listing.
44. determining whether the first device is authenticated; 33. The method of claim 32, further comprising: playing the media content on the second device in response to determining that the first device is authenticated.
45. identifying the user based on the speech input; determining the user intent of the speech input based on data associated with the identified user; 33. The method of claim 32, further comprising:
46. determining whether the user is authenticated based on the speech input; 33. The method of claim 32, further comprising: playing media content on the second device in response to determining that the user is an authenticated user.
47. 47. The method of claim 46, wherein determining whether the user is authenticated includes analyzing the speech input using voice recognition.
48. In response to determining that the user intent includes a request for information, displaying information associated with the media content on the first display of the first device; In response to determining that the user intent includes a request to play the media content, playing information associated with the media content on the second device; and 33. The method of claim 32, further comprising:
49. the speech input includes a request to play content on the second device; 33. The method of claim 32, further comprising playing the media content on the second device in response to the request to play content on the second device.
50. determining whether the determined media content should be displayed on the first display or the second display based on media format, user preferences, or default settings; displaying the media content on the second display in response to determining that the determined media content should be displayed on the second display; 33. The method of claim 32, further comprising displaying the media content on the first display in response to determining that the determined media content should be displayed on the first display.
51. determining a proximity of each of two or more devices, including the second device and a third device; 33. The method of claim 32, further comprising playing media content on the second device associated with the second display based on the proximity of the second device relative to the proximity of the third device.
52. 52. The method of claim 51, wherein determining the proximity of each of the two or more devices includes determining the proximity based on Bluetooth LE.
53. displaying a list of display devices including the second device associated with the second display; receiving a selection of the second device in the list of display devices; 33. The method of claim 32, further comprising: displaying the media content on the second display in response to receiving the selection of the second device.
54. determining whether headphones are attached to the first device; displaying the media content on the first display in response to determining that the headphones are attached to the first device; In response to determining that the headphones are not attached to the first device, displaying the media content on the second display; 33. The method of claim 32, further comprising:
55. determining alternative media content for display based on one or more of user preferences, program popularity, and the status of the live sporting event; displaying a notification including the determined alternative media content; 33. The method of claim 32, further comprising:
56. 56. A non-transitory computer-readable storage medium comprising computer-executable instructions for performing the method of any one of claims 32 to 55.
57. 57. A non-transitory computer-readable storage medium according to claim 56; and a processor capable of executing computer-executable instructions. A system comprising:
58. A system for controlling television interactions using a virtual assistant, the system comprising: means for receiving speech input from a user at a first device having a first display; means for determining a user's intention for the speech input based on content displayed on the first display; means for determining media content based on the user intent; means for playing the media content on a second device associated with a second display; A system comprising:
59. 1. A method for controlling television interactions using a virtual assistant, the method comprising: In electronic devices, receiving speech input from a user, the speech input including a query associated with content displayed on a television display; determining a user intent for the query based on one or more of the content displayed on the television display and a media content viewing history; displaying results of the query based on the determined user intent; and A method comprising:
60. 60. The method of claim 59, wherein the speech input is received in a remote control.
61. 60. The method of claim 59, wherein the speech input is received at a mobile phone.
62. 60. The method of claim 59, further comprising displaying the results of the query on the television display.
63. 60. The method of claim 59, wherein the content displayed on the television display includes a movie.
64. 60. The method of claim 59, wherein the content displayed on the television display comprises a television program.
65. 60. The method of claim 59, wherein the content displayed on the television display includes a sporting event.
66. 60. The method of claim 59, wherein the query includes a request for information about a person associated with the content displayed on the television display, and the results of the query include information about the person.
67. 67. The method of claim 66, wherein the results of the query include media content associated with the person.
68. 68. The method of claim 67, wherein the media content includes one or more of a movie, a television program, or a sporting event associated with the person.
69. 60. The method of claim 59, wherein the query includes a request for information about a character associated with the content displayed on the television display, and the results of the query include information about the character or an actor playing the character.
70. 70. The method of claim 69, wherein the results of the query include media content associated with the actor playing the character.
71. 71. The method of claim 70, wherein the media content includes one or more of a movie, a television program, or a sporting event associated with the actor playing the character.
72. 60. The method of claim 59, further comprising determining the results of the query based on metadata associated with the viewing history of the content or media content displayed on the television display.
73. 73. The method of claim 72, wherein the metadata includes one or more of a title, a description, a list of characters, a list of actors, a list of players, a genre, or a display schedule associated with the content or the viewing history of media content displayed on the television display.
74. 60. The method of claim 59, wherein the content displayed on the television display includes a list of media content, and the query includes a request to display one of the items in the list.
75. 75. The method of claim 74, wherein the content displayed on the television display further includes an item in the list of media content having focus, and determining the user intent of the query includes identifying the item having focus.
76. 60. The method of claim 59, further comprising determining the user intent of the query based on recently displayed menus or search content on the television display.
77. 77. The method of claim 76, wherein the content displayed on the television display includes a page of enumerated media and the recently displayed menu or search content includes a previous page of enumerated media.
78. 60. The method of claim 59, wherein the content displayed on the television display includes one or more categories of media, one of the one or more categories of media having a focus.
79. 80. The method of claim 78, further comprising determining the user intent of the query based on the one of the one or more categories of the media having a focus.
80. 79. The method of claim 78, wherein the media categories include movies, television programs, and music.
81. determining alternative media content for display based on one or more of user preferences, program popularity, and the status of the live sporting event; displaying a notification including the determined alternative media content; 60. The method of claim 59, further comprising:
82. 82. A non-transitory computer-readable storage medium comprising computer-executable instructions for performing the method of any one of claims 59 to 81.
83. 83. A non-transitory computer-readable storage medium according to claim 82; a processor capable of executing the computer-executable instructions; and A system comprising:
84. A system for controlling television interactions using a virtual assistant, the system comprising: means for receiving speech input from a user, the speech input including a query associated with content displayed on a television display; means for determining a user intent of the query based on one or more of the content displayed on the television display and a media content viewing history; means for displaying results of the query based on the determined user intent; A system comprising:
85. A method for recommending virtual assistant interactions to control media content, the method comprising: In electronic devices, Displaying media content on a display; receiving input from a user; Determining one or more virtual assistant queries based on one or more of the media content and browsing history with the media content; Displaying the one or more virtual assistant queries on the display; A method comprising:
86. 86. The method of claim 85, wherein the input is received from the user on a remote control.
87. 86. The method of claim 85, wherein the input is received from the user on a mobile phone.
88. 86. The method of claim 85, wherein the one or more virtual assistant queries are overlaid on a moving image.
89. 86. The method of claim 85, wherein the input comprises a double-click of a button.
90. 86. The method of claim 85, wherein the media content comprises a movie.
91. 86. The method of claim 85, wherein the media content comprises a television program.
92. 86. The method of claim 85, wherein the media content comprises a sporting event.
93. 86. The method of claim 85, wherein the one or more virtual assistant queries include queries about people appearing in the media content.
94. 86. The method of claim 85, wherein the one or more virtual assistant queries include queries regarding characters appearing in the media content.
95. 86. The method of claim 85, wherein the one or more virtual assistant queries include queries about media content associated with a person appearing in the media content.
96. 86. The method of claim 85, wherein the media content or the browsing history of media content includes an episode of a television program, and the one or more virtual assistant queries include a query for another episode of the television program.
97. 86. The method of claim 85, wherein the media content or the viewing history of media content includes episodes of a television program, and the one or more virtual assistant queries include a request to set a reminder to watch or record a subsequent episode of the media content.
98. 86. The method of claim 85, wherein the one or more virtual assistant queries include a query for descriptive details of the media content.
99. 99. The method of claim 98, wherein the descriptive details include one or more of a program title, a character list, an actor list, an episode description, a team roster, a team ranking, or a program synopsis.
100. Receiving a selection of one of the one or more virtual assistant queries; Receiving results of the selected one of the one or more virtual assistant queries; 86. The method of claim 85, further comprising:
101. 86. The method of claim 85, wherein determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on one or more of query history, user preferences, or query popularity.
102. 86. The method of claim 85, wherein determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on media content available for viewing by the user.
103. 86. The method of claim 85, wherein determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on a received notification.
104. 86. The method of claim 85, wherein determining the one or more virtual assistant queries includes determining the one or more virtual assistant queries based on an active application.
105. determining alternative media content for display based on one or more of user preferences, program popularity, and the status of the live sporting event; displaying a notification including the determined alternative media content; 86. The method of claim 85, further comprising:
106. 106. A non-transitory computer-readable storage medium comprising computer-executable instructions for performing the method of any one of claims 85 to 105.
107. 107. A non-transitory computer-readable storage medium according to claim 106; a processor capable of executing the computer-executable instructions; and A system comprising:
108. A system for recommending virtual assistant interactions to control media content, the system comprising: means for displaying media content on a display; means for receiving input from a user; A means for determining one or more virtual assistant queries based on one or more of the media content and browsing history of the media content; A means for displaying the one or more virtual assistant queries on the display; A system comprising:
Citation Information
Patent Citations
Remote operation control system and its program recording medium
JP2002034087A
Method, device, and program for command processing
JP2002287793A
Radio wave remote control system
JP2003284170A
Program information display apparatus with voice-recognition capability
JP2004260544A
Portable terminal, method for controlling the same, and program
JP2011087110A