Smart automated assistant for television user interaction

By receiving voice input and processing user intent through a virtual assistant system, the problem of complex interaction with devices such as televisions is solved, providing an intuitive interface and query suggestions, thus improving the user experience.

CN114093348BActive Publication Date: 2026-05-08APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2015-03-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The interaction between users and media control devices such as televisions is complex and difficult to learn, especially when faced with multiple media sources, making it difficult to quickly find the desired content, resulting in a poor user experience.

Method used

A virtual assistant system is used to receive voice input, determine the user's intent through natural language processing, and control the TV interaction based on this, providing an intuitive user interface and query suggestions to simplify access to media content.

Benefits of technology

It enables intuitive and simple interaction between users and devices such as televisions, improves the accessibility of media content and user experience, and provides convenient control options through voice commands and interface expansion/collapse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114093348B_ABST
    Figure CN114093348B_ABST
Patent Text Reader

Abstract

A system and method for controlling television user interactions using a virtual assistant is disclosed. The virtual assistant can interact with a television set-top box to control content displayed on a television. Voice input for the virtual assistant can be received from a device having a microphone. A user intent can be determined from the voice input, and the virtual assistant can perform a task in accordance with the user intent, including causing media to be played back on the television. Virtual assistant interactions can be displayed on the television in an interface that expands or contracts to occupy a minimal amount of space while conveying desired information. A user intent can be determined from the voice input, and information conveyed to the user using multiple devices associated with multiple displays. In some examples, virtual assistant query suggestions can be provided to the user based on media content displayed on the displays.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application filed on March 27, 2015, with international application number PCT / US2015 / 023089, national application number 201580029053.9, and entitled "Intelligent Automated Assistant for Television User Interaction".

[0002] Cross-reference to related applications

[0003] This application claims priority to U.S. Provisional Application No. 62 / 019,312, filed June 30, 2014, entitled “Intelligent Automated Assistant for TV User Interactions,” and U.S. Non-Provisional Application No. 14 / 498,503, filed September 26, 2014, entitled “Intelligent Automated Assistant for TV User Interactions,” both of which are incorporated herein by reference in their entirety for all purposes.

[0004] This application is also related to the following co-pending provisional application: U.S. Patent Application No. 62 / 019,292 (Agency Document No. 106843097900P22498USP1), filed June 30, 2014, entitled “REAL-TIME DIGITAL ASSISTANT KNOWLEDGE UPDATES”, the entire contents of which are incorporated herein by reference. Technical Field

[0005] The present invention relates generally to controlling the interaction of a television user, and more specifically, to processing voice for a virtual assistant to control the interaction of a television user. Background Technology

[0006] Intelligent automated assistants (or virtual assistants) provide an intuitive interface between users and electronic devices. These assistants allow users to interact with devices or systems using natural language in the form of spoken language and / or text. For example, a user can access the services of an electronic device by providing spoken user input in natural language to a virtual assistant associated with the device. The virtual assistant can perform natural language processing on the spoken user input to infer the user's intent and operationalize that intent into a task. The task can then be performed by executing one or more functions of the electronic device, and in some examples, relevant output can be returned to the user in natural language.

[0007] While mobile phones (e.g., smartphones) and tablets have benefited from virtual assistant control, many other user devices lack such convenient control mechanisms. For example, user interaction with media control devices (e.g., televisions, set-top boxes, cable boxes, gaming devices, streaming devices, digital video recorders, etc.) can be complex and difficult to learn. Furthermore, with the increasing number of media sources available through such devices (e.g., over-the-air television, subscription television services, streaming video services, on-demand cable video services, web-based video services, etc.), finding the desired media content to consume can be cumbersome or even impossible for some users. As a result, many media control devices may offer a poor user experience, leading to frustration for many users. Summary of the Invention

[0008] This invention discloses a system and process for controlling television interaction using a virtual assistant. In one embodiment, voice input can be received from a user. Media content can be determined based on the voice input. A first user interface having a first size can be displayed, which may include selectable links to the media content. A selection of one of the selectable links can be received. In response to the selection, a second user interface having a second size larger than the first size can be displayed, which may include the media content associated with the selection.

[0009] In another embodiment, voice input can be received from a user at a first device having a first display. The user's intent regarding the voice input can be determined based on the content displayed on the first display. Media content can be determined based on the user's intent. The media content can be played on a second device associated with a second display.

[0010] In another embodiment, voice input can be received from a user, which may include a query associated with content displayed on a television display. The user intent for the query can be determined based on the viewing history of the content displayed on the television display and / or media content. The query results can be displayed based on the determined user intent.

[0011] In another embodiment, media content can be displayed on a monitor. Input can be received from the user. The virtual assistant query can be determined based on the media content and / or the viewing history of the media content. The virtual assistant query can be displayed on the monitor. Attached Figure Description

[0012] Figure 1 An exemplary system for using a virtual assistant to control user interaction on a television set is shown.

[0013] Figure 2 A block diagram of an exemplary user device according to various embodiments is shown.

[0014] Figure 3 A block diagram of an exemplary media control device in a system for controlling user interaction on a television set is shown.

[0015] Figures 4A-4E An exemplary voice input interface is shown above the video content.

[0016] Figure 5 An exemplary media content interface is shown above the video content.

[0017] Figures 6A-6B An example media details interface is shown above the video content.

[0018] Figures 7A-7B An exemplary media switching interface is shown.

[0019] Figures 8A-8B An exemplary voice input interface is shown above the menu content.

[0020] Figure 9 An exemplary virtual assistant results interface is shown above the menu content.

[0021] Figure 10 An exemplary process is shown for using a virtual assistant to control television interaction and for using different interfaces to display related information.

[0022] Figure 11 An example of television media content on a mobile user device is shown.

[0023] Figure 12 An example of television control using a virtual assistant is shown.

[0024] Figure 13 Exemplary screen and video content on a mobile user device are shown.

[0025] Figure 14 An example media display control using a virtual assistant is shown.

[0026] Figure 15 An exemplary virtual assistant interaction with results is shown on a mobile user device and a media display device.

[0027] Figure 16 An exemplary virtual assistant interaction with media results is shown on a media display device and a mobile user device.

[0028] Figure 17 An example of proximity-based media device control is shown.

[0029] Figure 18An exemplary process for controlling television interaction using a virtual assistant and multiple user devices is shown.

[0030] Figure 19 An exemplary voice input interface with a virtual assistant query for background video content is shown.

[0031] Figure 20 An example informational virtual assistant is shown above the video content.

[0032] Figure 21 An exemplary voice input interface is shown, featuring a virtual assistant for querying media content associated with background video content.

[0033] Figure 22 An exemplary virtual assistant response interface with selectable media content is shown.

[0034] Figures 23A-23B An example page of the program menu is shown.

[0035] Figure 24 An example media menu is shown, divided into categories.

[0036] Figure 25 An exemplary process is shown for controlling television interaction and viewing history of media content using media content displayed on the display.

[0037] Figure 26 An exemplary interface with virtual assistant query suggestions based on background video content is shown.

[0038] Figure 27 An exemplary interface for confirming the suggested query selection is shown.

[0039] Figures 28A-28B An exemplary virtual assistant response interface based on the selected query is shown.

[0040] Figure 29 An exemplary interface is shown, displaying media content notifications and virtual assistant query suggestions based on those notifications.

[0041] Figure 30 A mobile user device with exemplary graphics and video content playable on a media control device is shown.

[0042] Figure 31 An exemplary mobile user device interface is shown, which provides virtual assistant query suggestions based on playable user device content and video content displayed on a separate display.

[0043] Figure 32An exemplary interface is shown, featuring virtual assistant query suggestions based on playable content from individual user devices.

[0044] Figure 33 An exemplary process for suggesting interactions between a virtual assistant and another user to control media content is shown.

[0045] Figure 34 A functional block diagram of an electronic device is shown, which is configured to use a virtual assistant to control television interaction and display related information using different interfaces according to various embodiments.

[0046] Figure 35 A functional block diagram of an electronic device is shown, which is configured to control television interaction using a virtual assistant and multiple user devices according to various embodiments.

[0047] Figure 36 A functional block diagram of an electronic device is shown, which is configured to control television interaction using media content displayed on a display and viewing history of the media content according to various embodiments.

[0048] Figure 37 A functional block diagram of an electronic device is shown, which is configured to suggest virtual assistant interactions for controlling media content according to various embodiments. Detailed Implementation

[0049] The accompanying drawings will be referenced in the following description of the embodiments, in which specific examples that may be implemented are illustrated by way of example. It should be understood that other embodiments may be used and structural changes may be made without departing from the scope of the various embodiments.

[0050] This invention relates to systems and processes for controlling television user interaction using a virtual assistant. In one embodiment, the virtual assistant can be used to interact with a media control device, such as a set-top box that controls content displayed on a television display. Voice input to the virtual assistant can be received using a mobile user device or a remote control with a microphone. The user's intent can be determined from the voice input, and the virtual assistant can perform tasks based on the user's intent, including causing media playback on a connected television and controlling any other functions of the set-top box or similar device (e.g., managing video recording, searching media content, navigating in menus, etc.).

[0051] The virtual assistant interaction can be displayed on a connected television or other display. In one embodiment, media content can be determined based on voice input received from a user. A first user interface with a first small size can be displayed, including selectable links to the determined media content. After receiving a selection of a media link, a second user interface with a second larger size can be displayed, including the media content associated with that selection. In other embodiments, the interface for transmitting the virtual assistant interaction can expand or collapse to occupy a minimal amount of space while transmitting the desired information.

[0052] In some embodiments, multiple devices associated with multiple displays can be used to determine user intent from voice input and to transmit information to the user in different ways. For example, voice input can be received from the user at a first device having a first display. User intent can be determined from voice input based on the content displayed on the first display. Media content can be determined based on user intent and can be played on a second device associated with a second display.

[0053] The content displayed on the television screen can also be used as contextual input to determine the user's intent from the voice input. For example, voice input can be received from the user, including queries related to the content displayed on the television screen. The user's intent for the query can be determined based on the content displayed on the television screen and the viewing history of the media content on the television screen (e.g., to disambiguate the query based on actors in a television program). The query results can then be displayed based on the determined user intent.

[0054] In some embodiments, virtual assistant query suggestions can be provided to the user (e.g., informing the user that commands are available, suggesting content of interest, etc.). For example, media content can be displayed on the screen, and input requesting virtual assistant query suggestions can be received from the user. Virtual assistant query suggestions can be determined based on the media content displayed on the screen and the viewing history of the media content displayed on the screen (e.g., suggesting queries related to playing television programs). The suggested virtual assistant queries can then be displayed on the screen.

[0055] Using a virtual assistant to control a television user interaction, as described in the various embodiments discussed herein, can provide an efficient and enjoyable user experience. Using a virtual assistant capable of receiving natural language queries or commands, user interaction with media control devices can be intuitive and simple. Available functions can be suggested to the user as needed, including meaningful query suggestions based on the content being played, thus assisting the user in learning control capabilities. Furthermore, intuitive spoken commands can be used to make available media easily accessible. However, it should be understood that many other advantages can still be achieved according to the various embodiments discussed herein.

[0056] Figure 1 An exemplary system 100 for controlling television user interaction using a virtual assistant is illustrated. It should be understood that controlling television user interaction as described herein is merely one example of controlling media on a display technology and is for reference only. The concepts described herein can be generally applied to controlling any media content interaction, including any content on various devices and associated displays (e.g., monitors, laptop displays, desktop computer monitors, mobile user device displays, projector displays, etc.). Therefore, the term "television" can refer to any type of display associated with any of the various devices. Furthermore, the terms "virtual assistant," "digital assistant," "intelligent automated assistant," or "automatic digital assistant" can refer to any information processing system that interprets natural language input in spoken and / or textual form to infer user intent and performs actions based on the inferred user intent. For example, in order to perform an action according to the inferred user intent, the system may perform one or more of the following: identify a task flow by utilizing steps and parameters specifically designed to achieve the inferred user intent; input the specific requirements from the inferred user intent into the task flow; execute the task flow by calling programs, methods, services, APIs, etc.; and generate an output response in the form of auditory (e.g., voice) and / or visual responses to the user.

[0057] Virtual assistants can accept user requests, at least in part, in the form of natural language commands, requests, statements, narration, and / or inquiries. Typically, user requests either seek an informational response from the virtual assistant or seek to perform a task through the virtual assistant (e.g., causing specific media to be displayed). A satisfactory response to a user request may include providing the requested informational response, performing the requested task, or a combination of both. For example, a user may ask a virtual assistant a question such as, “Where am I now?” Based on the user’s current location, the virtual assistant might answer, “You are in Central Park.” A user may also request to perform a task, such as, “Please remind me to call my mother at 4 p.m. today.” In response, the virtual assistant can acknowledge the request and then create an appropriate reminder in the user’s electronic calendar. During the performance of the requested task, the virtual assistant may interact with the user in a continuous conversation involving multiple exchanges of information, sometimes for extended periods. Many other methods exist for interacting with a virtual assistant to request information or perform various tasks. In addition to providing verbal responses and taking programmed actions, virtual assistants may also provide responses in other visual or audio forms (e.g., as text, alerts, music, video, animation, etc.). Furthermore, as described herein, the exemplary virtual assistant is able to control the playback of media content (e.g., playing video on a television) and cause information to be displayed on a monitor.

[0058] An example of a virtual assistant is described in U.S. Utility Model Patent Application No. 12 / 987,982, filed January 10, 2011, entitled “Intelligent Automated Assistant,” the entire disclosure of which is incorporated herein by reference.

[0059] like Figure 1 As shown, in some embodiments, the virtual assistant may be implemented according to a client-server model. The virtual assistant may include a client-side portion executing on user equipment 102 and a server-side portion executing on server system 110. The client-side portion may also execute on set-top box 104 in conjunction with remote control 106. User equipment 102 may include any electronic device, such as a mobile phone (e.g., a smartphone), tablet computer, portable media player, desktop computer, laptop computer, PDA, wearable electronic device (e.g., digital glasses, wristband, watch, brooch, armband, etc.). Set-top box 104 may include any media control device, such as a cable box, satellite box, video player, video streaming device, digital video recorder, gaming system, DVD player, Blu-ray disc. TM Players, combinations of such devices, etc. The set-top box 104 can be connected to the display 112 and speakers 111 via wired or wireless connections. The display 112 (with or without speakers 111) can be any type of display, such as a television monitor, a projector, etc. In some embodiments, the set-top box 104 can be connected to an audio system (e.g., an audio receiver), and the speakers 111 can be independent of the display 112. In other embodiments, the display 112, speakers 111, and set-top box 104 can be integrated into a single device, such as a smart television with advanced processing and networking capabilities. In such embodiments, the functionality of the set-top box 104 can be performed as an application on a combined device.

[0060] In some embodiments, the set-top box 104 can act as a media control center for media content of various types and sources. For example, the set-top box 104 can provide convenient access to live television (e.g., over-the-air, satellite, or cable television). Thus, the set-top box 104 may include a cable tuner, a satellite tuner, etc. In some embodiments, the set-top box 104 may also record television programs for viewing at a later time. In other embodiments, the set-top box 104 may provide access to one or more streaming media services, such as cable-transmitted on-demand television programs, videos, and music, and internet-transmitted television programs, videos, and music (e.g., from various free, paid, and subscription-based streaming services). In other examples, the set-top box 104 can conveniently play back or display media content from any other source, such as displaying photos from a mobile user device, playing videos from a coupled storage device, playing music from a coupled music player, etc. The set-top box 104 may also include various other combinations of the media control features discussed herein, as needed.

[0061] User equipment 102 and set-top box 104 can communicate with server system 110 via one or more networks 108, which may include the Internet, intranet, or any other wired or wireless public or private network. Furthermore, user equipment 102 can communicate with set-top box 104 via network 108 or directly via any other wired or wireless communication mechanism (e.g., Bluetooth, Wi-Fi, RF, infrared transmission, etc.). As shown, remote control 106 can communicate with set-top box 104 using any type of communication, such as a wired connection or any type of wireless communication (e.g., Bluetooth, Wi-Fi, RF, infrared transmission, etc.), including via network 108. In some examples, the user can interact with set-top box 104 via user equipment 102, remote control 106, or interface elements integrated within set-top box 104 (e.g., buttons, microphone, camera, joystick, etc.). For example, voice input, including media-related queries or commands for a virtual assistant, can be received at user device 102 and / or remote control 106, and the voice input can be used to perform media-related tasks on the set-top box 104. Similarly, tactile commands for controlling media on the set-top box 104 can be received at user device 102 and / or remote control 106 (and from other devices not shown). Therefore, various functions of the set-top box 104 can be controlled in various ways, providing the user with multiple options for controlling media content from multiple devices.

[0062] The client-side portion of the exemplary virtual assistant, which executes on user equipment 102 and / or a set-top box 104 with remote control 106, is capable of providing client-side functionality, such as user-oriented input and output processing and communication with server system 110. Server system 110 can provide server-side functionality for any number of clients residing on the respective user equipment 102 or the respective set-top box 104.

[0063] Server system 110 may include one or more virtual assistant servers 114, each including a client-facing I / O interface 122, one or more processing modules 118, a data and model storage device 120, and an I / O interface 116 to external services. The client-facing I / O interface 122 facilitates client-facing input and output processing for the virtual assistant server 114. One or more processing modules 118 may utilize the data and model storage device 120 to determine the user's intent based on natural language input and perform task execution based on the inferred user intent. In some examples, the virtual assistant server 114 may communicate with external services 124 via network 108, such as telephone services, calendar services, information services, instant messaging services, navigation services, television program services, streaming media services, etc., to complete tasks or obtain information. The I / O interface 116 to external services facilitates such communication.

[0064] Server system 110 may be implemented on one or more standalone data processing devices or distributed networks. In some examples, server system 110 may utilize various virtual devices and / or services from third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing and / or infrastructure resources of server system 110.

[0065] although Figure 1 The virtual assistant's functionality is presented as including both client-side and server-side components. However, in some examples, the assistant's functions (or general voice recognition and media control) can be implemented as a standalone application installed on user devices, set-top boxes, smart TVs, etc. Furthermore, the functional division between the client and server components of the virtual assistant can vary in different examples. For instance, in some examples, the client running on user device 102 or set-top box 104 can be a thin client, providing only user-facing input and output processing functions and delegating all other virtual assistant functions to the backend server.

[0066] Figure 2A block diagram of an exemplary user equipment 102 according to various embodiments is shown. As shown, user equipment 102 may include a memory interface 202, one or more processors 204, and a peripheral device interface 206. Various components in user equipment 102 may be coupled together by one or more communication buses or signal lines. User equipment 102 may further include various sensors, subsystems, and peripheral devices coupled to the peripheral device interface 206. Sensors, subsystems, and peripheral devices acquire information and / or facilitate various functions of user equipment 102.

[0067] For example, user equipment 102 may include a motion sensor 210, a light sensor 212, and a proximity sensor 214 coupled to the peripheral device interface 206 to facilitate orientation, illumination, and proximity sensing functions. One or more other sensors 216, such as positioning systems (e.g., GPS receivers), temperature sensors, biometric sensors, gyroscopes, compasses, accelerometers, etc., may also be connected to the peripheral device interface 206 to facilitate related functions.

[0068] In some examples, camera subsystem 220 and optical sensor 222 can be used to facilitate camera functions such as taking photos and recording video clips. Communication functions can be facilitated by one or more wired and / or wireless communication subsystems 224, which may include various communication ports, radio frequency receivers and transmitters, and / or optical (such as infrared) receivers and transmitters. Audio subsystem 226 can be coupled to speaker 228 and microphone 230 to facilitate voice-enabled functions such as speech recognition, voice copying, digital recording, and telephone functions.

[0069] In some examples, user equipment 102 may further include an I / O subsystem 240 coupled to peripheral device interface 206. I / O subsystem 240 may include touchscreen controller 242 and / or other input controllers 244. Touchscreen controller 242 may be coupled to touchscreen 246. Touchscreen 246 and touchscreen controller 242 may use any of a variety of touch sensitivity technologies, such as capacitive, resistive, infrared and surface acoustic wave technologies, proximity sensor arrays, etc., to detect contact and movement or their interruption. Other input controllers 244 may be coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointing devices (such as styluses).

[0070] In some examples, user equipment 102 may also include a memory interface 202 coupled to memory 250. Memory 250 may include any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device; portable computer disk (magnetic), random access memory (RAM) (magnetic); read-only memory (ROM) (magnetic); erasable programmable read-only memory (EPROM) (magnetic); portable optical discs such as CD, CD-R, CD-RW, DVD, DVD-R, or DVD-RW; or flash memory such as compact flash memory cards, secure digital cards, USB storage devices, memory sticks, etc. In some examples, the non-transitory computer-readable storage medium of memory 250 may be used to store instructions (e.g., for performing some or all of the various processes described herein) for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-containing system, or other system capable of fetching and executing instructions from and from an instruction execution system, apparatus, or device. In other examples, instructions (e.g., for performing some or all of the various processes described herein) may be stored on a non-transitory computer-readable storage medium of server system 110, or may be allocated between a non-transitory computer-readable storage medium of memory 250 and a non-transitory computer-readable storage medium of server system 110. In the context of this document, "non-transitory computer-readable storage medium" can be any medium that may include or store programs for use by or in connection with an instruction execution system, apparatus, or device.

[0071] In some examples, memory 250 may store operating system 252, communication module 254, graphical user interface module 256, sensor processing module 258, telephone module 260, and applications 262. Operating system 252 may include instructions for processing basic system services and for performing hardware-related tasks. Communication module 254 facilitates communication with one or more additional devices, one or more computers, and / or one or more servers. Graphical user interface instructions 256 facilitate graphical user interface processing. Sensor processing module 258 facilitates sensor-related processing and functions. Telephone module 260 facilitates telephone-related processes and functions. Application module 262 facilitates various functions of user applications, such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions.

[0072] As described herein, memory 250 may also store client-side virtual assistant commands (e.g., in virtual assistant client module 264) and various user data 266 (e.g., user-specific vocabulary data, preference data, and / or other data such as the user's electronic address book, to-do list, shopping list, TV program collection, etc.) to, for example, provide client-side functionality for the virtual assistant. User data 266 may also be used for speech recognition when supporting the virtual assistant or for any other application.

[0073] In various examples, the virtual assistant client module 264 can accept sound input (e.g., voice input), text input, touch input, and / or gesture input through various user interfaces of the user device 102 (e.g., I / O subsystem 240, audio subsystem 226, etc.). The virtual assistant client module 264 can also provide output in the form of audio (e.g., voice output), visual, and / or haptic feedback. For example, the output can be provided as voice, sound, alarms, text messages, menus, graphics, video, animation, vibration, and / or a combination of both or more of these. During operation, the virtual assistant client module 264 uses the communication subsystem 224 to communicate with the virtual assistant server.

[0074] In some examples, the virtual assistant client module 264 may utilize various sensors, subsystems, and peripherals to collect additional information from the user device 102's surrounding environment to establish contextual information associated with the user, current user interaction, and / or current user input. This context may also include information from other devices, such as the set-top box 104. In some examples, the virtual assistant client module 264 may provide contextual information, or a subset thereof, along with user input to the virtual assistant server to help infer the user's intent. The virtual assistant may also use the contextual information to determine how to prepare output and deliver it to the user. This contextual information may also be used by the user device 102 or the server system 110 to support accurate speech recognition.

[0075] In some examples, contextual information accompanying user input may include sensor information such as lighting, ambient noise, ambient temperature, images or videos of the surrounding environment, distance to another object, etc. Contextual information may also include information associated with the physical state of user device 102 (e.g., device orientation, device location, device temperature, power level, speed, acceleration, motion pattern, cellular signal strength, etc.) or the software state of user device 102 (e.g., operation process, installed programs, past and present network activity, background services, error logs, resource usage, etc.). Contextual information may also include information associated with the state of connected devices or other devices associated with the user (e.g., media content displayed on set-top box 104, media content available for set-top box 104, etc.). Any of these types of contextual information may be provided to virtual assistant server 114 (or used on user device 102 itself) as contextual information associated with user input.

[0076] In some examples, the virtual assistant client module 264 may selectively provide information stored on the user device 102 (e.g., user data 266) in response to a request from the virtual assistant server 114 (or may use it on the user device 102 itself when performing speech recognition and / or virtual assistant functions). The virtual assistant client module 264 may also elicit additional input from the user via natural language dialogue or other user interfaces upon request from the virtual assistant server 114. The virtual assistant client module 264 may pass this additional input to the virtual assistant server 114 to assist the virtual assistant server 114 in intent inference and / or fulfilling the user intent expressed in the user request.

[0077] In various examples, memory 250 may include additional instructions or fewer instructions. Furthermore, various functions of user equipment 102 may be executed in hardware and / or firmware, including in one or more signal processing and / or application-specific integrated circuits.

[0078] Figure 3A block diagram of an exemplary set-top box 104 in a system 300 for controlling user interaction with a television set is shown. System 300 may include a subset of the elements of system 100. In some examples, system 300 may perform specific functions independently and be able to work with other elements of system 100 to perform other functions. For example, elements of system 300 may handle specific media control functions without interacting with server system 110 (e.g., playing back locally stored media, recording functions, channel tuning, etc.), while system 300 may combine server system 110 and other elements of system 100 to handle other media control functions (e.g., playing back remotely stored media, downloading media content, handling specific virtual assistant queries, etc.). In other examples, elements of system 300 may perform functions of the larger system 100, including accessing external service 124 over a network. It should be understood that functions can be partitioned between local devices and remote server devices in various other ways.

[0079] like Figure 3 As shown, in one example, a set-top box 104 may include a memory interface 302, one or more processors 304, and a peripheral device interface 306. Various components in the set-top box 104 may be coupled together by one or more communication buses or signal lines. The set-top box 104 may further include various subsystems and peripheral devices coupled to the peripheral device interface 306. The subsystems and peripheral devices may acquire information and / or facilitate various functions of the set-top box 104.

[0080] For example, a television set-top box 104 may include a communication subsystem 324. Communication functionality may be facilitated by one or more wired and / or wireless communication subsystems 324, which may include various communication ports, radio frequency receivers and transmitters, and / or optical (such as infrared) receivers and transmitters.

[0081] In some examples, the set-top box 104 may also include an I / O subsystem 340 coupled to the peripheral interface 306. The I / O subsystem 340 may include an audio / video output controller 370. The audio / video output controller 370 may be coupled to the display 112 and the speaker 111 or may otherwise provide audio and video output (e.g., via an audio / video port, wireless transmission, etc.). The I / O subsystem 340 may also include a remote controller 342. The remote controller 342 may be communicatively coupled to the remote controller 106 (e.g., via a wired connection, Bluetooth, Wi-Fi, etc.). The remote controller 106 may include a microphone 372 for capturing audio input (e.g., voice input from a user), buttons 374 for capturing tactile input, and a transceiver 376 for facilitating communication with the set-top box 104 via the remote controller 342. The remote controller 106 may also include other input mechanisms such as a keyboard, joystick, touchpad, etc. The remote controller 106 may also include output mechanisms such as lights, displays, speakers, etc. Input received at remote control 106 (e.g., user voice, button presses, etc.) can be transmitted to set-top box 104 via remote controller 342. I / O subsystem 340 may also include other input controllers 344. These other input controllers 344 can be coupled to other input / control devices 348, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointing devices (such as styluses).

[0082] In some examples, the set-top box 104 may also include a memory interface 302 coupled to the memory 350. The memory 350 may include any electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device; a portable computer disk (magnetic), random access memory (RAM) (magnetic); read-only memory (ROM) (magnetic); erasable programmable read-only memory (EPROM) (magnetic); portable optical discs such as CD, CD-R, CD-RW, DVD, DVD-R, or DVD-RW; or flash memory such as compact flash memory cards, secure digital cards, USB storage devices, memory sticks, etc. In some examples, the non-transitory computer-readable storage medium of the memory 350 may be used to store instructions (e.g., for performing some or all of the various processes described herein) for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-containing system, or other systems capable of fetching and executing instructions from and from an instruction execution system, apparatus, or device. In other examples, instructions (e.g., for performing some or all of the various processes described herein) may be stored on a non-transitory computer-readable storage medium of server system 110, or may be allocated between a non-transitory computer-readable storage medium of memory 350 and a non-transitory computer-readable storage medium of server system 110. In the context of this document, "non-transitory computer-readable storage medium" can be any medium that may include or store programs for use by or in connection with an instruction execution system, apparatus, or device.

[0083] In some examples, memory 350 may store operating system 352, communication module 354, graphical user interface module 356, on-device media module 358, off-device media module 360, and application 362. Operating system 352 may include instructions for handling basic system services and for performing hardware-related tasks. Communication module 354 facilitates communication with one or more additional devices, one or more computers, and / or one or more servers. Graphical user interface instructions 356 facilitate graphical user interface processing. On-device media module 358 facilitates the storage and playback of media content locally stored on set-top box 104 and other media content that can be obtained locally (e.g., cable channel tuning). Off-device media module 360 ​​facilitates the streaming or downloading of media content remotely stored (e.g., on a remote server, on user device 102, etc.). Application module 362 facilitates various functions of user applications, such as electronic messaging, web browsing, media processing, games, and / or other processes and functions.

[0084] As described herein, memory 350 may also store client-side virtual assistant commands (e.g., in virtual assistant client module 364) and various user data 366 (e.g., user-specific vocabulary data, preference data, and / or other data such as the user's electronic address book, to-do list, shopping list, TV program collection, etc.) to, for example, provide client-side functionality for the virtual assistant. User data 366 may also be used for speech recognition when supporting the virtual assistant or for any other application.

[0085] In various examples, the virtual assistant client module 364 can accept sound input (e.g., voice input), text input, touch input, and / or gesture input through various user interfaces of the set-top box 104 (e.g., I / O subsystem 340, etc.). The virtual assistant client module 364 can also provide output in the form of audio (e.g., voice output), visual, and / or haptic feedback. For example, the output can be provided as voice, sound, alarms, text messages, menus, graphics, video, animation, vibration, and / or a combination of both or more of these. During operation, the virtual assistant client module 364 uses the communication subsystem 324 to communicate with the virtual assistant server.

[0086] In some examples, the virtual assistant client module 364 may utilize various subsystems and peripherals to gather additional information from the surrounding environment of the set-top box 104 to establish a context associated with the user, current user interaction, and / or current user input. This context may also include information from other devices, such as user device 102. In some examples, the virtual assistant client module 364 may provide the contextual information, or a subset thereof, along with the user input to the virtual assistant server to help infer the user's intent. The virtual assistant may also use the contextual information to determine how to prepare output and deliver it to the user. This contextual information may also be used by the set-top box 104 or the server system 110 to support accurate speech recognition.

[0087] In some examples, contextual information accompanying user input may include sensor information such as lighting, ambient noise, ambient temperature, distance to another object, etc. Contextual information may also include information associated with the physical state of the set-top box 104 (e.g., device location, device temperature, power level, etc.) or the software state of the set-top box 104 (e.g., operating process, installed programs, past and present network activity, background services, error logs, resource usage, etc.). Contextual information may also include information associated with the state of connected devices or other devices associated with the user (e.g., content displayed on user device 102, content playable on user device 102, etc.). Any of these types of contextual information may be provided to the virtual assistant server 114 (or used on the set-top box 104 itself) as contextual information associated with user input.

[0088] In some examples, the virtual assistant client module 364 may selectively provide information stored on the set-top box 104 (e.g., user data 366) in response to a request from the virtual assistant server 114 (or it may be used on the set-top box 104 itself when performing voice recognition and / or virtual assistant functions). The virtual assistant client module 364 may also elicit additional input from the user via natural language dialogue or other user interfaces when requested by the virtual assistant server 114. The virtual assistant client module 364 may transmit this additional input to the virtual assistant server 114 to assist the virtual assistant server 114 in intent inference and / or fulfilling the user intent expressed in the user request.

[0089] In various examples, memory 350 may include additional instructions or fewer instructions. Furthermore, various functions of the set-top box 104 may be performed in hardware and / or firmware, including in one or more signal processing and / or application-specific integrated circuits.

[0090] It should be understood that System 100 and System 300 are not limited to Figure 1 and Figure 3 The components and configurations shown, including user equipment 102, set-top box 104, and remote control 106, are similarly not limited to these. Figure 2 and Figure 3 The components and configurations shown. System 100, system 300, user equipment 102, set-top box 104, and remote control 106 may all include fewer or other components in a variety of configurations according to various examples.

[0091] Throughout this disclosure, the term "system" may include system 100, system 300, or one or more elements of system 100 or system 300. For example, a typical system mentioned herein may include at least a television set-top box 104 that receives user input from remote control 106 and / or user equipment 102.

[0092] Figures 4A to 4E An exemplary voice input interface 484 is shown, which can be displayed on a display (e.g., display 112) to transmit voice input information to a user. In one embodiment, the voice input interface 484 can be displayed on video 480, which may include any moving image or paused video. For example, video 480 may include a live television program, playing video, streaming movie, playback of a recorded program, etc. The voice input interface 484 can be configured to occupy minimal space so as not to significantly interfere with the user's viewing of video 480.

[0093] In one embodiment, a virtual assistant can be triggered to listen for voice input containing commands or queries (or to begin recording voice input for subsequent processing, or to begin processing voice input in real time). Listening can be triggered in various ways, including by indication, such as a user pressing a physical button on remote control 106, a user pressing a physical button on user device 102, a user pressing a virtual button on user device 102, a user uttering a phrase recognizable by the always-listening device (e.g., uttering "Hello, Assistant" to initiate listening), a user performing a gesture detectable by sensors (e.g., monitoring in front of a camera), etc. In another embodiment, a user can press and hold a physical button on remote control 106 or user device 102 to initiate listening. In other embodiments, a user can press and hold a physical button on remote control 106 or user device 102 while speaking a query or command, and be able to release the button upon completion. Various other indications can similarly be received to initiate receiving voice input from the user.

[0094] In response to receiving an instruction to listen for voice input, the voice input interface 484 can be displayed. Figure 4A A notification area 482 is shown that expands upward from the bottom portion of the display 112. A voice input interface 484 can be displayed in the notification area 482 when an instruction to listen for voice input is received. This interface can be animated to slide upward from the bottom edge of the viewing area of ​​the display 112 shown. Figure 4B A voice input interface 484 is shown after sliding up to view it. The voice input interface 484 can be configured to occupy minimal space at the bottom of the display 112 to avoid significantly interfering with the video 480. In response to receiving an indication to listen for voice input, a readiness confirmation 486 can be displayed. The readiness confirmation 486 may include the microphone symbol shown, or may include any other image, icon, animation, or symbol to convey that the system (e.g., one or more elements of system 100) is ready to capture voice input from the user.

[0095] When the user begins to speak, it can display Figure 4C The listening confirmation 487 shown confirms that the system is capturing voice input. In some embodiments, listening confirmation 487 may be displayed in response to receiving voice input (e.g., capturing voice). In other embodiments, a ready confirmation 486 may be displayed for a predetermined amount of time (e.g., 500 milliseconds, 1 second, 3 seconds, etc.), after which listening confirmation 487 may be displayed. Listening confirmation 487 may include the waveform symbol shown, or may include an active waveform animation that moves (e.g., changes frequency) in response to user voice. In other embodiments, listening confirmation 487 may include any other image, icon, animation, or symbol to convey that the system is capturing voice input from the user.

[0096] When it is detected that the user has finished speaking (e.g., based on a pause, a voice interruption indicating the end of a query, or any other endpoint detection method), a message can be displayed. Figure 4D The processing confirmation 488 shown confirms that the system has completed capturing the voice input and is processing it (e.g., interpreting the voice input, determining the user's intent, and / or performing an associated task). Processing confirmation 488 may include the hourglass symbol shown, or may include any other image, icon, animation, or symbol to convey that the system is processing the captured voice input. In another embodiment, processing confirmation 488 may include an animation of a rotating circle or a colored / glowing dot moving around the circle.

[0097] After the captured voice input is interpreted as text (or in response to successfully converting the voice input to text), it can be displayed. Figure 4E The command reception confirmation 490 and / or transcript 492 shown acknowledge that the system has received and interpreted the voice input. Transcript 492 may include a transcript of the received voice input (e.g., “What sporting events are happening now?”). In some embodiments, transcript 492 may be animated to slide upwards from the bottom of display 112. Figure 4E The location shown briefly displays the transcript (e.g., for a few seconds) before it can be swiped to the top of the voice input interface 484 before disappearing from view (e.g., as if text scrolls up and eventually goes out of view). In other embodiments, the transcript may not be displayed, user commands or queries may be processed, and associated tasks may be performed without displaying the transcript (e.g., a simple channel change may be performed immediately without displaying a transcript of the user's voice).

[0098] In other embodiments, speech transcription can be performed in real time as the user speaks. When the text is transcribed, it can be displayed in the voice input interface 484. For example, the text can be displayed side-by-side with the listening confirmation 487. After the user finishes speaking, a brief command reception confirmation 490 can be displayed before performing the task associated with the user's command.

[0099] Furthermore, in other embodiments, command reception acknowledgment 490 can convey information about a command that has been received and understood. For example, for a simple request to change to another channel, a logo or number associated with the channel can be briefly displayed as command reception acknowledgment 490 (e.g., for a few seconds) when the channel is changed. In another embodiment, for a request to pause video (e.g., video 480), a pause symbol (e.g., two vertical parallel bars) can be displayed as command reception acknowledgment 490. The pause symbol can remain on the display until, for example, the user performs another action (e.g., issues a play command to resume playback). Symbols, logos, etc. (e.g., symbols for rewind, fast forward, stop, play, etc.) can be similarly displayed for any other command. Thus, command reception acknowledgment 490 can be used to convey command-specific information.

[0100] In some embodiments, the voice input interface 484 may be hidden after a user query or command is received. For example, the voice input interface 484 may be animated to slide downwards until it reaches outside the view at the bottom of the display 112. The voice input interface 484 may be hidden without needing to display further information to the user. For example, for common or direct commands (e.g., changing the channel to channel 10, changing to a sports channel, play, pause, fast forward, rewind, etc.), the voice input interface 484 may be hidden immediately after receiving a confirmation command, and the associated task may be executed immediately. Although various embodiments herein exemplify and describe interfaces at the bottom or top edge of the display, it should be understood that any of the various interfaces may be located at other locations around the display. For example, the voice input interface 484 may appear from the side edge of the display 112, in the middle of the display 112, or in a corner of the display 112, etc. Similarly, various other interface examples described herein may be arranged at various locations on the display in various orientations. Furthermore, although the interfaces described herein are illustrated as opaque, any of the interfaces may be transparent or otherwise allow images (blurred or fully) to be seen through the interface (e.g., interface content overlaying media content without completely blurring the media content underneath).

[0101] In other examples, query results can be displayed within the voice input interface 484 or in a different interface. Figure 5 An exemplary media content interface 510 is shown above video 480, having Figure 4E Examples of transcription queries are provided. In some examples, virtual assistant query results may include media content in place of or as a supplement to text content. For example, virtual assistant query results may include television programs, videos, music, etc. Some results may include media that can be played back immediately, while others may include media that can be purchased, etc.

[0102] As shown in the figure, the media content interface 510 can be larger than the voice input interface 484. In one embodiment, the voice input interface 484 can be a smaller first size to accommodate voice input information, while the media content interface 510 can be a larger second size to accommodate query results, which may include text, still images, and moving images. In this way, the interface used to transmit virtual assistant information can scale its size according to the content to be transmitted, thereby limiting intrusion into the screen footprint (e.g., minimizing obstruction of other content, such as video 480).

[0103] As shown, the media content interface 510 may include selectable video links 512, selectable text links 514, and additional content links 513 (as a virtual assistant query result). In some embodiments, links can be selected by navigating to a specific element with focus, cursor, etc., and selecting it using a remote control (e.g., remote control 106). In other embodiments, links can be selected using voice commands given to the virtual assistant (e.g., watch that football game, show details about the basketball game, etc.). Selectable video links 512 may include still or moving images and may be selectable to play back associated video. In one embodiment, selectable video links 512 may include a video playing associated video content. In another embodiment, selectable video links 512 may include live feeds of television channels. For example, as a virtual assistant query result about a current sporting event on a television, selectable video links 512 may include live feeds of a football game on a sports channel. Selectable video links 512 may also include any other videos, animations, images, etc. (e.g., triangle play symbols). In addition, Link512 can link to any type of media content such as movies, television programs, sports events, music, etc.

[0104] Selectable text link 514 may include text content associated with selectable video link 512 or may include text representations of virtual assistant query results. In one example, selectable text link 514 may include a description of media obtained from the virtual assistant query. For example, selectable text link 514 may include the name of a television program, a movie title, a description of a sporting event, a television channel name or number, etc. In one embodiment, selecting text link 514 may cause the associated media content to be played back. In another embodiment, selecting text link 514 may provide additional detailed information about media content or other virtual assistant query results. Additional content link 513 may link to additional results of the virtual assistant query and cause them to be displayed.

[0105] although Figure 5Examples of specific media content are shown, but it should be understood that any type of media content can be included as a result of a virtual assistant query targeting media content. For example, media content that can be returned as a result of a virtual assistant may include videos, television programs, music, television channels, etc. Furthermore, in some embodiments, category filters may be provided in any interface herein to allow users to filter search or query results or displayed media options. For example, selectable filters may be provided to filter results based on type (e.g., movies, music albums, books, television programs, etc.). In other embodiments, selectable filters may include genre or content descriptors (e.g., comedy, interviews, specific programs, etc.). In other embodiments, selectable filters may include time (e.g., this week, last week, last year, etc.). It should be understood that filters may be provided in any of the various interfaces described herein to allow users to filter results based on categories related to the displayed content (e.g., filtering by type when media results have various types, filtering by genre when media results have various genres, filtering by time when media results have various times, etc.).

[0106] In other embodiments, the media content interface 510 may include query paraphrases in addition to the media content results. For example, a paraphrase of the user's query may be displayed above the media content results (above the selectable video link 512 and the selectable text link 514). Figure 5 In the example, such a user query paraphrase could include something like, "Here are some sporting events that are currently underway." Other text describing the media content results could be displayed similarly.

[0107] In some embodiments, after any interface, including interface 510, is displayed, a user can initiate the capture of additional voice input using a new query (which may be related to or unrelated to a previous query). The user query may include a command to act on interface elements, such as a command to select video link 512. In another embodiment, the user's voice may include a query associated with displayed content, such as displayed menu information, playing a video (e.g., video 480), etc. A response to such a query may be determined based on displayed information (e.g., displayed text) and / or metadata associated with the displayed content (e.g., metadata associated with playing a video). For example, a user may query for media results displayed in an interface (e.g., interface 510) and may search for metadata associated with that media to provide an answer or result. Such an answer or result may then be provided in another interface or within the same interface (e.g., in any interface discussed herein).

[0108] As described above, in one embodiment, additional details about the media content may be displayed in response to the selection of text link 514. Figure 6A and Figure 6B An exemplary media details interface 618 is shown above video 480 after text link 514 is selected. In one embodiment, media content interface 510 can be extended to media details interface 618 when providing additional details, such as... Figure 6A The interface extension transformation is shown in 616. Specifically, as shown in... Figure 6A As shown, the selected content can be expanded in size, and additional text information can be provided by expanding the interface upwards on display 112 to occupy more screen space. The interface can be expanded to accommodate additional detailed information desired by the user. In this way, the size of the interface can scale with the amount of content desired by the user, thereby minimizing screen space intrusion while still delivering the expected content.

[0109] Figure 6B The fully expanded details interface 618 is shown. As illustrated, the details interface 618 can be larger than the media content interface 510 or the voice input interface 484 to accommodate the desired detailed information. The details interface 618 may include detailed media information 622, which may include various details associated with another result of the media content or virtual assistant query. Detailed media information 622 may include program titles, program descriptions, program broadcast times, channels, episode summaries, movie descriptions, actor names, character names, sports event participants, producer names, director names, or any other details associated with the virtual assistant query results.

[0110] In one embodiment, the details interface 618 may include an optional video link 620 (or another link to play media content), which may include a larger version of the corresponding optional video link 512. Thus, the optional video link 620 may include still or moving images and may be optional for playing associated video. The optional video link 620 may include playback video of associated video content, live feeds from television channels (e.g., live feeds of a football match from a sports channel), etc. The optional video link 620 may also include any other video, animation, image, etc. (e.g., a triangle play symbol).

[0111] As described above, a video can be played in response to the selection of a video link, such as video link 620 or video link 512. Figure 7A and Figure 7B An exemplary media transition interface is shown that can be displayed in response to selecting a video link (or other command to play video content). As shown, video 726 can be used instead of video 480. In one embodiment, video 726 can be extended to catch up with or cover video 480, such as... Figure 7AThe interface extension transformation 724 is shown in the diagram. The result of this transformation may include... Figure 7B An extended media interface 728. Like other interfaces, the extended media interface 728 can be large enough to provide the user with the desired information; here, this information may include an extension to fill the display 112. The extended media interface 728 can thus be larger than any other interface because the desired information may include the playing media content across the entire display. Although not shown, in some embodiments, descriptive information may be briefly overlaid on the video 726 (e.g., along the bottom of the screen). Such descriptive information may include the name of the associated program, video, channel, etc. The descriptive information can then be hidden from view (e.g., after a few seconds).

[0112] Figures 8A to 8B An exemplary voice input interface 836 is shown, which can be displayed on display 112 to transmit voice input information to a user. In one embodiment, the voice input interface 836 may be displayed above a menu 830. Menu 830 may include various media options 832, and the voice input interface 836 may similarly be displayed above any other type of menu (e.g., content menu, category menu, control menu, settings menu, program menu, etc.). In one embodiment, the voice input interface 836 may be configured to occupy a larger portion of the screen area of ​​display 112. For example, the voice input interface 836 may be larger than the voice input interface 484 described above. In one embodiment, the size of the voice input interface to be used (e.g., a smaller interface 484 or a larger interface 836) may be determined based on the background content. When the background content includes moving images, for example, a smaller voice input interface (e.g., interface 484) may be displayed. On the other hand, when the background content includes static images (e.g., paused video) or menus, for example, a larger voice input interface (e.g., interface 836) may be displayed. In this way, a smaller voice input interface can be displayed if the user is watching video content, taking up minimal screen space; while a larger voice input interface can be displayed if the user is navigating a menu or watching a paused video or other static image, conveying more information or having a more immersive effect by occupying additional space. Other interfaces discussed in this paper can similarly be sized based on the background content.

[0113] As described above, a virtual assistant can be triggered to listen for voice input containing commands or queries (or to begin recording voice input for subsequent processing, or to begin processing voice input in real time). Listening can be triggered in various ways, including by indication, such as a user pressing a physical button on remote control 106, a user pressing a physical button on user device 102, a user pressing a virtual button on user device 102, a user uttering a phrase recognizable by the always-listening device (e.g., uttering "Hello, Assistant" to initiate listening), a user performing a gesture detectable by sensors (e.g., monitoring in front of a camera), etc. In other embodiments, a user can press and hold a physical button on remote control 106 or user device 102 to initiate listening. In other embodiments, a user can press and hold a physical button on remote control 106 or user device 102 while speaking a query or command, and be able to release the button upon completion. Various other indications can similarly be received to initiate the reception of voice input from the user.

[0114] In response to receiving an instruction to listen for voice input, a voice input interface 836 can be displayed above menu 830. Figure 8A A large notification area 834, expanding upwards from the bottom portion of display 112, is shown. A voice input interface 836 can be displayed in the large notification area 834 upon receiving an indication to listen for voice input. This interface can be animated to slide upwards from the bottom edge of the viewing area of ​​the illustrated display 112. In some embodiments, when displaying overlapping interfaces (e.g., in response to receiving an indication to listen for voice input), background menus, paused videos, still images, or other background content can be minimized and / or moved backwards in the z-direction (as if moving further into display 112). A background interface shrinkage transition 831 and associated inward-pointing arrows illustrate how background content (e.g., menu 830) can be minimized—the displayed menu, image, text, etc. This can provide the visual effect that the background content appears to leave the user and move into a new foreground interface (e.g., interface 836). Figure 8B A collapsed (shrinkable) background interface 833, including a menu 830, is shown. As illustrated, the collapsed background interface 833 (which may include a border) can appear to move further away from the user while shifting focus to the foreground interface 836. Background content (including background video content) in any of the other embodiments described herein can similarly collapse and / or shift backward in the z-direction when displaying overlapping interfaces.

[0115] Figure 8B The voice input interface 836 is shown after sliding up the view. As described above, various confirmations can be displayed while receiving voice input. Although not shown here, the voice input interface 836 can be understood by referring to the above description. Figure 4B , Figure 4Cand Figure 4D The voice input interface 484 displays larger versions of the readiness confirmation 486, listening confirmation 487, and / or processing confirmation 488 in a similar manner.

[0116] like Figure 8B As shown, a command reception acknowledgment 838 (such as the smaller command reception acknowledgment 490 described above) can be displayed to confirm that the system has received and interpreted the speech input. Transcription 840 can also be displayed and may include a transcription of the received speech input (e.g., “What’s the weather like in New York?”). In some embodiments, the transcript 840 can be animated to slide upwards from the bottom of the display 112. Figure 8B The location shown briefly displays the transcript (e.g., for a few seconds) and can then be swiped to the top of the voice input interface 836 before disappearing from view (e.g., as if text scrolls up and eventually goes out of view). In other embodiments, the transcript may not be displayed, user commands or queries may be processed, and associated tasks may be performed without displaying the transcript.

[0117] In other embodiments, speech transcription can be performed in real time as the user speaks. When the text is transcribed, it can be displayed in the voice input interface 836. For example, the text can be displayed side-by-side with a larger version of the aforementioned listening confirmation 487. After the user finishes speaking, a brief command reception confirmation 838 can be displayed before performing the task associated with the user's command.

[0118] Furthermore, in other embodiments, command reception acknowledgment 838 can convey information about a command that has been received and understood. For example, for a simple request to tune to a specific channel, a logo or number associated with the channel may be briefly displayed as command reception acknowledgment 838 (e.g., for a few seconds) while tuning. In another embodiment, for a request to select a displayed menu item (e.g., one of menu options 832), an image associated with the selected menu item may be displayed as command reception acknowledgment 838. Thus, command reception acknowledgment 838 can be used to convey command-specific information.

[0119] In some embodiments, the voice input interface 836 can be hidden after receiving a user query or command. For example, the voice input interface 836 can be animated to slide down until it reaches outside the view at the bottom of the display 112. The voice input interface 836 can be hidden when no further information needs to be displayed to the user. For example, for ordinary or direct commands (e.g., changing the channel to channel 10, changing to a sports channel, playing the movie, etc.), the voice input interface 836 can be hidden immediately after receiving a confirmation command, and the associated task can be executed immediately.

[0120] In other embodiments, the query results may be displayed within the voice input interface 836 or on a different interface. Figure 9 An exemplary virtual assistant results interface 942 is shown above menu 830 (especially above the collapsed background interface 833), featuring... Figure 8B The virtual assistant query results are exemplary results of the transcription query. In some embodiments, the virtual assistant query results may include a text answer, such as text answer 944. The virtual assistant query results may also include media content that resolves the user's query, such as content associated with optional video link 946 and purchase link 948. Specifically, in this embodiment, a user can ask for weather information for a specified location in New York. The virtual assistant can provide a text answer 944 that directly answers the user's query (e.g., indicating that the weather looks good and providing temperature information). As an alternative to or supplement to text answer 944, the virtual assistant may provide optional video link 946 along with purchase link 948 and associated text. The media associated with links 946 and 948 may also provide a response to the user's query. Here, the media associated with links 946 and 948 may include a ten-minute clip of weather information for the specified location—specifically, a five-day forecast for New York from a television channel called the Weather Forecast Channel.

[0121] In one embodiment, the clip resolving a user query may include a time-indicated portion of previously broadcast content (available from recordings or streaming services). In one embodiment, the virtual assistant may identify such content based on the user's intent associated with voice input and by searching for detailed information about available media content (e.g., including metadata for recorded programs, along with detailed timing information or details about streaming content). In some embodiments, the user may not be able to access or subscribe to specific content. In such cases, content may be offered for purchase, for example, via purchase link 948. The cost of the content may be automatically deducted from or charged to the user's account when purchase link 948 or video link 946 is selected.

[0122] Figure 10 An exemplary process 1000 for using a virtual assistant to control television interaction and displaying related information using different interfaces is illustrated. In block 1002, voice input can be received from a user. For example, voice input can be received at a user device 102 or remote control 106 of system 100. In some embodiments, voice input (or some or all of a data representation of the voice input) can be transmitted to and received by a server system 110 and / or a set-top box 104. In response to a user initiating the reception of voice input, various notifications can be displayed on a display (e.g., display 112). For example, they can be referenced above. Figures 4A-4EThe system displays confirmations of readiness, listening, processing, and / or command reception. Furthermore, it can transcribe received user voice input and display the transcription.

[0123] Refer again Figure 10 In process 1000, at box 1004, media content can be determined based on voice input. For example, media content that addresses a user's query and is directed to a virtual assistant can be determined (e.g., by searching for available media content). For example, it can be determined that... Figure 4E Transcription of 492 related media content (“What sporting events are happening now?”). Such media content may include live sporting events being broadcast on one or more television channels that a user can watch.

[0124] In box 1006, a first user interface with selectable media links can be displayed at a first size. For example, it can be as follows: Figure 5 As shown, a media content interface 510 with selectable video links 512 and selectable text links 514 is displayed on the monitor 112. As described above, the media content interface 510 can have a smaller size to avoid interfering with the background video content.

[0125] In box 1008, selection of one of the links can be received. For example, selection of one of links 512 and / or 514 can be received. In box 1010, a larger second user interface with media content associated with the selection can be displayed. For example, it can be as follows: Figure 6B As shown, a details interface 618 is displayed, featuring a selectable video link 620 and detailed media information 622. As mentioned above, the details interface 618 can be larger to convey desired additional detailed media information. Similarly, when selecting video link 620, it can be as follows... Figure 7B The diagram shows an extended media interface 728 with video 726. As mentioned above, the extended media interface 728 can still be larger in size to provide the user with the desired media content. In this way, the sizes of the various interfaces discussed herein can be set to accommodate the desired content (including expanding to a larger size or shrinking to a smaller size) while otherwise occupying a limited screen area. Therefore, process 1000 can be used to use a virtual assistant to control television interaction and utilize different interfaces to display relevant information.

[0126] In another embodiment, a larger interface can be displayed above the control menu instead of above the background video content. For example, it can be as follows: Figure 8B As shown, a voice input interface 836 is displayed above menu 830, which can be used as follows: Figure 9 As shown, the assistant results interface 942 is displayed above menu 830, and can be viewed as follows: Figure 5The image shows a smaller media content interface 510 displayed above video 480. In this way, the size of the interface (e.g., the amount of screen space occupied by the interface) can be determined at least in part by the type of background content.

[0127] Figure 11 Exemplary television media content on user equipment 102 is shown. The user equipment may include a mobile phone, tablet computer, remote control, etc., with a touch screen 246 (or another display). Figure 11 An interface 1150 is shown, comprising a list of television programs 1152. Interface 1150 may correspond, for example, to a specific application on user device 102, such as a television control application, a television content list application, an internet application, etc. In some embodiments, a user intent can be determined from voice input associated with content displayed on user device 102 (e.g., on touchscreen 246), and the user intent can be used to cause content to be played back or displayed on another device and display (e.g., on set-top box 104 and display 112 and / or speaker 111). For example, the content displayed in interface 1150 on user device 102 can be used to dispel ambiguity in a user request and determine the user intent from voice input, and then the determined user intent can be used to play or display media via set-top box 104.

[0128] Figure 12 An example of television control using a virtual assistant is shown. Figure 12 An interface 1254 is shown, which may include a virtual assistant interface formatted as a conversational dialogue between an assistant and a user. For example, interface 1254 may include an assistant greeting 1256 prompting the user to make a request. Subsequently received user voice can then be transcribed, such as transcribed user voice 1258, displaying the conversation back and forth. In some embodiments, interface 1254 may appear on user device 102 in response to a trigger initiating the reception of voice input (such as a button press, key phrase, etc.).

[0129] In one embodiment, a user request to play content via a set-top box 104 (e.g., on display 112 and speaker 111) may include an ambiguous reference to content displayed on user device 102. For example, transcribed user speech 1258 may include a reference to “that” football match (“Show that football match.”). The specific football match desired may not be clear from the voice input alone. However, in some embodiments, the content displayed on user device 102 can be used to dispel ambiguity in the user request and determine user intent. In one embodiment, content displayed on user device 102 before the user makes a request (e.g., before interface 1254 appears on touchscreen 246) can be used to determine user intent (content may appear within interface 1254, such as previous queries and results). In the illustrated embodiment, content can be used... Figure 11 The content displayed in interface 1150 determines the user's intent from the command to display "that" football match. The television program list 1152 includes various programs, one of which is titled "Football" appearing on channel 5. The appearance of the football list can be used to determine the user's intent from mentioning "that" football match. Specifically, the user mentioning "that" football match can be interpreted as a football program appearing in the television program list of interface 1150. Therefore, the virtual assistant can enable playback of the specific football match the user desires (e.g., by causing the set-top box 104 to tune to the appropriate channel and display the match).

[0130] In other embodiments, users can invoke the television programs shown in interface 1150 (e.g., programs on Channel 8, news, dramas, advertisements, premieres, etc.) in various other ways, and user intent can be similarly determined based on the displayed content. It should be understood that user intent can also be determined further by combining the displayed content with metadata associated with the displayed content (e.g., television program descriptions), fuzzy matching techniques, synonym matching, etc. For example, the term "advertisement" can be matched to the description "pay-per-view" (e.g., using synonym and / or fuzzy matching techniques) to determine user intent from a request to display "advertisement." Similarly, the description of a particular television program can be analyzed when determining user intent. For example, the term "law" can be identified in the detailed description of a courtroom drama, and user intent can be determined from a user request to watch a "law" program based on the detailed description associated with the content displayed in interface 1150. Therefore, the displayed content and the data associated with it can be used to dispel ambiguity in user requests and determine user intent.

[0131] Figure 13 Exemplary pictures and video content on user equipment 102 are shown. The user equipment may include a mobile phone, tablet computer, remote control, etc., with a touch screen 246 (or another display). Figure 13An interface 1360 is shown, including a list of photos and videos. Interface 1360 may correspond to a specific application on user device 102, such as a media content application, a file navigation application, a storage application, a remote storage management application, a camera application, etc. As shown, interface 1360 may include videos 1362, a photo album 1364 (e.g., a group of multiple pictures), and photos 1366. (See above reference...) Figure 11 and Figure 12 The user intent can be determined from voice input associated with the content displayed on user equipment 102. The user intent can then be used to cause the content to be played back or displayed on another device and display (e.g., on set-top box 104 and display 112 and / or speaker 111). For example, the content displayed in interface 1360 on user equipment 102 can be used to dispel ambiguity in the user request and determine the user intent from voice input, and then the determined user intent can be used to play or display media via set-top box 104.

[0132] Figure 14 An exemplary media display control using a virtual assistant is shown. Figure 14 An interface 1254 is shown, which may include a virtual assistant interface formatted as a conversational dialogue between an assistant and a user. As shown, interface 1254 may include an assistant greeting 1256 prompting the user to make a request. Within the dialogue, then... Figure 14 The example illustrates the transcription of user speech. In some embodiments, interface 1254 may appear on user device 102 in response to a trigger initiating the reception of voice input (such as a button press, key phrase, etc.).

[0133] In one embodiment, a user request to play or display media content via a set-top box 104 (e.g., on display 112 and speaker 111) may include an ambiguous reference to content displayed on user device 102. For example, transcribed user voice 1468 may include a reference to “that” video (“Show that video.”). The specific video referenced may not be clear from the perspective of voice input alone. However, in some embodiments, the content displayed on user device 102 can be used to dispel ambiguity in the user request and determine user intent. In one embodiment, content displayed on user device 120 before the user makes a request (e.g., before interface 1254 appears on touchscreen 246) can be used to determine user intent (content may appear within interface 1254, such as previous queries and results). In the embodiment of user voice 1468, the following can be used: Figure 13The content displayed in interface 1360 determines the user's intent to display "that" video from the command. The list of photos and videos in interface 1360 includes various different photos and videos, including video 1362, album 1354, and photo 1366. Since only one video appears in interface 1360 (e.g., video 1362), the user's intent to cite "that" video can be determined by the presence of video 1362 in interface 1360. Specifically, the user's citation of "that" video can be interpreted as video 1362 (titled "Graduation Video") appearing in interface 1360. Therefore, the virtual assistant can cause video 1362 to play back (e.g., by causing video 1362 to be transmitted from user device 102 or a remote storage device to set-top box 104 and causing playback to begin).

[0134] In another embodiment, the transcribed user voice 1470 includes referencing "that" album ("Play the slideshow of that album."). The specific album being referenced may not be clear from the voice input alone. The content displayed on user device 102 can again be used to dispel ambiguity in the user's request. Specifically, it can use... Figure 13 The content displayed in interface 1360 determines the user's intent from the command to play a slideshow of "that" album. The list of photos and videos in interface 1360 includes album 1354. The appearance of album 1364 in interface 1360 can be used to determine the user's intent from speaking "that" album. Specifically, the user's reference to "that" album can be resolved to album 1364 (titled "Graduation Album") appearing in interface 1360. Therefore, in response to the user's voice 1470, the virtual assistant can cause the slideshow to be displayed, including photos from album 1364 (e.g., by causing the photos of album 1364 to be transferred from user device 102 or a remote storage device to set-top box 104 and causing the photo slideshow to start).

[0135] In another embodiment, the transcribed user voice 1472 includes a reference to a "last" photo ("Show the last photo on the kitchen TV."). The specific photo referenced may not be clear from the voice input alone. The content displayed on user device 102 can again be used to dispel ambiguity in the user's request. Specifically, it can be used... Figure 13The content displayed in interface 1360 determines the user's intent to display the "last" photo from the command. The list of photos and videos in interface 1360 includes two separate photos 1366. The appearance of photo 1366 in interface 1360, and especially the order in which photo 1366 appears within the interface, can be used to determine the user's intent from mentioning the "last" photo. Specifically, the user's reference to the "last" photo can be interpreted as photo 1366 (date: June 21, 2014) appearing at the bottom of interface 1360. Therefore, in response to the user's voice 1472, the virtual assistant can cause the display of the last photo 1366 shown in interface 1360 (e.g., by causing the last photo 1366 to be transferred from user device 102 or a remote storage device to set-top box 104 and causing the photo to be displayed).

[0136] In other embodiments, users can invoke the media content shown in interface 1360 through various other means (e.g., last pair of photos, all videos, all photos, graduation album, graduation video, photos from June 21st, etc.), and similarly determine user intent based on the displayed content. It should be understood that user intent can also be determined further by combining the displayed content with associated metadata (e.g., timestamps, location information, titles, descriptions, etc.), fuzzy matching techniques, synonym matching, etc. Therefore, the displayed content and its associated data can be used to dispel ambiguity in user requests and determine user intent.

[0137] It should be understood that any type of displayed content from any application interface of any application can be used to determine user intent. For example, an image displayed on a webpage in an internet browser application can be referenced in voice input, and the displayed webpage content can be analyzed to identify the desired image. Similarly, a music track from a music playlist in a music application can be referenced in voice input based on title, genre, artist, band name, etc., and the user intent can be determined from the voice input using content displayed in the music application (and associated metadata in some embodiments). As described above, the determined user intent can then be used to display or play media via another device, such as a television set-top box 104.

[0138] In some embodiments, user identification, user authentication, and / or device authentication may be used to determine whether media control can be authorized, whether media content can be displayed, and access permission, etc. For example, it may be determined whether a specific user device (e.g., user device 102) is authorized to control media, such as on a set-top box 104. Authorization of the user device may be based on registration, pairing, trust determination, password, security questions, system settings, etc. In response to determining that a specific user device has been authorized, permission may be granted to attempt to control the set-top box 104 (e.g., authorization may be granted in response to determining that the device requests control of the media to play media content). Conversely, media control commands or requests from unauthorized devices may be ignored and / or users of such devices may be prompted to register their devices for use when controlling a specific set-top box 104.

[0139] In another embodiment, a specific user can be identified, and personal data associated with that user can be used to determine the user's intent in making the request. For example, the user can be identified based on voice input, such as through speech recognition utilizing the user's voiceprint. In some embodiments, the user can speak a specific phrase, which is analyzed for speech recognition. In other embodiments, speech recognition can be used to analyze voice input requests directed to a virtual assistant to identify the speaker. The user can also be identified based on the source of a voice input sample (e.g., on the user's personal device 102). The user can also be identified based on passwords, PINs, menu selections, etc. The voice input received from the user can then be interpreted based on the identified user's personal data. For example, the user's intent in the voice input can be determined based on previous requests from the user, media content owned by the user, media content stored on the user's device, user preferences, user settings, user demographics (e.g., languages ​​spoken, etc.), user profile information, user payment methods, or various other personal information associated with the specifically identified user. For example, ambiguity in voice input referencing a favorites list can be eliminated based on personal data, and the user's personal favorites list can be identified. Similarly, user-identified voice inputs referencing "my" photos, "my" videos, "my" performances, etc., can be eliminated to correctly identify photos, videos, and performances associated with the identified user (e.g., photos stored on the user's personal device). Likewise, voice input requesting to purchase content can be eliminated to determine if the identified user's payment method should be used to pay for the purchase (as opposed to another user's payment method).

[0140] In some embodiments, user authentication can be used to determine whether a user is allowed access to media content, to purchase media content, etc. For example, voice recognition can be used to verify the identity of a specific user (e.g., using their voiceprint) to authorize the user to make purchases using the user's payment method. Similarly, passwords or other methods can be used to authenticate users to authorize shopping. In another embodiment, voice recognition can be used to verify the identity of a specific user to determine whether the user is authorized to watch specific programs (e.g., programs with specific parental guidance ratings, movies with specific age-appropriate ratings, etc.). For example, a child's request for a specific program can be denied based on voice recognition indicating that the requester is not an authorized user (e.g., a parent) capable of watching such content. In other embodiments, voice recognition can be used to determine whether a user is capable of accessing specific subscription content (e.g., restricting access to premium channel content based on voice recognition). In some embodiments, a user can speak a specific phrase, which is analyzed for voice recognition. In other embodiments, voice recognition can be used to analyze voice input requests directed to a virtual assistant to identify the speaker. Specific media content can then be played in response to the initial identification of an authorized user through various means.

[0141] Figure 15 An exemplary virtual assistant interaction with results is illustrated on a mobile user device and a media display device. In some embodiments, the virtual assistant can provide information and control about more than one device, such as user device 102 and set-top box 104. Furthermore, in some embodiments, the same virtual assistant interface used for control and information on user device 102 can be used to issue requests for control of media on set-top box 104. In this way, the virtual assistant system can determine whether to display results or perform tasks on user device 102 or set-top box 104. In some embodiments, when controlling set-top box 104 using user device 102, the intrusion of the virtual assistant interface on the display (e.g., display 112) associated with set-top box 104 can be minimized by displaying information on user device 102 (e.g., on touchscreen 246). In other embodiments, virtual assistant information can be displayed solely on display 112, or virtual assistant information can be displayed on both user device 102 and display 112.

[0142] In some embodiments, it can be determined whether the virtual assistant query results should be displayed directly on user device 102 or on display 112 associated with set-top box 104. In one embodiment, in response to determining that the user intent of the query includes a request for information, an information response can be displayed on user device 102. In another embodiment, in response to determining that the user intent of the query includes a request for playing media content, media content in response to the query can be played via set-top box 104.

[0143] Figure 15 A virtual assistant interface 1254 is shown, illustrating an example of a conversational dialogue between the virtual assistant and the user. An assistant greeting 1256 may prompt the user to make a request. In a first query, transcribed user voice 1574 (which may also be typed or entered otherwise) includes a request for an informational answer related to the displayed media content. Specifically, transcribed user voice 1574 queries who is participating, possibly on an interface on user device 102 (e.g., ...). Figure 11 (as listed in interface 1150) or on display 112 (e.g., on Figure 5 Listed in interface 510 or as video 726 Figure 7B A football match is displayed on the monitor 112. The user intent of the transcribed user speech 1574 can be determined based on the displayed media content. For example, the specific football match being discussed can be identified based on the content displayed on the user device 102 or the monitor 112. The user intent of the transcribed user speech 1574 may include obtaining information about the teams involved in the identified football match based on the displayed content. In response to determining that the user intent includes a request for information, the system can determine the... Figure 15 The response is displayed within interface 1254 (opposite to display 112). In some embodiments, the response to a query may be determined based on metadata associated with the displayed content (e.g., based on a description of a football match in a list of televisions). As shown, assistant response 1576 can therefore be displayed on touchscreen 246 of user device 102 in interface 1254, indicating that Alpha and Zeta teams are playing a match. Thus, in some embodiments, an information response may be displayed within interface 1254 on user device 102 based on the determination that the query includes an information request.

[0144] However, the second query in interface 1254 includes a media request. Specifically, the transcribed user voice 1578 requests that the displayed media content be changed to "match". The user intent of the transcribed user voice 1578 can be determined based on the displayed content (e.g., identifying which match the user wants to watch), for example, Figure 5 The matches listed in interface 510, Figure 11The matches listed in interface 1150, matches referenced in previous queries (e.g., in transcribed user voice 1574), etc. The user intent of transcribed user voice 1578 can therefore include changing the displayed content to a specific match—here, a football match between Alpha and Zeta. In one embodiment, the match can be displayed on user device 102. However, in other embodiments, based on a query including a request to play media content, the match can be displayed via set-top box 104. Specifically, in response to determining that the user intent includes a request to play media content, the system can determine to display media content on display 112 via set-top box 104 (as per...). Figure 15 (The opposite is true within interface 1254). In some embodiments, a response or paraphrase confirming the desired action of the virtual assistant (e.g., "Change to football match.") may be displayed in interface 1254 or on display 112.

[0145] Figure 16 An exemplary virtual assistant interaction with media results displayed on a media display device and a mobile user device is illustrated. In some embodiments, the virtual assistant can provide access to media on user device 102 and set-top box 104. Furthermore, in some embodiments, the same virtual assistant interface used for media on user device 102 can be used to make a request for media on set-top box 104. In this way, the virtual assistant system can determine whether to display the media results on user device 102 or on display 112 via set-top box 104.

[0146] In some embodiments, whether to display media on device 102 or display 112 can be determined based on media result format, user preferences, default settings, explicit commands in the request itself, etc. For example, the format of the queried media results can be used to determine which device the media results should be displayed on by default (e.g., without specific instructions). Television programs may be more suitable for display on a television set, larger format videos may be more suitable for display on a television set, thumbnail photos may be more suitable for display on a user device, small format web videos may be more suitable for display on a user device, and various other media formats may be more suitable for display on a larger television screen or a smaller user device display. Therefore, in response to (e.g., based on media format) determining that media content should be displayed on a particular display, the media content may be displayed on that particular display by default.

[0147] Figure 16A virtual assistant interface 1254 is shown, with an example of a query related to playing or displaying media content. An assistant greeting 1256 may prompt the user to make a request. In a first query, the transcribed user voice 1680 includes a request to display a football match. As in the embodiments described above, the user intent of the transcribed user voice 1680 can be determined based on the displayed content (e.g., identifying which match the user wants to watch), for example, Figure 5 The matches listed in interface 510, Figure 11 The matches listed in interface 1150, matches referenced in previous queries, etc. The transcribed user intent of user voice 1680 can therefore include displaying a specific football match, for example, broadcast on a television set. In response to determining that the user intent includes a request to display media formatted for a television set (e.g., a football match broadcast on a television set), the system can automatically determine the desired media to be displayed on display 112 via set-top box 104 (as opposed to user device 102 itself). The virtual assistant system can then cause set-top box 104 to tune to the football match and display it on display 112 (e.g., by performing necessary tasks and / or sending appropriate commands).

[0148] However, in the second query, the transcribed user voice 1682 includes a request to display images of team players (e.g., "images of the Alpha team"). As in the embodiments described above, the user intent of the transcribed user voice 1682 can be determined. The user intent of the transcribed user voice 1682 may include searching for images associated with "Alpha team" (e.g., a web search) and displaying the resulting images. In response to determining that the user intent includes a request to display media that can be presented in thumbnail format, media associated with a web search, or other non-specific media without a specific format, the system can automatically determine to display the desired media results on the touchscreen 246 in the interface 1254 of the user device 102 (as opposed to displaying the resulting images on the display 112 via the set-top box 104). For example, as shown, thumbnail photos 1684 may be displayed within the interface 1254 on the user device 102 in response to a user query. Thus, the virtual assistant system can make it possible to display media in a specific format or media that can be presented in a specific format (e.g., in a set of thumbnails) on the user device 102 by default.

[0149] It should be understood that in some embodiments, a football match referenced in user voice 1680 may be displayed on user device 102, and a photo 1684 may be displayed on display 112 via set-top box 104. However, the default display device can be automatically determined based on the media format, thereby simplifying media commands for the user. In other embodiments, the default device for displaying requested media content may be determined based on user preferences, default settings, the device most recently used to display content, voice recognition identifying the user and the device associated with the user, etc. For example, the user may set preferences or default configurations to display specific types of content (e.g., videos, slideshows, television programs, etc.) on display 112 via set-top box 104 and other types of content (e.g., thumbnails, photos, web videos, etc.) on touchscreen 246 of user device 102. Similarly, preferences or default configurations may be set to respond to a specific query by displaying content on one device or another. In another embodiment, all content may be displayed on user device 102 unless the user instructs otherwise.

[0150] In other embodiments, a user query may include a command to display content on a specific display. For example, Figure 14 User voice 1472 includes a command to display a photo on a kitchen television. Therefore, the system can cause a photo to be displayed on a television monitor associated with the user's kitchen, as opposed to displaying a photo on user device 102. In other embodiments, the user can indicate which display device to use in various other ways (e.g., on a television, on a large screen, in the living room, in the bedroom, on my tablet, on my phone, etc.). Therefore, the display device used to display the results of a virtual assistant query can be determined in various different ways.

[0151] Figure 17Exemplary media device control based on proximity is illustrated. In some embodiments, a user may have multiple televisions and set-top boxes within the same household or on the same network. For example, a household may have a television and set-top box in the living room, another set in the bedroom, and yet another set in the kitchen. In other embodiments, multiple set-top boxes may be connected to the same network, such as a public network in an apartment or office building. While a user can pair, connect, or otherwise authorize a remote control 106 and user device 102 for a particular set-top box to prevent unauthorized access, in other embodiments, a remote control and / or user device may be used to control more than one set-top box. A user may, for example, use a single user device 102 to control set-top boxes in the bedroom, living room, and kitchen. A user may also, for example, use a single user device 102 to control their own set-top box in their own apartment, as well as a neighbor's set-top box in a neighbor's apartment (e.g., sharing content from user device 102 with a neighbor, such as displaying a slideshow of photos stored on user device 102 on a neighbor's television). Because a user can control multiple different set-top boxes using a single user device 102, the system can determine which set-top box to send a command to. Similarly, because a home can have multiple remote controls 106 capable of operating multiple set-top boxes, the system can similarly determine which set-top box to send a command to.

[0152] In one embodiment, proximity of the devices can be used to determine which of the multiple set-top boxes to send a command to (or on which display the requested media content is displayed). Proximity can be determined between the user equipment 102 or remote controller 106 and each of the multiple set-top boxes. The issued command can then be sent to the nearest set-top box (or the requested media content can be displayed on the nearest display). Proximity can be determined by any of a variety of methods, such as time-of-flight measurement (e.g., using radio frequency), Bluetooth LE, electronic verification signals, proximity sensors, sound travel measurement, etc. The measured or approximate distances can then be compared, and a command can be issued to the nearest device (e.g., the nearest set-top box).

[0153] Figure 17A multi-device system 1790 is illustrated, including a first set-top box 1792 with a first display 1786 and a second set-top box 1794 with a second display 1788. In one embodiment, a user can issue a command from a user device 102 to display media content (e.g., without specifying where or on which device). A distance 1795 to the first set-top box 1792 and a distance 1796 to the second set-top box 1794 can then be determined (or approximately determined). As shown, distance 1796 can be greater than distance 1795. Based on proximity, a command from the user device 102 can be issued to the first set-top box 1792, which is the nearest device and most likely to match the user's intent. In some embodiments, a single remote control 106 can also be used to control more than one set-top box. The desired device to be controlled at a given time can be determined based on proximity. A distance 1797 to the second set-top box 1794 and a distance 1798 to the first set-top box 1792 can then be determined (or approximately determined). As shown, distance 1798 can be greater than distance 1797. Based on proximity, commands from the remote control 106 can be sent to the second set-top box 1794, which is the closest device and most likely to match the user's intent. The distance measurement results can be refreshed periodically or in conjunction with each command, for example, to adapt to the user entering different rooms and wishing to control different devices.

[0154] It should be understood that users can specify different devices for commands, in some cases, bypassing proximity. For example, a list of available display devices can be displayed on user device 102 (e.g., listing the first display 1786 and the second display 1788 by setting a name, specifying a room, etc., or listing the first set-top box 1792 and the second set-top box 1794 by setting a name, specifying a room, etc.). The user can select one of the devices from the list and then send a command to the selected device. The request for media content issued at user device 102 can then be processed by displaying the desired media on the selected device. In other embodiments, the user can speak the desired device as part of a spoken command (e.g., displaying a game on the kitchen television, changing to a cartoon channel in the living room, etc.).

[0155] In other embodiments, a default device for displaying the requested media content can be determined based on state information associated with a particular device. For example, it can be determined whether headphones (or headsets) are attached to user device 102. In response to determining that headphones are attached to user device 102 when a request to display media content is received, the requested content can be displayed on user device 102 by default (e.g., assuming the user is consuming content on user device 102 rather than a television). In response to determining that headphones are not attached to user device 102 when a request to display media content is received, the requested content can be displayed on either user device 102 or a television according to any of the various determination methods discussed herein. Other device status information can be similarly used to determine whether the requested content should be displayed on user equipment 102 or set-top box 104, such as ambient lighting around user equipment 102 or set-top box 104, proximity of other devices to user equipment 102 or set-top box 104, orientation of user equipment 102 (e.g., landscape orientation may more likely indicate the desired view on user equipment 102), display status of set-top box 104 (e.g., in sleep mode), time elapsed since the last interaction on a particular device, or any of a variety of other status indicators for user equipment 102 and / or set-top box 104.

[0156] Figure 18 An exemplary process 1800 for controlling television interaction using a virtual assistant and multiple user devices is illustrated. At block 1802, voice input can be received from a user at a first device having a first display. For example, voice input can be received from a user at user device 102 or remote control 106 of system 100. The first display may include a touchscreen 246 of user device 102 or a display associated with remote control 106 in some embodiments.

[0157] In box 1804, the user's intent can be determined from the voice input based on the content displayed on the first display. For example, analysis such as... Figure 11 TV programs in interface 1150 or Figure 13 The content of photos and videos in the interface 1360 is used to determine the user's intent regarding the voice input. In some embodiments, the user may refer to the content displayed on the first display in an ambiguous manner, and the ambiguity of the reference can be eliminated by analyzing the content displayed on the first display to resolve the reference (e.g., determining the user's intent regarding "that" video, "that" album, "that" game, etc.), as referenced above. Figure 12 and Figure 14 As stated above.

[0158] Refer again Figure 18In process 1800, at box 1806, media content can be determined based on user intent. For example, specific videos, photos, albums, television programs, sports events, music tracks, etc., can be identified based on user intent. (As described above...) Figure 11 and Figure 12 In examples, for instance, it can be based on citations Figure 11 The user intent identifier for "that" football match shown in interface 1150 is the specific football match displayed on channel five. (As described above...) Figure 13 and Figure 14 In the example, it can be based on from Figure 14 The voice input example determines the user intent and identifies a specific video 1362 titled "Graduation Video", a specific album 1364 titled "Graduation Album", or a specific photo 1366.

[0159] Refer again Figure 18 In process 1800, at block 1808, media content can be played on a second device associated with the second display. For example, the determined media content can be played on a display 112 with speakers 111 via a set-top box 104. Playing media content may include tuning to a specific television channel on the set-top box 104 or another device, playing a specific video, displaying a slideshow of photos, displaying a specific photograph, playing a specific audio track, etc.

[0160] In some embodiments, it can be determined whether a response to voice input directed at a virtual assistant should be displayed on a first display associated with a first device (e.g., user equipment 102) or a second display associated with a second device (e.g., a set-top box 104). For example, as referenced above... Figure 15 and Figure 16 The aforementioned means that informational answers or media content suitable for display on a smaller screen can be displayed on user equipment 102, while media responses or media content suitable for display on a larger screen can be displayed on a display associated with set-top box 104. (See above reference) Figure 17 In some embodiments, the distance between user equipment 102 and multiple set-top boxes can be used to determine which set-top box to play media content on or to which set-top box to issue a command. Various other determinations can be made similarly to provide a convenient and user-friendly experience when multiple devices can interact.

[0161] In some embodiments, since the interpretation of voice input can be communicated using the content displayed on user equipment 102 as described above, the interpretation of voice input can also be communicated using the content displayed on display 112. Specifically, the content displayed on the display associated with the set-top box 104, along with metadata associated with that content, can be used to determine user intent from voice input, dispel ambiguity in user queries, and respond to content-related queries, etc.

[0162] Figure 19 An exemplary voice input interface 484 (as described above) is shown, with a virtual assistant query about video 480 displayed in the background. In some embodiments, a user query may include a question about media content displayed on display 112. For example, transcription 1916 includes a query requesting identification of actresses (“Who are those actresses?”). The content displayed on display 112—along with metadata or other descriptive information about that content—can be used to determine the user’s intent from the voice input associated with that content and to determine a response to the query (responses include informative responses and media responses that offer the user media options). For example, video 480, a description of video 480, a list of characters and actors in video 480, rating information for video 480, genre information for video 480, and various other descriptive information associated with video 480 can be used to disambiguate the user’s request and determine a response to the user’s query. Associated metadata may include, for example, identification information for characters 1910, 1912, and 1914 (e.g., character names along with the names of the actresses who play the characters). Metadata for any other content may similarly include titles, descriptions, lists of characters, lists of actors, lists of crew members, genres, producer names, director names, or display schedules associated with the content displayed on the monitor or the viewing history of media content on the monitor (e.g., recently displayed media).

[0163] In one embodiment, a user query directed to a virtual assistant may include an ambiguous reference to something displayed on display 112. For example, transcription 1916 includes a reference to “those” actresses (“Who are those actresses?”). From the voice input alone, the specific actress the user is asking about may not be clear. However, in some embodiments, the content displayed on display 112 and associated metadata can be used to dispel the ambiguity of the user request and determine the user's intent. In the illustrated embodiment, the user's intent can be determined from the reference to “those” actresses using the content displayed on display 112. In one embodiment, set-top box 104 may identify the content being played along with details associated with that content. In this case, set-top box 104 may identify the title of video 480 along with various descriptive content. In other embodiments, television performances, sporting events, or other content may be shown, which can be used in conjunction with associated metadata to determine the user's intent. Furthermore, in any of the various embodiments discussed herein, the voice recognition results and intent determination may be weighted more heavily on items associated with the displayed content than alternative methods. For example, the actor's name of a character on screen can be given higher weight when those actors are on screen (or when a performance featuring them is playing), which can provide accurate speech recognition and intent determination of possible user requests associated with the displayed content.

[0164] In one embodiment, all or most of the main actresses appearing in video 480 can be identified using a list of characters and / or actors associated with video 480, which may include actresses 1910, 1912, and 1914. The identified actresses can be returned as possible results (including fewer or additional actresses if the metadata resolution is coarse). However, in another embodiment, the metadata associated with video 480 may include identifiers of which male and female actors appeared on screen at a given time, and the actresses appearing at the time of the query can be determined from the metadata (e.g., specifically identifying actresses 1910, 1912, and 1914). In another embodiment, actresses 1910, 1912, and 1914 can be identified from images displayed on display 112 using a facial recognition application. In other embodiments, various other metadata associated with video 480 and various other identification methods can be used to identify the user's possible intent when inferring "those" actresses.

[0165] In some embodiments, the content displayed on display 112 may change during query submission and response determination. This allows the viewing history of the media content to be used to determine user intent and response to the query. For example, if video 480 moves to another view (e.g., with different characters) before a response to the query is generated, the query result can be determined based on the user's view at the time the query is spoken (e.g., the character displayed on the screen when the user initiated the query). In some cases, the user may pause media playback to issue a query; the content displayed during the pause can be used in conjunction with associated metadata to determine user intent and response to the query.

[0166] Given a defined user intent, query results can be provided to the user. Figure 20 An exemplary assistant response interface 2018 is shown, including an assistant response 2020, which may include responses from... Figure 19 The response determined by the query transcribed in 1916. As shown, the assistant response 2020 may include the name of each actress and a list of their associated roles in video 480 ("Actress Jennifer Jones plays the role of Blanche; Actress Elizabeth Arnold plays the role of Julia; Actress Whitney Davidson plays the role of Melissa."). The actresses and roles listed in response 2020 may correspond to roles 1910, 1912, and 1914 appearing on display 112. As mentioned above, in some embodiments, the content displayed on display 112 may change during the submission of the query and the determination of the response. In this way, response 2020 may include information about content or roles that may no longer appear on display 112.

[0167] Like other interfaces displayed on display 112, the assistant response interface 2018 can occupy a minimal amount of screen space while providing ample space to convey the desired information. In some embodiments, like other text displayed on the interface on display 112, the assistant response 2020 can scroll up from the bottom of display 112 to... Figure 20 The position shown displays a specific amount of time (e.g., a delay based on response length) and scrolls up out of view. In other embodiments, interface 2018 can slide down out of view after a delay.

[0168] Figure 21 and Figure 22 Another example is shown where user intent is determined and a query is responded to based on what is displayed on display 112. Figure 21An exemplary voice input interface 484 is shown, featuring a virtual assistant for querying media content associated with video 480. In some embodiments, a user query may include a request for media content associated with media displayed on display 112. For example, a user may request other movies, television programs, sporting events, etc., associated with a particular media based on, for example, a character, actor, genre, etc. For example, transcription 2122 includes a query requesting other media associated with an actress in video 480, referenced by her character name in video 480 (“What other shows has Blanche appeared in?”). The content displayed on display 112—along with metadata or other descriptive information about the content—can again be used to determine the user’s intent and the response to the query (informative or leading to media selection) from the voice input associated with that content.

[0169] In some embodiments, a user query directed to a virtual assistant may include ambiguous references using character names, actor names, program names, team member names, etc. Without the context of the content displayed on display 112 and its associated metadata, such references may be difficult to resolve precisely. Transcription 2122 may include, for example, a reference to a character named “Blanche” from video 480. From the voice input alone, it may not be clear which actress or other individual the user is asking about. However, in some embodiments, the content displayed on display 112 and the associated metadata can be used to dispel ambiguity in the user request and determine the user's intent. In an exemplary embodiment, the user intent can be determined from the character name “Blanche” using the content displayed on display 112 and the associated metadata. In this case, the list of characters associated with video 480 can be used to determine that “Blanche” likely refers to the character “Blanche” in video 480. In another embodiment, detailed metadata and / or facial recognition can be used to determine that the character named “Blanche” appears on the screen (or appears on the screen when the user query is initiated), making the actress associated with that character the most likely intent of the user query. For example, it can be determined that characters 1910, 1912, and 1914 appear on display 112 (or appear on display 112 when a user query is initiated), and then their associated characters can be invoked to determine the user intent of the query invoking character Blanche. The actress who plays Blanche can then be identified using the list of actors, and a search can be performed to identify other media featuring the identified actress.

[0170] Given the determined user intent (e.g., the resolution of the character referencing "Blanche") and the determination of the query results (e.g., other media associated with the actress who plays "Blanche"), a response can be provided to the user. Figure 22An exemplary assistant response interface 2224 is shown, including an assistant text response 2226 and an optional video link 2228, which can be in response to... Figure 21 The response is given in response to the query transcribed from 2122. As shown, the assistant text response 2226 may include a paraphrase of the user's request introducing the optional video link 2228. The assistant text response 2226 may also include instructions to dispel ambiguity in the user's query—specifically, identifying actress Jennifer Jones as the character Blanche in video 480. Such paraphrasing can confirm to the user that the virtual assistant has correctly interpreted the user's query and provided the expected results.

[0171] The assistant response interface 2224 may also include selectable video links 2228. In some embodiments, various media content may be provided as virtual assistant query results, including movies (e.g., Movie A and Movie B of interface 2224). The media content displayed as query results may include media available for user consumption (free, purchased, or as part of a subscription). The user can select the displayed media to watch or consume the resulting content. For example, the user can select one of the selectable video links 2228 (e.g., using a remote control, voice command, etc.) to watch one of other movies featuring actress Jennifer Jones. In response to the selection of one of the selectable video links 2228, the video associated with that selection can be played, replacing video 480 on display 112. Thus, the user's intent can be determined from voice input using the displayed media content and associated metadata, and in some embodiments, playable media may be provided as results.

[0172] It should be understood that users may, when formulating a query, cite actors, players, roles, locations, teams, sporting event details, movie themes, or various other information associated with the displayed content. The virtual assistant system can similarly disambiguate such requests and determine the user's intent based on the displayed content and associated metadata. Similarly, it should be understood that in some embodiments, the results may include media recommendations associated with the query, such as movies, television shows, or sporting events associated with the person as the subject of the query (regardless of whether the user specifically requested such media content).

[0173] Furthermore, in some embodiments, user queries may include requests for information associated with the media content itself, such as queries about characters, episodes, movie plots, previous scenes, etc. As in the embodiments described above, the user intent and response can be determined from such queries using the displayed content and associated metadata. For example, a user might request a description of a character (e.g., “What did Blanche do in this movie?”). The virtual assistant system can then identify the requested information about the character from the metadata associated with the displayed content, such as a character description or person (e.g., “Blanche is one of a group of lawyers considered the troublemaker for Hartford.”). Similarly, a user might request an episode summary (e.g., “What happened in the last episode?”), and the virtual assistant system can search for and provide a description of the episode.

[0174] In some embodiments, the content displayed on the display 112 may include menu content, which may be similarly used to determine the user's intent regarding the voice input and the response to the user's query. Figures 23A-23B An exemplary page of program menu 830 is shown. Figure 23A The first page shows media option 832. Figure 23B The second page media option 832 is shown (which may include a successive next page in the content list that extends beyond a single page).

[0175] In one embodiment, a user request to play content may include an ambiguous reference in menu 830 to content displayed on display 112. For example, a user viewing menu 830 may request to watch "that" football game, "that" basketball game, a vacuum cleaner commercial, a legal program, etc. The specific program desired may not be clear from voice input alone. However, in some embodiments, the content displayed on display 112 can be used to dispel ambiguity in the user request and determine the user's intent. In an illustrated embodiment, media options in menu 830 (along with metadata associated with the media options in some embodiments) can be used to determine the user's intent from a command including an ambiguous reference. For example, "that" football game may be resolved to a football game on a sports channel. "That" basketball game may be resolved to a basketball game on a college sports channel. A vacuum cleaner commercial may be resolved to a pay-per-view show (e.g., based on metadata associated with a program describing a vacuum cleaner). A legal program may be resolved to courtroom drama based on metadata associated with the program and / or synonym matching, fuzzy matching, or other matching techniques. Various media options 832 appear in the menu 830 on the display 112, which can be used to eliminate ambiguity in the user's request.

[0176] In some embodiments, the displayed menu can be navigated using a cursor, joystick, arrow, button, gesture, etc. In such cases, focus can be displayed on the selected item. For example, the selected item can be displayed using bold text, underline, bounded outline, larger size than other menu items, shadow, reflection, halo, and / or any other features to highlight which menu item is selected and has focus. Figure 23A The media option 2330 selected in the image can have the same focus as the currently selected media option, and is displayed with a large underline type and border.

[0177] In some embodiments, a request to play content or select a menu item may include an ambiguous reference to a menu item that has focus. For example, Figure 23A The user viewing menu 830 can request to play "that" show (e.g., "Play that show"). Similarly, the user can request various other commands associated with the focused menu item, such as play, delete, hide, remind me to watch, record, etc. From the perspective of voice input alone, the specific menu item or show desired may not be clear. However, the content displayed on display 112 can be used to dispel ambiguity in the user's request and determine the user's intent. Specifically, the fact that the selected menu option 2330 is focused in menu 830 can be used to identify the desired media subject, whether it refers to "that" show, a command without a subject (e.g., play, delete, hide, etc.), or any other ambiguous command referring to focused media content. Therefore, the focused menu item can be used when determining the user's intent from voice input.

[0178] Just as viewing history of media content can be used to dispel ambiguity in user requests (e.g., content displayed when a user made a request but since it was submitted), similarly, previously displayed menus or search results can be used to dispel ambiguity in later user requests, for example, after navigating to a later menu or search results. Figure 23B The second page of menu 830, featuring additional media options 832, is shown. Users can proceed to... Figure 23B The second page is shown, but please refer to it again. Figure 23A The content shown on the first page (for example, Figure 23A (See media option 832 shown). For example, even though the user has moved to the second page of menu 830, they can still request to watch "that" football game, "that" basketball game, or a legal program—all of which are media option 832 most recently displayed on the previous page of menu 830. Such references can be ambiguous, but the user's intent can be determined using the most recently displayed menu content from the first page of menu 830. Specifically, analysis can be performed... Figure 23AThe recently displayed media options 832 are used to identify a specific football match, basketball game, or courtroom drama referenced in an ambiguous example request. In some embodiments, results may be biased based on how recently the content was displayed (e.g., weighting recently viewed results pages relative to earlier viewed results). In this way, the viewing history of what was recently displayed on display 112 can be used to determine user intent. It should be understood that any recently displayed content can be used, such as previously displayed search results, previously displayed programs, previously displayed menus, etc. This allows users to refer back to content they saw earlier without having to find and navigate to the specific view where they saw it.

[0179] In other embodiments, various display prompts shown in the menus or result lists on display 112 can be used to dispel ambiguity in the user's request and determine the user's intent. Figure 24 An exemplary media menu is shown, divided into separate sections, one of which has a focus (movie). Figure 24 A category interface 2440 is shown, which may include a disc-shaped interface for categorizing media options, including a TV option 2442, a movie option 2444, and a music option 2446. As shown, only the music category is partially displayed. The disc-shaped interface can be offset to display additional content on the right (e.g., as indicated by the arrow), as if rotating media within a disc. In the illustrated embodiment, the movie category has a focus as indicated by the underlined title and boundary, although the focus can be represented in any of a variety of other ways (e.g., making the category larger to appear closer to the user than other categories, adding a halo, etc.).

[0180] In some embodiments, a request to play content or select a menu item may include an ambiguous reference to a menu item within a set of items (e.g., categories). For example, a user viewing Category Interface 2440 might request to play a football program (“Play a football program.”). The specific menu item or performance desired may not be clear from the voice input alone. Furthermore, the query may resolve to more than one program displayed on Display 112. For example, a request for a football program might reference a football match listed in the Television Programs category or a football movie listed in the Movies category. The content displayed on Display 112—including display prompts—can be used to dispel ambiguity in the user's request and determine the user's intent. Specifically, the fact that the Movies category is focused in Category Interface 2440 can be used to identify the specific football program desired; given focus on a Movies category, it is likely a football movie. Therefore, a category (or any other media grouping) of media with focus as shown on Display 112 can be used when determining the user's intent from the voice input. It should also be understood that a user can make various other requests associated with categories, such as requesting to display content from a specific category (e.g., show me a comedy movie, show me a horror movie, etc.).

[0181] In other embodiments, the user may invoke menus or media items displayed on display 112 in various other ways, and the user's intent may be similarly determined based on the displayed content. It should be understood that the user's intent can also be determined from voice input using metadata associated with the displayed content (e.g., television program descriptions, movie descriptions, etc.), fuzzy matching techniques, synonym matching, etc., in conjunction with the displayed content. Therefore, various forms of user requests—including natural language requests—can be accommodated, and the user's intent can be determined according to the various embodiments discussed herein.

[0182] It should be understood that, in determining user intent, the content displayed on display 112 can be used alone or in combination with the content displayed on user equipment 102 or on a display associated with remote control 106. Similarly, it should be understood that virtual assistant queries can be received at any of various devices communicatively coupled to the set-top box 104, and the content displayed on display 112 can be used to determine user intent, regardless of which device receives the query. Query results can similarly be displayed on display 112 or on another display (e.g., on user equipment 102).

[0183] Furthermore, in any of the various embodiments discussed herein, the virtual assistant system is capable of navigating menus and selecting menu options without requiring the user to explicitly open the menu and navigate to a menu item. For example, an options menu may appear after selecting media content or a menu button, such as selecting... Figure 24 Following the movie option 2444, menu options can include playing media and alternatives to simple media playback, such as setting reminders to watch media later, setting media recording, adding media to the favorites list, and hiding media from further viewing. While the user is viewing content above the menu or content with submenu options, the user can issue virtual assistant commands that would normally require navigation to the menu or submenu for selection. For example, watching... Figure 24 The user of category interface 2440 can issue any menu command associated with movie option 2444 without manually opening the associated menu. For example, a user might request to add a football movie to their favorites list, record the evening news, and set a reminder to watch movie B without having to navigate to the menus or submenus associated with those media options, where such commands might be present. The virtual assistant system can therefore navigate menus and submenus to execute commands on behalf of the user, regardless of whether those menu options appear on display 112. This simplifies user requests and reduces the number of clicks or selections the user must make to achieve the desired menu functionality.

[0184] Figure 25An exemplary process 2500 for controlling television interaction and viewing history of media content displayed on a monitor is shown. In block 2502, voice input can be received from a user, including queries related to content displayed on the television monitor. For example, the voice input may include queries about characters, actors, movies, television programs, sports events, players, etc., appearing on display 112 of system 100 (displayed by set-top box 104). Figure 19 Transcription 1916 includes queries associated with the actress shown in video 480 on display 112. Similarly, Figure 21 The transcription 2122 includes queries associated with characters in the video 480 on display 112. Voice input may also include queries associated with menus or search terms appearing on display 112, such as selecting a specific menu item or obtaining information about a specific search result. For example, the displayed menu content may include... Figure 23A and Figure 23B The media options 832 in the middle menu 830. The displayed menu content may similarly include those appearing in... Figure 24 The category interface 2440 includes the TV option 2442, movie option 2444, and / or music option 2446.

[0185] Refer again Figure 25 In process 2500, at box 2504, the user intent for the query can be determined based on the displayed content and the viewing history of the media content. For example, the user intent can be determined based on the context of a displayed or recently displayed television program, sporting event, movie, etc. The user intent can also be determined based on the displayed or recently displayed menu or search content. The displayed content can also be analyzed together with metadata associated with the content to determine the user intent. For example, references can be used alone or in combination with metadata associated with the displayed content. Figure 19 , Figure 21 , Figure 23A , Figure 23B Based on the content shown and described, determine the user's intent.

[0186] At box 2506, query results can be displayed based on a defined user intent. For example, a display similar to [example display] can be shown on monitor 112. Figure 20 The result of Assistant Response 2020 in the Assistant Response Interface 2018. In another embodiment, it can be... Figure 22The assistant response interface 2224 shown provides text and selectable media as results, such as assistant text response 2226 and selectable video link 2228. In another embodiment, displaying query results may include displaying or playing the selected media content (e.g., playing the selected video on display 112 via television set-top box 104). Thus, user intent can be determined from voice input in various ways by utilizing the displayed content and associated metadata as context.

[0187] In some embodiments, a virtual assistant may provide query suggestions to a user, such as informing the user of available queries, suggesting content the user can enjoy, teaching the user how to use the system, encouraging the user to find additional media content to consume, and so on. In some embodiments, query suggestions may include general suggestions of possible commands (e.g., find a comedy, show me a TV guide, search for action movies, turn on closing captions, etc.). In other embodiments, query suggestions may include targeted suggestions related to the displayed content (e.g., add this show to your watchlist, share this show via social media, show me the audio track for this movie, show me the books this customer is selling, show me the movie trailer this customer is inserting, etc.), user preferences (e.g., use of closing captions, etc.), content owned by the user, content stored on the user's device, notifications, prompts, viewing history of media content (e.g., recently displayed menu items, recently displayed program scenes, recent appearances of actors, etc.), etc. Suggestions may be displayed on any device, including on display 112 via set-top box 104, on user device 102, or on a display associated with remote control 106. Furthermore, recommendations can be determined based on which devices are nearby and / or communicating with the set-top box 104 at a specific time (e.g., recommended content for devices of a user watching television in the room at a specific time). In other embodiments, recommendations can be determined based on a variety of other contextual information, including time of day, information from mass sources (e.g., popular programs watched at a given time), live programming (e.g., live sports events), viewing history of media content (e.g., the last few shows watched, a collection of recently viewed search results, a group of recently viewed media options, etc.), or any of a variety of other contextual information.

[0188] Figure 26An exemplary suggestion interface 2650 is shown, including content-based virtual assistant query suggestion 2652. In one embodiment, query suggestions can be provided in an interface such as interface 2650 in response to input received from a user requesting suggestions. For example, input requesting query suggestions can be received from user device 102 or remote control 106. In some embodiments, input may include button presses, double-clicks, menu selections, voice commands (e.g., show me some suggestions, what can you do for me, what are some options, etc.) received at user device 102 or remote control 106. For example, while viewing an interface associated with a set-top box 104, a user can double-click a physical button on remote control 106 to request query suggestions, or they can double-click a physical or virtual button on user device 102 to request query suggestions.

[0189] The suggestion interface 2650 can be displayed above a moving image, such as video 480, or any other background content (e.g., a menu, a still image, a paused video, etc.). Like other interfaces discussed herein, the suggestion interface 2650 can be animated by sliding upwards from the bottom of the display 112 and can occupy a minimal amount of space while effectively conveying the desired information to limit interference with the video 480 in the background. In other embodiments, a larger suggestion interface (e.g., a paused video, menu, image, etc.) can be provided when the background content is static.

[0190] In some embodiments, virtual assistant query suggestions can be determined based on the displayed media content or the viewing history of the media content (e.g., movies, television programs, sporting events, recently watched programs, recently viewed menus, recently watched movie scenes, recent scenes from television series, etc.). For example, Figure 26 Content-based suggestion 2652 is shown, which can be determined based on the displayed video 480 shown in the background, with characters 1910, 1912, and 1914 appearing on monitor 112. Metadata associated with the displayed content (e.g., descriptive details of the media content) can also be used to determine the query suggestion. Metadata may include various information associated with the displayed content, including program title, list of characters, list of actors, episode description, team roster, team ranking, program summary, movie details, plot description, director's name, producer's name, time of actors' appearances, sports stands, sports scores, genre, list of seasons of a series, related media content, or various other related information. For example, metadata associated with video 480 may include the character names of characters 1910, 1912, and 1914 along with the actresses who played those roles. Metadata may also include descriptions of the plot of video 480, descriptions of the previous or next episode (where video 480 is a television series in a series), etc.

[0191] Figure 26 Various content-based suggestions 2652 are illustrated, which can be displayed in the suggestion interface 2650 based on video 480 and metadata associated with video 480. For example, character 1910 in video 480 could be named "Blanche," and query suggestions about character Blanche or the actress who plays the character could be written using this character name (e.g., "Who is the actress who plays Blanche?"). Character 1910 can be identified from metadata associated with video 480 (e.g., a list of characters, a list of actors, the time associated with the actor's appearance, etc.). In other embodiments, facial recognition can be used to identify actresses and / or characters appearing on display 112 at a given time. Various other query suggestions associated with characters within the media itself can be provided, such as queries related to a character's personality, profile, relationships with other characters, etc.

[0192] In another embodiment, the male or female actor appearing on display 112 can be identified (e.g., based on metadata and / or facial recognition) and query suggestions can be provided associated with that actor or actress. Such query suggestions may include any of the following: the character they have played, acting awards, age, other media appearances, history, family members, relationships, or various other details about the actor or actress. For example, character 1914 could be played by an actress named Whitney Davidson, and the actress's name, Whitney Davidson, could be used to craft query suggestions to identify other films, television programs, or other media appearances of the actress Whitney Davidson (e.g., "What other films or television shows has Whitney Davidson appeared in?").

[0193] In other embodiments, details about the program can be used to write query suggestions. Episode summaries, plot summaries, episode lists, episode titles, series titles, etc., can be used to write query suggestions. For example, a suggestion can be provided to describe what happened in the previous episode of a television program (e.g., “What happened in the last episode?”), to which the virtual assistant system can respond with an episode summary identified from previous episodes based on the episode currently displayed on display 112 (and its associated metadata). In another embodiment, a suggestion can be provided to set a record for the next episode, which can be accomplished by the system identifying the next episode based on the currently playing episode displayed on display 112. In yet another embodiment, a suggestion can be provided to obtain information about the current episode or the program appearing on display 112, using program titles obtained from metadata to write query suggestions (e.g., “What is this episode of ‘Their Show’ about?” or “What is ‘Their Show’ about?”).

[0194] In another embodiment, query suggestions can be crafted using categories, genres, ratings, awards, descriptions, etc., associated with the displayed content. For example, video 480 might correspond to a television program described as a comedy with a female lead. Query suggestions can be crafted from this information to identify other programs with similar characteristics (e.g., “Find other comedies with female leads.”). In other embodiments, suggestions can be determined based on the user’s subscription, content available for playback (e.g., content on set-top box 104, content on user device 102, content available for streaming, etc.). For example, potential query suggestions can be filtered based on the presence of informational or media results. Query suggestions that may not yield playable media content or informational answers can be excluded, and / or those that provide easily accessible informational answers or playable media content (or are given greater weight when determining which suggestions to provide). Thus, query suggestions can be determined using the displayed content and associated metadata in various ways.

[0195] Figure 27 An exemplary selection interface 2754 for confirming a suggested query is shown. In some examples, the user can select the displayed query suggestions by speaking the query, selecting them using buttons, navigating to them using the cursor, etc. In response to the selection, the selected suggestions can be briefly displayed on a confirmation interface, such as selection interface 2754. In one example, the selected suggestions 2756 can be animated to move from wherever they appear in suggestion interface 2650. Figure 27 The location shown is adjacent to the command receiving confirmation 490 (e.g., as indicated by the arrow), and other unselected suggestions can be hidden from the display.

[0196] Figures 28A-28B An exemplary virtual assistant response interface 2862 based on a selected query is shown. In some examples, an informative answer to the selected query can be displayed in the answer interface, such as answer interface 2862. A transition interface 2858 can be displayed when switching from the suggestion interface 2650 or the selection interface 2754, such as... Figure 28A As shown. Specifically, as the next piece of content scrolls up from the bottom of display 112, previously displayed content within the interface can be scrolled upwards to leave that interface. For example, the selected suggestion 2756 can be swiped or scrolled upwards until it disappears at the top edge of the virtual assistant interface, and the assistant result 2860 can be swiped or scrolled from the bottom of display 112 until it reaches... Figure 28B The location shown.

[0197] Answer interface 2862 may include informative answers and / or media results in response to a selected query suggestion (or in response to any other query). For example, in response to a selected query suggestion 2756, an assistant result 2860 may be identified and provided. Specifically, in response to a request for a summary of the previous episode, the previous episode may be identified based on the displayed content, and an associated description or summary may be identified and provided to the user. In the illustrated example, assistant result 2860 may describe the previous episode of the program corresponding to video 480 on display 112 (e.g., “In episode 203 of ‘Their Show,’ Blanche is invited as a guest speaker in a university psychology class. Julia and Melissa make a surprise appearance, causing a commotion.”). Informative answers and media results may also be presented in any of the other ways described herein (e.g., a selectable video link), or results may be presented in a variety of other ways (e.g., the answer is spoken aloud, content is played immediately, animations are displayed, images are displayed, etc.).

[0198] In another example, notifications or prompts can be used to determine the virtual assistant's query suggestions. Figure 29 The suggestion interface 2650 shows media content notification 2964 (although any notification may be considered when making a suggestion) and a suggestion interface that includes both notification-based suggestions 2966 and content-based suggestions 2652 (which may include, as referenced above). Figure 26 (Some of the same concepts described above). In some examples, the content of the notification can be analyzed to identify relevant media-related names, titles, themes, actions, etc. In the illustrated example, notification 2964 includes a prompt to inform the user of alternative media content to be displayed—especially if the sports event is live and the match content may be of interest to the user (e.g., "Five minutes left in the match, Zeta and Alpha are still tied."). In some examples, the notification may be briefly displayed at the top of display 112. The notification can be swiped down from the top of display 112 (as indicated by the arrow) to... Figure 29 The time displayed at the indicated position is a certain amount, and it slides up and down, disappearing at the top of the display 112.

[0199] Notifications or prompts can inform users of various information, such as available alternative media content (e.g., alternatives to content that might be displayed on the current monitor 112), available live television programs, newly downloaded media content, recently added subscription content, suggestions received from friends, media received from another device, etc. Notifications can also be personalized based on the user's family or media viewing identity (e.g., based on user authentication, using account selection, voice recognition, passwords, etc.). In one example, the system can interrupt the display and show a notification based on potentially desired content, such as notification 2964 for users who might want the notification content—based on user profile, supported teams, preferred sports, viewing history, etc. For example, sports scores, match status, remaining time, etc., obtained from sports data feeds, news outlets, social media discussions, etc., can be used to identify possible alternative media content for notification to the user.

[0200] In other examples, popular media content can be offered via prompts or notifications (e.g., among many users) to suggest alternative content to the currently viewed material (e.g., notifying the user that a popular show or a show from a genre the user likes has just started or is available to watch by other means). In the illustrated example, a user might be following either or both of the Zeta and Alpha teams (or might be following football or a specific sport, league, etc.). The system can determine that available live content matches the user's preferences (e.g., a match on another channel matches the user's preferences, with little time remaining and a close score). The system can then decide to prompt the user via notification 2964, which may indicate the content they expect. In some examples, the user can select notification 2964 (or a link within notification 2964) to switch to the suggested content (e.g., using a remote control button, cursor, spoken request, etc.).

[0201] Based on notifications, virtual assistant query suggestions can be determined by analyzing the notification content to identify terms, names, headlines, themes, actions, etc., related to the relevant media. Appropriate virtual assistant query suggestions can then be written using the identified information, such as notification-based suggestion 2966 based on notification 2964. For example, a notification about the exciting ending of a live sports event could be displayed. If the user then requests query suggestions, a suggestion interface 2650 can be displayed, including query suggestions for watching sports events, querying team statistics, or finding content related to the notification (e.g., changing to a Zeta / Alpha match, Zeta team statistics, what other football matches are currently underway, etc.). Similarly, various other query suggestions can be determined and provided to the user based on specific terms of interest identified in the notification.

[0202] Virtual assistant query suggestions related to media content (e.g., for consumption via set-top box 104) can also be determined from the content on the user's device, and suggestions can also be provided on the user's device. In some examples, playable device content can be identified on the user's device that is connected to or communicates with set-top box 104. Figure 30 A user device 102 with exemplary image and video content in interface 1360 is shown. It is possible to determine what content on the user device can be played back, or what content might be desired to be played back. For example, playable media 3068 can be identified based on an active application (e.g., photo and video applications), or it can be identified based on stored content displayed or not displayed on interface 1360 (e.g., in some examples, content can be identified from an active application, or in other examples, it may not be displayed at a given time). Playable media 3068 may include, for example, videos 1362, albums 1364, and photos 1366, each of which may include personal user content that can be sent to a set-top box 104 for display or playback. In other examples, any photos, videos, music, game interfaces, application interfaces, or other media content stored or displayed on user device 102 can be identified and used to determine query suggestions.

[0203] After identifying playable media 3068, the virtual assistant can determine query suggestions and provide them to the user. Figure 31 An exemplary TV assistant interface 3170 on user device 102 is shown, which features virtual assistant query suggestions based on video content playable on the user device and displayed on a separate display (e.g., display 112 associated with TV set-top box 104). TV assistant interface 3170 may include a virtual assistant interface specifically for interacting with media content and / or TV set-top box 104. A user can request query suggestions on user device 102, for example, by double-clicking a physical button while viewing interface 3170. Other inputs can be similarly used to indicate a request for query suggestions. As shown, assistant greeting 3172 may introduce the provided query suggestions (e.g., "Here are some suggestions for controlling your TV experience.").

[0204] The virtual assistant query suggestions provided on user device 102 may include suggestions based on various source devices as well as general suggestions. For example, device-based suggestions 3174 may include query suggestions based on content stored on user device 102 (including content displayed on user device 102). Content-based suggestions 2652 may be based on content displayed on display 112 associated with television set-top box 104. General suggestions 3176 may include general suggestions that may not be associated with specific media content or a specific device having media content.

[0205] For example, device-based recommendations 3174 can be determined based on playable content (e.g., videos, music, photos, game interfaces, application interfaces, etc.) identified on user device 102. In the illustrated example, it is possible to Figure 30 The playable media 3068 shown determines a device-based suggestion 3174. For example, given that album 1364 is identified as playable media 3068, a query can be written using details of album 1364. The system can identify content as an album of multiple photos that can be displayed in a slideshow, and then use the album's title (in some cases) to write a query suggestion to display a slideshow of a specific photo album (e.g., "Show slideshow of 'Graduation Album' from your photos."). In some examples, the suggestion may include an indication of the content source (e.g., "From your photos," "From Jennifer's phone call," "From Daniel's tablet," etc.). The suggestion may also use other details to refer to specific content, such as a suggestion to view photos from a specific date (e.g., show photos from June 21st). In another example, video 1362 can be identified as playable media 3068, and a query suggestion can be written using the video's title (or other identifying information) to play the video (e.g., "Show 'Graduation Video' from your video.").

[0206] In other examples, content available on other connected devices can be identified and used to craft virtual assistant query suggestions. For instance, content from each of the two user devices 102 connected to a public TV set-top box 104 can be identified and used to craft virtual assistant query suggestions. In some examples, users can choose which content can be seen by the system for sharing and can hide other content from the system so as not to include it in query suggestions or otherwise make it available for playback.

[0207] For example, it can be determined based on the content displayed on the display 112 associated with the set-top box 104. Figure 31 The content-based suggestions 2652 shown in interface 3170. In some examples, this can be achieved through references above. Figure 26 The same method is used to determine content-based recommendations 2652. In the illustrated example, Figure 31 The content-based recommendations shown may be based on the video 480 displayed on the display 112 (e.g., such as...). Figure 26 (As in the example). In this way, virtual assistant query suggestions can be derived based on content displayed or available on any number of connected devices. In addition to targeted suggestions, general suggestions can be scheduled and provided (e.g., show me a guide, what sports are happening, what's playing on Channel 3, etc.).

[0208] Figure 32An exemplary suggestion interface 2650 is shown, which has a connected device-based suggestion 3275, along with a content-based suggestion 2652 displayed on a display 112 associated with a TV set-top box 104. In some examples, suggestions can be made through reference above. Figure 26 The same method is used to determine content-based recommendations 2652. As described above, virtual assistant query recommendations can be written based on content on any number of connected devices, and recommendations can be provided on any number of connected devices. Figure 32 A suggestion 3275 based on the connected device is shown, which can be derived from content on user device 102. For example, playable content, such as the photos and videos shown in interface 1360, can be identified on user device 102 as... Figure 30 Playable media 3068. The playable content recognized on user device 102 can then be used to write suggestions that can be displayed on display 112 associated with TV set-top box 104. In some examples, this can be done via reference above. Figure 31 The device-based suggestion 3174 is determined in the same manner as the connected device-based suggestion 3275. Furthermore, as described above, in some examples, source identification information may be included in the suggestion, such as "a call from Jake" as shown in the connected device-based suggestion 3275. Thus, a virtual assistant query suggestion provided on a device can be derived based on content from another device (e.g., displayed content, stored content, etc.). It should be understood that the connected device may include a remote storage device accessible to the set-top box 104 and / or user equipment 102 (e.g., accessing media content stored in the cloud to compose suggestions).

[0209] It should be understood that any combination of virtual assistant query suggestions from various sources can be provided in response to a request for suggestions. For example, suggestions from various sources can be randomly combined, or suggestions can be presented based on popularity, user preferences, selection history, etc. Furthermore, queries can be determined and presented in various other ways based on various other factors, such as query history, user preferences, query popularity, etc. Additionally, in some examples, query suggestions can be automatically cycled by replacing the displayed suggestions with new alternative suggestions after a delay. It should also be understood that users can select suggestions on any interface, for example, by tapping on a touchscreen, speaking the query, selecting a query using navigation keys, selecting a query using a button, selecting a query using a cursor, etc., and then an associated response (e.g., an informative and / or media response) can be provided.

[0210] In any of the examples, virtual assistant query suggestions can also be filtered based on available content. For example, potential query suggestions with informative answers that would result in unavailable media content (e.g., no cable subscription) or potentially irrelevant content can be disqualified and kept from being displayed. On the other hand, potential query suggestions that would result in immediately playable media content that the user can access can be weighted or otherwise biased for display relative to other potential suggestions. In this way, the availability of media content available for the user to watch can also be used when determining which virtual assistant query suggestions to display.

[0211] Furthermore, in any of the examples, pre-loaded query answers may be provided to replace or supplement suggestions (e.g., in suggestion interface 2650). Such pre-loaded query answers may be selected and provided based on personal use and / or the current context. For example, a user watching a particular show may tap a button, double-tap a button, long-press a button, etc., to receive suggestions. As an alternative to or supplement to query suggestions, context-based information may be automatically provided, such as identifying the song or audio track being played (e.g., "This song is a performance show"), identifying the cast members of the currently playing episode (e.g., "Actress Janet Quinn plays Genevieve"), identifying similar media (e.g., "Show Q is similar to this show"), or providing any results of other queries discussed herein.

[0212] Furthermore, a capability representation can be provided in any of the various interfaces for users to rate media content, informing the virtual assistant of the user's preferences (e.g., selectable rating scales). In other examples, the user can speak the rating information as a natural language command (e.g., "I like it," "I don't like it," "I don't like this show," etc.). In other examples, various other functionalities and informational elements can be provided in any of the various interfaces described herein. For example, the interface may also include links to important functions and places, such as search links, shopping links, media links, etc. In another example, the interface may also include recommendations for what to watch next based on the currently playing content (e.g., selecting similar content). In yet another example, the interface may also include recommendations for what to watch next based on personalized tastes and / or recent activity (e.g., selecting content based on user ratings, user-inputted preferences, recently watched shows, etc.). In other examples, the interface may also include instructions for user interaction (e.g., "Press and hold to speak to the virtual assistant," "Tap once for suggestions," etc.). In some examples, providing pre-loaded answers, suggestions, etc. can provide a pleasant user experience, while also making the content easily accessible to a wide range of users (e.g., users of various skill levels, regardless of their language or other control medium).

[0213] Figure 33An exemplary process 3300 is shown for suggesting virtual assistant interactions to control media content (e.g., virtual assistant queries). In block 3302, media content can be displayed on a monitor. For example, such as... Figure 26 As shown, video 480 can be displayed on monitor 112 via TV set-top box 104, or interface 1360 can be displayed on touchscreen 246 of user equipment 102, such as... Figure 30 As shown in the diagram. In box 3304, input can be received from the user. This input may include a request for suggestions from the virtual assistant. This input may include button presses, double-clicks, menu selections, spoken queries for suggestions, etc.

[0214] In box 3306, a virtual assistant query can be determined based on media content and / or the viewing history of that media content. For example, a virtual assistant query can be determined based on displayed programs, menus, applications, media content lists, notifications, etc. In one example, it can be based on video 480 and references. Figure 26 The associated metadata is used to determine content-based recommendations 2652. In another example, recommendations may be based on references. Figure 29 The notification 2964 determines the recommendation 2966 based on the notification. In yet another example, it may be based on the above reference. Figure 30 and 31 Playable media 3068 on user equipment 102 determines device-based recommendations 3174. In other examples, recommendations may be based on the above references. Figure 32 Playable media 3068 on the user equipment 102 is determined based on the recommendation 3275 of the connected device.

[0215] Refer again Figure 33 In process 3300, at box 3308, a virtual assistant query can be displayed on the monitor. For example, as shown in the reference... Figure 26 , 27 The query suggestions determined by the display shown in Figures 29, 31, and 32 are as described above. As mentioned above, query suggestions can be determined and displayed based on various other information. Furthermore, virtual assistant query suggestions provided on the display can be derived based on content from another device with another display. Thus, targeted virtual assistant query suggestions can be provided to the user, thereby assisting the user in understanding potential queries and providing desired content suggestions, among other benefits.

[0216] Furthermore, in any of the examples described herein, aspects can be personalized for a specific user. User data, including contacts, preferences, location, favorite media, etc., can be used to interpret voice commands and facilitate user interaction with the various devices discussed herein. The various processes discussed herein can also be modified in various ways based on user preferences, contacts, text, usage history, profiling data, demographic information, etc. Moreover, such preferences and settings can be updated over time based on user interactions (e.g., frequently spoken commands, frequently selected applications, etc.). The collection and use of user data available from various sources can be used to improve the delivery of invitations or any other content that may be of interest to users. This disclosure anticipates that, in some instances, the collected data may include personal information that uniquely identifies or can be used to contact or locate specific individuals. Such personal information may include demographic data, location-based data, telephone numbers, email addresses, home addresses, or any other identifying information.

[0217] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, the personal information data can be used to deliver targeted content that is of interest to the user. Therefore, the use of such personal information data enables planned control over the delivered content. Furthermore, this disclosure also anticipates other uses of the personal information data that are beneficial to the user.

[0218] This disclosure also anticipates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of these legitimate uses. Furthermore, such collection should only be conducted with the user's informed consent. Additionally, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others with access to such personal information data comply with their privacy policies and procedures. Furthermore, such entities may be subject to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices.

[0219] Regardless of the foregoing, this disclosure also anticipates examples of users selectively blocking the use or access to personal information data. That is, this disclosure anticipates providing hardware and / or software components to prevent or block access to such personal information data. For example, with respect to advertising delivery services, the technology of this invention can be configured to allow users to choose to "join" or "opt out" of the collection of personal information data during service registration. As another example, users may choose not to provide location information for a targeted content delivery service. Yet another example is that users may choose not to provide precise location information but allow the transmission of location area information.

[0220] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not be rendered inoperable due to the absence of all or part of such personal information data. For example, preferences can be inferred based on non-personal information data or an absolute minimum of personal information, such as content requested by a user's associated device, other non-personal information available to the content delivery service, or publicly available information, thereby selecting content and delivering it to the user.

[0221] Based on some examples, Figure 34 A functional block diagram of an electronic device 3400 is shown, configured according to the principles of various described examples, for example, to control television interaction using a virtual assistant and display associated information using different interfaces. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software that execute the principles of the various described embodiments. Those skilled in the art will understand that... Figure 34 The functional blocks described herein can be combined or separated into sub-blocks to implement the principles of the various embodiments described herein. Therefore, the specific implementation herein optionally supports any possible combination or separation or further limitation of the functional blocks described herein.

[0222] like Figure 34 As shown, the electronic device 3400 may include a display unit 3402 (e.g., a display 112, a touchscreen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3400 may also include an input unit 3404 (e.g., a microphone, receiver, touchscreen, button, etc.) configured to receive information, such as voice input, tactile input, gesture input, etc. The electronic device 3400 may also include a processing unit 3406 coupled to the display unit 3402 and the input unit 3404. In some examples, the processing unit 3406 may include a voice input receiving unit 3408, a media content determination unit 3410, a first user interface display unit 3412, a selection receiving unit 3414, and a second user interface display unit 3416.

[0223] Processing unit 3406 may be configured to receive voice input from a user (e.g., via input unit 3404). Processing unit 3406 may also be configured to determine media content based on the voice input (e.g., using media content determination unit 3410). Processing unit 3406 may also be configured to display a first user interface having a first size on display unit 3402 (e.g., using first user interface display unit 3412 on display unit 3402), wherein the first user interface includes one or more selectable links to media content. Processing unit 3406 may also be configured to receive a selection of one or more selectable links from input unit 3404 (e.g., using selection receiving unit 3414 from input unit 3404). Processing unit 3406 may also be configured to, in response to the selection, display a second user interface having a second size larger than the first size on display unit 3402 (e.g., using second user interface display unit 3416 on display unit 3402), wherein the second user interface includes media content associated with the selection.

[0224] In some examples, the first user interface (e.g., of the first user interface display unit 3412) expands into the second user interface (e.g., of the second user interface display unit 3416) in response to a selection (e.g., of the selection receiving unit 3414). In other examples, the first user interface overlaps the playing media content. In one example, the second user interface overlaps the playing media content. In another example, the voice input (e.g., from the input unit 3404 to the voice input receiving unit 3408) includes a query, and the media content (e.g., of the media content determining unit 3410) includes query results. In yet another example, the first user interface includes links to query results in addition to one or more selectable links to the media content. In other examples, the query includes a weather query, and the first user interface includes links to media content associated with the weather query. In another example, the query includes a location, and the links to the media content associated with the weather query include links to a portion of the media content associated with the weather at that location.

[0225] In some examples, in response to the selection, processing unit 3406 can be configured to play media content associated with the selection. In one example, the media content includes a movie. In another example, the media content includes a television program. In yet another example, the media content includes a sporting event. In some examples, the second user interface (e.g., of the second user interface display unit 3416) includes a description of the media content associated with the selection. In other examples, the first user interface includes a link to purchase the media content.

[0226] Processing unit 3406 may also be configured to receive additional voice input from a user (e.g., via input unit 3404), wherein the additional voice input includes a query associated with the displayed content. Processing unit 3406 may also be configured to determine a response to the query associated with the displayed content based on metadata associated with the displayed content. Processing unit 3406 may also be configured to display a third user interface (e.g., on display unit 3402) in response to receiving additional voice input, wherein the third user interface includes the determined response to the query associated with the displayed content.

[0227] Processing unit 3406 may also be configured to receive an instruction for initiating (e.g., via input unit 3404) reception of voice input. Processing unit 3406 may also be configured to display a readiness confirmation (e.g., on display unit 3402) in response to receiving the instruction. Processing unit 3406 may also be configured to display a listening confirmation in response to receiving voice input. Processing unit 3406 may also be configured to detect the end of voice input and display a processing confirmation in response to detecting the end of voice input. In some examples, processing unit 3406 may also be configured to display a transcription of the voice input.

[0228] In some examples, electronic device 3400 includes a television set. In other examples, electronic device 3400 includes a television set-top box. In other examples, electronic device 3400 includes a remote control. In other examples, electronic device 3400 includes a mobile phone.

[0229] In one example, one or more selectable links in the first user interface (e.g., the first user interface display unit 3412) include moving images associated with media content. In some examples, the moving images associated with media content include live feeds of the media content. In other examples, one or more selectable links in the first user interface include still images associated with media content.

[0230] In some examples, processing unit 3406 may also be configured to determine whether the currently displayed content includes a moving image or a control menu; in response to determining that the currently displayed content includes a moving image, selecting a smaller size as a first size of the first user interface (e.g., the first user interface display unit 3412); and in response to determining that the currently displayed content includes a control menu, selecting a larger size than the smaller size as a first size of the first user interface (e.g., the first user interface display unit 3412). In other examples, processing unit 3406 may also be configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of live sports events, and to display a notification including the determined alternative media content.

[0231] Based on some examples, Figure 35 A functional block diagram of an electronic device 3500 is shown, configured according to the principles of various described examples, for example, to control television interaction using a virtual assistant and multiple user devices. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software that execute the principles of the various described embodiments. Those skilled in the art will understand that... Figure 35 The functional blocks described herein can be combined or separated into sub-blocks to implement the principles of the various embodiments described herein. Therefore, the specific implementation herein optionally supports any possible combination or separation or further limitation of the functional blocks described herein.

[0232] like Figure 35 As shown, the electronic device 3500 may include a display unit 3502 (e.g., a display 112, a touchscreen 246, etc.) configured to display media, an interface, and other content. The electronic device 3500 may also include an input unit 3504 (e.g., a microphone, a receiver, a touchscreen, a button, etc.) configured to receive information, such as voice input, tactile input, gesture input, etc. The electronic device 3500 may also include a processing unit 3506 coupled to the display unit 3502 and the input unit 3504. In some examples, the processing unit 3506 may include a voice input receiving unit 3508, a user intent determination unit 3510, a media content determination unit 3512, and a media content playback unit 3514.

[0233] Processing unit 3506 may be configured (e.g., using voice input receiving unit 3508 from input unit 3504) to receive voice input from a user at a first device (e.g., device 3500) having a first display (e.g., display unit 3502 in some examples). Processing unit 3506 may also be configured to determine the user intent of the voice input based on the content displayed on the first display (e.g., using user intent determination unit 3510). Processing unit 3506 may also be configured (e.g., using media content determination unit 3512) to determine media content based on the user intent. Processing unit 3506 may also be configured (e.g., using media content playback unit 3514) to play media content on a second device associated with a second display (e.g., display unit 3502 in some examples).

[0234] In one example, the first device includes a remote control. In another example, the first device includes a mobile phone. In yet another example, the first device includes a tablet computer. In some examples, the second device includes a set-top box. In other examples, the second display includes a television.

[0235] In some examples, the content displayed on the first display includes an application interface. In one example, (e.g., from the voice input receiver 3508 of input unit 3504) the voice input includes a request to play media associated with the application interface. In one example, the media content includes media associated with the application interface. In another example, the application interface includes a photo album, and the media includes one or more photos from the album. In yet another example, the application interface includes a list of one or more videos, and the media includes one of the videos. In other examples, the application interface includes a list of television programs, and the media includes television programs from the list.

[0236] In some examples, processing unit 3506 may also be configured to determine whether a first device is authorized; wherein media content is played on a second device in response to determining that the first device is authorized. Processing unit 3506 may also be configured to identify a user based on voice input and determine the user intent of the voice input based on data associated with the identified user (e.g., using user intent determination unit 3510). Processing unit 3506 may also be configured to determine whether a user is authorized based on voice input; wherein media content is played on a second device in response to determining that the user is an authorized user. In one example, determining whether a user is authorized includes using speech recognition to analyze the voice input.

[0237] In other examples, processing unit 3506 may also be configured to display information associated with media content on a first display of a first device in response to determining that a user intent includes a request for information. Processing unit 3506 may also be configured to play media content on a second device in response to determining that a user intent includes a request to play media content.

[0238] In some examples, voice input includes a request to play content on a second device, and in response to the request, media content is played on the second device. Processing unit 3506 can also be configured to determine whether the determined media content should be displayed on a first or second display based on media format, user preferences, or default settings. In some examples, in response to determining that the determined media content should be displayed on the second display, the media content is displayed on the second display. In other examples, in response to determining that the determined media content should be displayed on the first display, the media content is displayed on the first display.

[0239] In other examples, processing unit 3506 may also be configured to determine the proximity of each of two or more devices, including a second device and a third device. In some examples, media content is played on a second device associated with a second display based on the proximity of the second device relative to the proximity of the third device. In some examples, determining the proximity of each of the two or more devices includes determining proximity based on Bluetooth LE.

[0240] In some examples, processing unit 3506 may also be configured to display a list of display devices including a second device associated with the second display, and receive a selection of the second device from the list of display devices. In one example, media content is displayed on the second display in response to receiving a selection of the second device. Processing unit 3506 may also be configured to determine whether headphones are attached to the first device. Processing unit 3506 may also be configured to display media content on the first display in response to determining that headphones are attached to the first device. Processing unit 3506 may also be configured to display media content on the second display in response to determining that headphones are not attached to the first device. In other examples, processing unit 3506 may also be configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of live sports events, and display a notification including the determined alternative media content.

[0241] Based on some examples, Figure 36 A functional block diagram of an electronic device 3600 is shown, configured according to the principles of various described examples, for example, to control television interaction using media content displayed on a screen and the viewing history of that media content. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software that execute the principles of the various described embodiments. Those skilled in the art will understand that... Figure 36 The functional blocks described herein can be combined or separated into sub-blocks to implement the principles of the various embodiments described herein. Therefore, the specific implementation herein optionally supports any possible combination or separation or further limitation of the functional blocks described herein.

[0242] like Figure 36As shown, the electronic device 3600 may include a display unit 3602 (e.g., a display 112, a touchscreen 246, etc.) configured to display media, interfaces, and other content. The electronic device 3600 may also include an input unit 3604 (e.g., a microphone, receiver, touchscreen, button, etc.) configured to receive information, such as voice input, tactile input, gesture input, etc. The electronic device 3600 may also include a processing unit 3606 coupled to the display unit 3602 and the input unit 3604. In some examples, the processing unit 3606 may include a voice input receiving unit 3608, a user intent determination unit 3610, and a query result display unit 3612.

[0243] Processing unit 3606 may be configured to receive voice input from a user (e.g., from input unit 3604 using voice input receiving unit 3608), wherein the voice input includes a query associated with content displayed on a television display (e.g., display unit 3602 in some examples). Processing unit 3606 may also be configured to determine the user intent of the query based on one or more of the content displayed on the television display and the viewing history of media content (e.g., using user intent determination unit 3610). Processing unit 3606 may also be configured to display query results based on the determined user intent (e.g., using query result display unit 3612).

[0244] In one example, voice input is received at a remote control. In another example, voice input is received at a mobile phone. In some examples, query results are displayed on a television screen. In another example, the content displayed on the television screen includes movies. In yet another example, the content displayed on the television screen includes television programs. In yet another example, the content displayed on the television screen includes sporting events.

[0245] In some examples, the query includes a request for information about a person associated with content displayed on a television display, and the query results (e.g., in query results display unit 3612) include information about that person. In one example, the query results include media content associated with that person. In another example, the media content includes one or more movies, television programs, or sporting events associated with that person. In some examples, the query includes a request for information about a character in content displayed on a television display, and the query results include information about that character or information about the actor playing that character. In one example, the query results include media content associated with the actor playing that character. In another example, the media content includes one or more movies, television programs, or sporting events associated with the actor playing that character.

[0246] In some examples, processing unit 3606 may also be configured to determine query results based on metadata or viewing history of media content associated with the content displayed on the television display. In one example, the metadata includes titles, descriptions, character lists, actor lists, team lists, genres, or display schedules or viewing history of media content associated with the content displayed on the television display. In another example, the content displayed on the television display includes a list of media content, and the query includes a request to display one item from the list. In yet another example, the content displayed on the television display also includes a focused item from the list of media content, and determining the user intent of the query (e.g., using user intent determination unit 3610) includes identifying the focused item. In some examples, processing unit 3606 may also be configured to determine the user intent of the query based on recently displayed menus or search terms on the television display (e.g., using user intent determination unit 3610). In one example, the content displayed on the television display includes a page of listed media, and the recently displayed menus or search terms include media listed on the previous page. In another example, the content displayed on the television display includes one or more categories of media, and one of these categories is in focus. In one example, processing unit 3606 may also be configured to determine the user intent of the query based on the one or more categories of media that is in focus (e.g., using user intent determination unit 3610). In another example, the media categories include movies, television programs, and music. In other examples, processing unit 3606 may also be configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of live sports events, and to display a notification including the determined alternative media content.

[0247] Based on some examples, Figure 37 A functional block diagram of an electronic device 3700 is shown, configured according to the principles of various described examples, for example, to suggest a virtual assistant interaction for controlling media content. The functional blocks of the device may be implemented by hardware, software, or a combination of hardware and software that execute the principles of the various described embodiments. Those skilled in the art will understand that... Figure 37 The functional blocks described herein can be combined or separated into sub-blocks to implement the principles of the various embodiments described herein. Therefore, the specific implementation herein optionally supports any possible combination or separation or further limitation of the functional blocks described herein.

[0248] like Figure 37As shown, the electronic device 3700 may include a display unit 3702 (e.g., a display 112, a touchscreen 246, etc.) configured to display media, an interface, and other content. The electronic device 3700 may also include an input unit 3704 (e.g., a microphone, a receiver, a touchscreen, a button, etc.) configured to receive information, such as voice input, tactile input, gesture input, etc. The electronic device 3700 may also include a processing unit 3706 coupled to the display unit 3702 and the input unit 3704. In some examples, the processing unit 3706 may include a media content display unit 3708, an input receiving unit 3710, a query determination unit 3712, and a query display unit 3714.

[0249] Processing unit 3706 can be configured to display media content on a display (e.g., display unit 3702) (e.g., using media content display unit 3708). Processing unit 3706 can also be configured to receive input from a user (e.g., from input unit 3704 using input receiving unit 3710). Processing unit 3706 can also be configured to determine one or more virtual assistant queries based on media content and one or more of the media content's viewing history (e.g., using query determination unit 3712). Processing unit 3706 can also be configured to display one or more virtual assistant queries on a display (e.g., using query display unit 3714).

[0250] In one example, input is received from the user on a remote control. In another example, input is received from the user on a mobile phone. In some examples, one or more virtual assistant queries are overlaid on a moving image. In another example, input includes a double-click on a button. In one example, the media content includes a movie. In another example, the media content includes television programming. In yet another example, the media content includes a sporting event.

[0251] In some examples, one or more virtual assistant queries include queries about people appearing in media content. In other examples, one or more virtual assistant queries include queries about characters appearing in media content. In yet another example, one or more virtual assistant queries include queries about media content associated with people appearing in media content. In some examples, the media content or the viewing history of the media content includes episodes of a television program, and one or more virtual assistant queries include queries about another episode of the television program. In another example, the media content or the viewing history of the media content includes episodes of a television program, and one or more virtual assistant queries include requests to set reminders to watch or record subsequent episodes of the media content. In yet another example, one or more virtual assistant queries include queries about descriptive details of the media content. In one example, descriptive details include one or more of the following: program title, list of characters, list of actors, episode description, cast list, cast rating, or program summary.

[0252] In some examples, processing unit 3706 may also be configured to receive a selection of one of one or more virtual assistant queries. Processing unit 3706 may also be configured to display the result of the selected virtual assistant query among one or more virtual assistant queries. In one example, determining one or more virtual assistant queries includes determining one or more virtual assistant queries based on one or more of the following: query history, user preferences, or query popularity. In another example, determining one or more virtual assistant queries includes determining one or more virtual assistant queries based on media content that the user can watch. In yet another example, determining one or more virtual assistant queries includes determining one or more virtual assistant queries based on received notifications. In yet another example, determining one or more virtual assistant queries includes determining one or more virtual assistant queries based on active applications. In other examples, processing unit 3706 may also be configured to determine alternative media content for display based on one or more of user preferences, program popularity, and the status of live sports events, and display a notification including the determined alternative media content.

[0253] Although examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art (e.g., modifications to any system or process described herein based on the concepts of any other system or process discussed herein). It should be understood that such changes and modifications are intended to be included within the scope of the various examples defined by the appended claims.

Claims

1. A method for suggesting virtual assistant interactions for controlling media content, the method comprising: At electronic devices: Display media content on a monitor, wherein the media content includes video associated with metadata; Receive input from the user: One or more virtual assistant query suggestions are determined based on one or more of the media content, the viewing history of the media content, and the metadata, as well as the content of the second electronic device; as well as The one or more virtual assistant query suggestions are displayed on the display, overlaid on the media content, wherein the one or more query suggestions include source information corresponding to the second electronic device.

2. The method of claim 1, wherein the input is received from the user on a remote control.

3. The method of claim 1, wherein the input is received from the user on the mobile phone.

4. The method according to any one of claims 1-3, wherein the one or more virtual assistant query suggestions are overlaid on the moving image.

5. The method according to any one of claims 1-4, wherein the input includes double-clicking a button.

6. The method according to any one of claims 1-5, wherein the media content includes a film.

7. The method according to any one of claims 1-6, wherein the media content includes television programs.

8. The method according to any one of claims 1-7, wherein the media content includes sports events.

9. The method according to any one of claims 1-8, wherein the one or more virtual assistant query suggestions include query suggestions about people appearing in the media content.

10. The method according to any one of claims 1-9, wherein the one or more virtual assistant query suggestions include query suggestions about characters appearing in the media content.

11. The method according to any one of claims 1-10, wherein the one or more virtual assistant query suggestions include query suggestions for media content associated with a person appearing in the media content.

12. The method according to any one of claims 1-11, wherein the media content or the viewing history of the media content includes episodes of a television program, and the one or more virtual assistant query suggestions include query suggestions for another episode of the television program.

13. The method according to any one of claims 1-12, wherein the media content or the viewing history of the media content includes episodes of a television program, and the one or more virtual assistant query suggestions include requests to set reminders to watch or record subsequent episodes of the media content.

14. The method according to any one of claims 1-13, wherein the one or more virtual assistant query suggestions include query suggestions for descriptive details of the media content.

15. The method of claim 14, wherein the descriptive details include one or more of the following: program title, list of characters, list of actors, episode description, cast list, cast rating, or program summary.

16. The method according to any one of claims 1-15, further comprising: Receive a selection of one of the one or more virtual assistant query suggestions; as well as Displays the result of the selected virtual assistant query suggestion from the one or more virtual assistant query suggestions.

17. The method of any one of claims 1-16, wherein determining the one or more virtual assistant query suggestions comprises determining the one or more virtual assistant query suggestions based on one or more of the following: query suggestion history, user preferences, or query suggestion popularity.

18. The method of any one of claims 1-17, wherein determining the one or more virtual assistant query suggestions comprises determining the one or more virtual assistant query suggestions based on media content available to the user for viewing.

19. The method of any one of claims 1-18, wherein determining the one or more virtual assistant query suggestions comprises determining the one or more virtual assistant query suggestions based on a received notification.

20. The method of any one of claims 1-19, wherein determining the one or more virtual assistant query suggestions comprises determining the one or more virtual assistant query suggestions based on the active application.

21. The method according to any one of claims 1-20, further comprising: Alternative media content is determined for display based on one or more of user preferences, program popularity, and the status of live sports events. as well as Display a notification that includes the identified alternative media content.

22. A non-transitory computer-readable storage medium comprising computer-executable instructions for performing the method according to any one of claims 1-21.

23. A system comprising: The non-transitory computer-readable storage medium according to claim 22; and A processor capable of executing the computer-executable instructions.

24. A system for suggesting virtual assistant interactions for controlling media content, the system comprising: A means for displaying media content on a display, wherein the media content includes video associated with metadata; A device for receiving input from a user; A means for determining one or more virtual assistant query suggestions based on one or more of the media content, the viewing history of the media content, and the metadata, as well as the content of a second electronic device; and A means for displaying one or more virtual assistant query suggestions overlaid on the media content on the display, wherein the one or more query suggestions include source information corresponding to the second electronic device.

Citation Information

Patent Citations

  • Intelligent automated assistant

    US9318108B2

  • Program guide system with video-on-demand browsing

    CN1585479A

  • Voice control of television-related information

    US20060075429A1