Focus Sessions in Speech Interface Devices

By integrating the sound assistant system and the sound assistant server system into the voice-activated electronic devices, the function of automatically detecting or allocating the target device is solved, and the cumbersome problem of users needing to clearly specify the target device is improved.

JP7675690B2Active Publication Date: 2025-05-13GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022133320
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-11-01
Filing Date
2022-08-24
Publication Date
2025-05-13
Estimated Expiration
2037-11-03

AI Technical Summary

Technical Problem

In the prior art, users need to explicitly specify the target device to execute voice commands, which in some cases can cause trouble and inconvenience, especially if the target device is unknown or unclear.

Method used

By integrating the sound assistant system and the sound assistant server system into the voice-activated electronic device, the function of automatically detecting or distributing the target device is realized. When a target device is not included or unclear in the user's voice command, the system can automatically identify and assign the appropriate target device and send a voice request to the device.

Benefits of technology

It solves the cumbersome problem that users need to clearly specify the target device when providing voice commands, and improves the user experience, especially when the target device is unclear or unclear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675690000001
    Figure 0007675690000001
  • Figure 0007675690000002
    Figure 0007675690000002
  • Figure 0007675690000003
    Figure 0007675690000003
Patent Text Reader

Abstract

A method, an apparatus and a storage medium for sending voice commands to a target device when the target is unknown or ambiguous. The method includes receiving a first voice command including a request for a first operation, determining a first target device from a local group for the first operation, establishing a focus session with the first target device, and causing the first operation to be performed by the first target device. The method further includes receiving a second voice command including a request for a second operation, determining that the second voice command does not include an explicit designation of the second target device, determining that the second operation can be performed by the first target device, determining whether the second voice command satisfies one or more focus session maintenance criteria, and causing the second operation to be performed by the first target device if the second voice command satisfies the focus session maintenance criteria.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Technical Field The disclosed embodiments relate generally to voice interfaces and related devices, including, but not limited to, methods and systems for sending voice commands to a target device when the target device is unknown or ambiguous from the voice command itself. [Background technology]

[0002] background Electronic devices with voice interfaces have been widely used to collect voice input from users and perform different voice-activated functions according to the voice input. These voice-activated functions may include instructing or commanding a target device to perform an operation. For example, a user may issue voice input to a voice interface device to instruct the target device to turn on or off or to control media playback on the target device.

[0003] Typically, when a user wishes to provide speech input to instruct a target device to perform an operation, the user will specify the target device in the speech input. However, it is tedious and annoying for the user to have to explicitly specify a target device for every such speech input. It is desirable for a speech interface device to have a target device for speech input even when the speech input does not specify a target or specifies an ambiguous target. Summary of the Invention [Means for solving the problem]

[0004] overview Therefore, there is a need for an electronic device having a voice assistant system and / or a voice assistant server system incorporating a method and system for determining or assigning a target device for a voice input, even when the target device designation in the voice input is absent or ambiguous. In various embodiments described in this application, an operating environment includes a voice-activated electronic device that provides an interface to a voice assistant service, and a number of devices (e.g., cast devices, smart home devices) that can be controlled by voice input via the voice assistant service. The voice-activated electronic device is configured to record a voice input from which a voice assistance service (e.g., a voice assistance server system) determines a user's voice request (e.g., a media playback request, a power state change request). The voice assistance server system then conveys the user's voice request to the target device indicated by the voice input. The voice-activated electronic device is configured to record a subsequent voice input in which the target device designation is absent or ambiguous. The electronic device or the voice assistance server system assigns a target device for the voice input, determines the user's voice request contained in the voice input, and sends the user's voice request to the assigned target device.

[0005] In accordance with some embodiments, a method is performed on a first electronic device having one or more microphones, a speaker, one or more processors, and a memory storing one or more programs for execution by the one or more processors. The first electronic device is a member of a local group of connected electronic devices that are communicatively coupled to a common network service. The method includes receiving a first voice command including a request for a first operation; selecting a first target device from among the local group of connected electronic devices for the first operation; the first target device performing the first operation via operation of the common network service; receiving a second voice command including a request for a second operation; determining that the second voice command does not include an explicit designation of the second target device; determining that the second operation may be performed by the first target device; determining whether the second voice command satisfies one or more focus session maintenance criteria; and causing the first target device to perform the second operation via operation of the common network service in accordance with a determination that the second voice command satisfies the focus session maintenance criteria.

[0006] According to some embodiments, an electronic device includes one or more microphones, a speaker, one or more processors, and a memory that stores one or more programs executed by the one or more processors, the one or more programs including instructions for performing the above-described method.

[0007] According to some embodiments, a non-transitory computer-readable storage medium stores one or more programs that include instructions that, when executed by an electronic device having one or more microphones, a speaker, and one or more processors, cause the electronic device to perform operations of the methods described above.

[0008] For a better understanding of the various embodiments described above, reference should be made to the following description of implementations in conjunction with the accompanying drawings, in which like reference numerals refer to corresponding parts throughout and in which: [Brief description of the drawings]

[0009] [Figure 1] 1 illustrates an exemplary operating environment according to some embodiments. [Diagram 2] 1 illustrates an exemplary voice-activated electronic device according to some embodiments. [Figure 3A]1 illustrates an exemplary voice assistance server system according to some embodiments. [Figure 3B] 1 illustrates an exemplary voice assistant server system according to some embodiments. [Figure 4A] 1 illustrates an example of a focus session according to some embodiments. [Figure 4B] 1 illustrates an example of a focus session according to some embodiments. [Figure 4C] 1 illustrates an example of a focus session according to some embodiments. [Figure 4D] 1 illustrates an example of a focus session according to some embodiments. [Diagram 5] 1 shows a flow diagram of an example process for establishing a focus session and responding to audio input according to the focus session, according to some embodiments. [Figure 6A] FIG. 1 is a front view of a voice-activated electronic device according to some embodiments. [Figure 6B] FIG. 2 is a rear view of a voice-activated electronic device according to some embodiments. [Figure 6C] FIG. 1 is a perspective view of a voice-activated electronic device 190 showing a speaker included in the base of the electronic device 190 in an open configuration according to some embodiments. [Figure 6D] FIG. 1 is a side view of a voice-activated electronic device showing electronic components contained therein, according to some embodiments. [Figure 6E] 6E(1)-(4) show four touch events detected on a touch sense array of a voice-activated electronic device according to some embodiments, and FIG. 6E(5) shows a user pressing a button on the back of the voice-activated electronic device according to some embodiments. [Figure 6F] FIG. 1 is a top view of a voice-activated electronic device according to some embodiments. [Figure 6G] 1A-1C are diagrams illustrating example visual patterns displayed by an array of full-color LEDs to indicate audio processing states, according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Throughout the drawings, like reference numbers refer to corresponding parts. Description of the embodiments While the digital revolution has brought many benefits to date, from open sharing of information to a sense of global togetherness, new technologies often generate confusion, doubt, and fear among consumers, thus preventing them from benefiting from the technology. Electronic devices are conveniently used as voice interfaces with the ability to receive voice input from users and initiate voice actions. Thereby, electronic devices provide an eyes-free and hands-free solution to approach both existing and new technologies. Specifically, voice input received by the electronic device can convey instructions and information even when the user's gaze is unclear and hands are occupied. To enable a hands-free and eyes-free experience, voice-activated electronic devices listen to the surroundings (i.e., constantly process voice signals collected from the surroundings) either all the time or only when triggered. Meanwhile, the identity of the user is associated with the user's voice and the language used. To protect the user's identity, these voice-activated electronic devices are usually used in private locations, which are protected, controlled, and intimate spaces (e.g., home and car).

[0011] According to some embodiments, if the designation of the target device in the voice command is absent or ambiguous, the voice-activated electronic device determines the target device or assigns the request made in the voice command to the target device. The voice-activated electronic device establishes a focus session with respect to the target device explicitly designated or designated in the voice command. If the voice-activated electronic device receives a subsequent voice command in which the designation or designation of the target device is absent or ambiguous, the voice-activated electronic device assigns the voice command to the target device of the focus session if the voice command meets one or more criteria.

[0012] In some embodiments, when a user interacts with the voice interface device to control another device, the voice interface device remembers which device has been targeted by the user (e.g., in a focus session). From then on, the default target device for control is the remembered device. For example, if the user first issues a voice command "Turn on the kitchen lights" and then "Turn off the lights," the target device for the second voice command will default to "Kitchen lights" if the second command is received immediately after the first command. As another example, if the first command is "Play music on the living room speakers" and the subsequent command is "Stop music," the target device for the second voice command will default to "Living room speakers" if the second command is received immediately after the first command.

[0013] WARNING 9 Furthermore, in some embodiments, if there is a longer time interval between voice inputs, the user may be prompted to confirm or verify that the last used target device is the intended target device. For example, if a first voice command is "Play music on the living room speaker" and a subsequent command received after a longer time interval from the first voice command is "Stop music", the voice interface device may ask the user "Do you want to stop the music on the living room speaker?" to confirm that the target device is the "living room speaker".

[0014] In this way, a user is relieved of the burden of having to specify the complete context of their request in each and every voice input (e.g., having to include a target device designation in each and every voice input requesting an operation to be performed). can be relieved from such burdens.

[0015] Voice assistant experience 1 is an exemplary operating environment according to some embodiments. Operating environment 100 includes one or more voice-activated electronic devices 104 (e.g., voice-activated electronic devices 104-1 through 104-N, hereafter referred to as "voice-activated device(s)"). The one or more voice-activated devices 104 may be located in one or more locations (e.g., throughout multiple spaces within a structure, or all within rooms or spaces of a structure spread across multiple structures (e.g., one in a home and one in a user's car)).

[0016] The environment 100 also includes one or more controllable electronic devices 106 (e.g., electronic devices 106-1 through 106-N, hereafter referred to as "controllable device(s)"). Examples of controllable devices 106 include media devices (smart televisions, speaker systems, wireless speakers, set-top boxes, media streaming devices, casting devices) and smart home devices (e.g., smart cameras, smart thermostats, smart lights, smart hazard detectors, smart door locks).

[0017] The voice-activated devices 104 and the controllable devices 106 are communicatively coupled to the voice assistant service 140 (e.g., to the voice assistance server system 112 of the voice assistant service 140) through a communication network 110. In some embodiments, one or more of the voice-activated devices 104 and the controllable devices 106 are communicatively coupled to a local network 108, which is communicatively coupled to the communication network 110; the voice-activated device(s) 104 and / or the controllable device(s) 106 are communicatively coupled to the communication network(s) 110 through the local network 108 (and to the voice assistance server system 112 through the communication network 110). In some embodiments, the local network 108 is a local area network implemented with a network interface (e.g., a router). The voice-activated devices 104 and the controllable devices 106 communicatively coupled to the local network 108 may also communicate with each other through the local network 108.

[0018] Optionally, one or more of the voice-activated devices 104 are communicatively coupled to the communications network 110 and are not on the local network 108. For example, these voice-activated devices are not on a Wi-Fi network corresponding to the local network 108, but are connected to the communications network 110 via a cellular connection. In some embodiments, communication between the voice-activated devices 104 on the local network 108 and the voice-activated devices 104 not on the local network 108 occurs through the voice assistance server system 112. The voice-activated devices 104 (whether on the local network 108 or on the network 110) are known to the voice assistance server system 112 because they are registered in the device registry 118 of the voice assistant service 140. Similarly, the voice-activated devices 104 not on the local network 108 can communicate with the controllable devices 106 through the voice assistant server system 112. The controllable devices 106 (whether on the local network 108 or on the network 110) are also registered in the device registry 118. In some embodiments, communication between the voice-activated devices 104 and the controllable devices 106 is via the voice assistance server system 112.

[0019] In some embodiments, the environment 100 also includes one or more content hosts 114. The content hosts 114 are remote content hosts from which content is streamed or otherwise obtained according to requests contained in a user's voice input or commands. The content host 114 may be a source from which the voice assistance server system 112 retrieves information according to a user's voice request.

[0020] In some embodiments, the controllable device 106 may receive an instruction or request (e.g., from the voice-activated device 104 and / or the voice assistance server system 112) to perform a specified operation or transition to a specified state, and may perform the operation or transition to a state according to the received instruction or request.

[0021] In some embodiments, one or more of the controllable devices 106 are media devices deployed in the operating environment 100 to provide media content, news, and / or other information to one or more users. In some embodiments, the content provided by the media devices is stored in a local content source, streamed from a remote content source (e.g., content host(s) 114), or generated locally (e.g., from local text to a voice processor that reads customized news briefs, emails, documents, local weather forecasts, etc., to one or more people using the operating environment 100). In some embodiments, the media devices include media output devices that output media content directly to an audience (e.g., one or more users) and networked casting devices that stream media content to the media output devices. Examples of media output devices include, but are not limited to, television (TV) displays and music players. Examples of casting devices include, but are not limited to, set-top boxes (STBs), DVD players, TV boxes, and media streaming devices such as Google's Chromescast® media streaming device.

[0022] In some embodiments, the controllable device 106 is also a voice-activated device 104. In some embodiments, the voice-activated device 104 is also a controllable device 106. For example, the controllable device 106 may include a voice interface to a voice assistance service 140 (e.g., a media device that can also receive, process, and respond to a user's voice input). As another example, the voice-activated device 104 may also perform certain operations and transition to certain states according to requests or commands in the voice input (e.g., a voice interface device that can also play streaming music).

[0023] In some embodiments, the voice-activated devices 104 and the controllable devices 106 are associated with a user with a respective account, or with multiple users (e.g., a group of related users, such as users in a family or organization; more generally, a primary user and one or more authorized additional users, etc.) with respective user accounts in a user domain. The users can input voice inputs or voice commands into the voice-activated devices 104. The voice-activated devices 104 receive these voice inputs from the users (e.g., user 102), and the voice-activated devices 104 and / or the voice assistance server system 112 proceed to determine the request in the voice input and to generate a response to the request.

[0024] In some embodiments, the request included in the voice input is a command or request to cause the controllable device 106 to perform an operation (e.g., play media, pause media, fast forward or rewind media, change the volume, change screen brightness, change light brightness) or transition to another state (e.g., change operating modes, turn on or off, go into or out of sleep mode).

[0025] In some embodiments, the voice-activated electronic device 104 generates and provides voice responses to voice commands (e.g., the current time in response to the question "What time is it?"). speak the time); stream requested media content to the user (e.g., "Play a Bach Boys song"); read out news articles or daily news summaries prepared for the user; play media items stored on the personal assistant device or on the local network; respond to voice input by changing the state of or operating one or more other devices connected within the operating environment 100 (e.g., turning lights, appliances or media devices on / off, locking / opening locks, opening windows, etc.); or by issuing a corresponding request to a server over the network 110.

[0026] In some embodiments, one or more voice-activated devices 104 are deployed in the operating environment 100 to collect voice inputs to initiate various functions (e.g., media playback functions on a media device). In some embodiments, these voice-activated devices 104 (e.g., devices 104-1 through 104-N) are deployed near controllable devices 104 (e.g., media devices), for example, in the same room as cast devices and media output devices. Alternatively, in some embodiments, the voice-activated devices 104 are deployed in a structure that has one or more smart home devices but no media devices. Alternatively, in some embodiments, the voice-activated devices 104 are deployed in a structure that has one or more smart home devices and one or more media devices. Alternatively, in some embodiments, the voice-activated devices 104 are deployed in locations that do not have networked electronic devices. Additionally, in some embodiments, a room or space in a structure may have multiple voice-activated devices 104.

[0027] In some embodiments, the voice-activated device 104 includes at least one or more microphones, a speaker, a processor, and a memory that stores at least one program for execution by the processor. The speaker is configured to enable the voice-activated device 104 to communicate voice messages and other sounds (e.g., audible tones) to the location where the voice-activated device 104 is located in the operating environment 100, thereby broadcasting music, reporting the status of a voice input process, conversing with a user of the voice input device 104, or providing instructions to a user of the voice input device 104. As an alternative to voice messages, visual signals may also be used to provide feedback to a user of the voice-activated device 104 regarding the status of the voice input process. When the voice-activated device 104 is a mobile device (e.g., a mobile phone or tablet computer), its display screen is configured to display notifications regarding the status of the voice input process.

[0028] In some embodiments, the voice-activated device 104 is a voice interface device that is networked to provide voice recognition functionality using the voice assistance server system 112. For example, the voice-activated device 104 includes a smart speaker that provides music to the user and allows eyes-free and hands-free access to a voice assistant service (e.g., Google Assistant). Optionally, the voice-activated device 104 is one of a desktop or laptop computer, a tablet, a mobile phone including a microphone, a casting device including a microphone and optionally a speaker, an audio system including a microphone and a speaker (e.g., a stereo system, a speaker system, a portable speaker, etc.), a television including a microphone and a speaker, and a user interface system of an automobile including a microphone, a speaker, and optionally a display. Optionally, the voice-activated device 104 is a simple, low-cost voice interface device. In general, the voice-activated device 104 can be any device that is network-enabled and includes a microphone, a speaker, and programs, modules, and data for interacting with a voice assistant service. Given the simplicity and low cost of the voice-activated device 104, the voice-activated device 104 may include an array of light-emitting diodes (LEDs) rather than a full display screen, with visual indications on the LEDs to indicate the status of the voice input process. In some embodiments, the LEDs are full color LEDs and the color of the LED may be employed as part of the visual pattern displayed on the LED. For example, several examples of using LEDs to display visual patterns to convey information or device status (e.g., indicating whether a focus session has been initiated, a status associated with it is active is extended, and / or which individual user of multiple users is associated with a particular focus session) are described below with reference to FIG. 6. In some embodiments, the visual pattern indicating the status of the voice processing operation is displayed using a distinctive image shown on a conventional display associated with the voice-activated device performing the voice processing operation.

[0029] In some embodiments, LEDs or other visual displays are used to communicate the collective voice processing status of multiple participating electronic devices. For example, in an operating environment with multiple voice processing or voice interface devices (e.g., multiple electronic devices 104 as shown in FIG. 6A; multiple voice-activated devices 104 of FIG. 1), a group of colored LEDs (e.g., LED 604 as shown in FIG. 6) associated with each electronic device can be used to communicate which electronic devices are listening to the user and which of the listening devices is the leader (the "leader" device typically takes the lead role in responding to voice requests issued by the user).

[0030] More generally, the following discussion with reference to FIG. 6 describes an “LED design language” for visually indicating various voice processing states of an electronic device, such as a hot word detection state, a listening state, a think mode, a work mode, a reply mode, and / or a busy mode, using a collection of LEDs. In some embodiments, unique states of the voice processing operations described herein are represented using groups of LEDs in accordance with one or more aspects of the “LED design language.” These visual indicators may also be combined with one or more audible indicators generated by the electronic device performing the voice processing operations. The resulting audio and / or visual indicators enable a user in a voice interaction environment to understand the state of various voice processing electronic devices in the environment and to effectively interact with those devices in a natural and intuitive manner.

[0031] In some embodiments, when voice input to the voice-activated device 104 is used to control media output devices via a cast device, the voice-activated device 104 effectively enables a new level of control of cast-enabled media devices. In a specific example, the voice-activated device 104 includes a casual enjoyment speaker with far-field voice access capability and serves as a voice interface device for a voice assistant service. The voice-activated device 104 can be deployed in any area in the operating environment 100. When multiple voice-activated devices 104 are distributed in multiple rooms, they are synchronized to become cast audio receivers providing audio input from these rooms.

[0032] Specifically, in some embodiments, the voice-activated device 104 includes a Wi-Fi speaker with a microphone that is connected to a voice-activated voice assistant service (e.g., Google Assistant). A user can issue a media playback request via the microphone of the voice-activated device 104, asking the voice assistant service to play media content on the voice-activated device 104 itself or on another connected media output device. For example, a user can issue a media playback request by saying to the Wi-Fi speaker, "Okay, Google, play cat videos on my living room TV." The voice assistant service then fulfills the media playback request by playing the requested media content on the requested device using a default or specified media application.

[0033] In some embodiments, a user can issue voice requests via the microphone of the voice-activated device 104 regarding media content that is already playing or is playing on the display device (e.g., the user can request information about the media content, purchase the media content at an online store, or create and publish a social post about the media content).

[0034] In some embodiments, users may want to take their current media session with them as they move around the house and may request such a service from one or more of the voice-activated devices 104. This requests that the voice assistant service 140 transfer the current media session from a first cast device to a second cast device that is not directly connected to the first cast device or is unaware of the existence of the first cast device. Following the transfer of the media content, a second output device coupled to the second cast device continues playing the media content that was previously playing on the first output device coupled to the first cast device from the exact point in the music track or video clip where the media content was playing on the first output device. In some embodiments, a voice-activated device 104 that receives a request to transfer a media session can fulfill the request. In some embodiments, a voice-activated device 104 that receives a request to transfer a media session relays the request to another device or system (e.g., the voice assistance server system 112) for processing.

[0035] Additionally, in some embodiments, a user may issue a request for information or a request to perform an action or operation via the microphone of the voice-activated device 104. The requested information may be personal (e.g., the user's email, the user's calendar events, the user's flight information, etc.), non-personal (e.g., sports scores, news articles, etc.), or somewhere in between (e.g., the scores of the user's favorite teams or sports, news articles from the user's favorite sources, etc.). The requested information or action / operation may include access to personal information (e.g., purchasing digital media items with payment information provided by the user, purchasing physical goods). The voice-activated device 104 responds to the request with a voice message response to the user, which may include, for example, a request for additional information to fulfill the request, confirmation that the request has been fulfilled, a notice that the request cannot be fulfilled, etc.

[0036] In some embodiments, in addition to the voice-activated devices 104 and the media devices among the controllable devices 106, the operating environment 100 may also include one or more smart home devices among the controllable devices 106. The integrated smart home devices include intelligent, multi-sensor, networked devices that seamlessly integrate with each other and / or with a central server or cloud computing system in a smart home network to provide a variety of useful smart home functions. In some embodiments, the smart home devices are deployed in the same location of the operating environment 100 as the cast device and / or output device, and thus are located in close proximity or at a known distance from the cast device and output device.

[0037] The smart home devices in the operating environment 100 may include one or more intelligent, multi-sensor, networked thermostats, one or more intelligent, multi-sensor, networked hazard detectors, one or more intelligent, multi-sensor, networked interface devices (hereinafter referred to as "smart doorbells" and "smart door locks"), one or more intelligent, multi-sensor, networked alarm systems, one or more intelligent, multi-sensor, networked The smart home devices may include, but are not limited to, one or more camera systems, one or more intelligent, multi-sensor networked wall switches, one or more intelligent, multi-sensor networked power sockets, and one or more intelligent, multi-sensor networked lights. In some embodiments, the smart home devices in the operating environment 100 of FIG. 1 may include a plurality of intelligent, multi-sensor networked appliances (hereinafter referred to as "smart appliances"), such as refrigerators, stoves, ovens, televisions, washers, dryers, lights, stereos, intercom systems, garage door openers, floor fans, ceiling fans, wall mounted air conditioners, pool heaters, irrigation systems, security systems, heating appliances, window AC units, powered duct vents, etc. In some embodiments, any one of these smart home device types may be equipped with a microphone and one or more voice processing capabilities described herein to respond in whole or in part to voice requests from a current occupant or user.

[0038] In some embodiments, each of the controllable devices 104 and voice-activated devices 104 can communicate data and share information with other controllable devices 106, voice-activated electronic devices 104, a central server or cloud computing system, and / or other networked devices (e.g., client devices). Data communication can be performed using any of a variety of conventional or standard wireless protocols (e.g., IEEE 802.15.4, Wi-Fi, ZigBee, 6LoWPAN, Thread, Z-Wave, Bluetooth Smart, ISA100.1la, WirelessHART, MiWi, etc.) and / or any of a variety of conventional or standard wired protocols (e.g., Ethernet, HomePlug, etc.), or any other suitable communication protocol, including communication protocols yet to be developed as of the filing date of this document.

[0039] Through a communication network (e.g., the Internet) 110, the controllable devices 106 and the voice-activated devices 104 can communicate with a server system (also referred to herein as a central server system and / or a cloud computing system). Optionally, the server system may be associated with a manufacturer, support entity, or service provider associated with the controllable devices and the media content displayed to the user. Thus, the server system includes a voice assistance server 112 that processes voice input collected by the voice-activated devices 104, one or more content hosts 114 that provide the displayed media content, optionally a crowdcast service server that creates a virtual user domain based on the distributed device terminals, and a device registry 118 that maintains a record of the distributed device terminals in the virtual user environment. Examples of distributed device terminals include, but are not limited to, the controllable devices 106, the voice-activated devices 104, and the media output devices. In some embodiments, these distributed device terminals are linked to user accounts (e.g., Google user accounts) in the virtual user domain. It should be understood that the processing of voice inputs collected by the voice-activated device 104, including generating responses to those inputs, can be performed locally at the voice-activated device 104, at the voice assistance server 112, at another smart home device (e.g., a hub device or a controllable device 106), or a combination of all or a subset of the above.

[0040] It will be appreciated that in some embodiments, the voice-activated device(s) 104 can function in environments without smart home devices. For example, the voice-activated device 104 can respond to user requests for information or to perform actions and / or initiate or control various media playback functions without the presence of smart home devices. The voice-activated device 104 can also function in a wide range of environments, including, but not limited to, in a vehicle, boat, business, or manufacturing environment.

[0041] In some embodiments, the voice-activated device 104 is "woke" by a voice input that includes a hot word (also referred to as a "wake word") (e.g., activating an interface for a voice assistant service on the voice-activated device 104 and placing the voice-activated device 104 in a state in which the voice-activated device 104 is ready to receive a voice request to the voice assistant service). In some embodiments, the voice-activated device 104 is required to wake up if it has been inactive for at least a predetermined time (e.g., 5 minutes) with respect to receiving voice input; the predetermined time corresponds to the amount of inactivity allowed before a voice interface session or conversation times out. The hot word may be a word or phrase, may be a predetermined default, and / or may be customized by the user (e.g., a user may set a nickname for a particular voice-activated device 104 as the device's hot word). In some embodiments, there may be multiple hot words that can wake up the voice-activated device 104. The user can speak the hotword and wait for an acknowledgment response from the voice-activated device 104 (e.g., the voice-activated device 104 outputs a greeting) and they make the first voice request. Alternatively, the user can combine the hotword and the first voice request into one voice input (e.g., the voice input includes the hotword followed by the voice request).

[0042] In some embodiments, the voice-activated device 104 interacts with controllable devices 106 (e.g., media devices, smart home devices), client devices, or server systems of the operating environment 100 according to some embodiments. The voice-activated device 104 is configured to receive voice input from an environment proximate to the voice-activated device 104. Optionally, the voice-activated device 104 stores the voice input and processes the voice input at least partially locally. Optionally, the voice-activated device 104 communicates the received voice input, or the partially processed voice input, to a voice assistance server system 112 via a communication network 110 for further processing. The voice-activated device 104, or the voice assistance server system 112, determines whether there is a request in the voice input and what the request is, determines and generates a response to the request, and communicates the request to one or more controllable device(s) 106. The controllable device(s) 106 that receive the response are configured to perform an action or change state according to the response. For example, the media device may be configured to obtain media or Internet content from one or more content hosts 114 for display on an output device coupled to the media device in response to requests in the audio input.

[0043] In some embodiments, the controllable device(s) 106 and the voice-activated device(s) 104 are linked to each other in a user domain, and more specifically, associated with each other through user accounts in the user domain. Information about the controllable devices 106 (whether on the local network 108 or the network 110) and the voice-activated devices 104 (whether on the local network 108 or the network 110) is stored in a device registry 118 in association with the user account. In some embodiments, there is a device registry for the controllable devices 106 and a device registry for the voice-activated devices 104. The controllable device registry can reference devices in the voice-activated device registry that it is associated with in the user domain, and vice versa.

[0044] In some embodiments, one or more voice-activated devices 104 (and one or more cast devices) and one or more controllable devices 106 are commissioned to the voice assistant service 140 via the client device 103. In some embodiments, the voice-activated device 104 does not include a display screen at all and relies on the client device 103 to provide a user interface during the commissioning process. And, for the controllable device 106, The same is true for the voice-activated device 104 and / or the controllable device 106 located near the client device. Specifically, an application is installed on the client device 103 that enables a user interface to facilitate the delegation of authority of a new voice-activated device 104 and / or a controllable device 106 located near the client device. A user may send a request on the user interface of the client device 103 to initiate the delegation process for the new electronic device 104 / 106 that needs to be delegated. After receiving the delegation request, the client device 103 establishes a short-range communication link with the new electronic device 104 / 103 that needs to be delegated. Optionally, the short-range communication link is established based on Near Field Communication (NFC), Bluetooth, Bluetooth Low Energy (BLE), and the like. Then, the client device 103 transmits wireless setting data related to a wireless local area network (WLAN) (e.g., local network 108) to the new device or electronic device 104 / 106. The wireless setting data includes at least a WLAN security code (i.e., a service set identifier (SSID) password), and optionally includes an SSID, an Internet Protocol (IP) address, proxy settings, and gateway settings. After receiving the wireless setting data via the short-range communication link, the new electronic device 104 / 106 decodes and recovers the wireless setting data, and joins the WLAN based on the wireless setting data.

[0045] In some embodiments, the additional user domain information is entered into a user interface displayed on the client device 103 and used to link the new electronic device 104 / 106 to an account in the user domain. Optionally, the additional user domain information is communicated to the new electronic device 104 / 106 along with wireless communication data via a short-range communications link. Optionally, the additional user domain information is communicated to the new electronic device 104 / 106 via the WLAN after the new device joins the WLAN.

[0046] Once the electronic device 104 / 106 has been authorized into a user domain, other devices and their associated operations can be controlled via multiple control paths. According to one control path, applications installed on the client device 103 are used to control the other devices and their associated operations (e.g., media playback operations). Or, according to another control path, the electronic device 104 / 106 is used to enable eyes-free and hands-free control of the other devices and their associated operations.

[0047] In some embodiments, the voice-activated devices 104 and the controllable devices 106 may be assigned nicknames by the user (e.g., by the primary user with which the devices are associated in the user domain). For example, a speaker device in the living room may be assigned the nickname "Living Room Speaker." In this way, the user can more easily refer to the device with voice input by speaking the device nickname. In some embodiments, the device nicknames and corresponding mappings to devices are stored in the voice-activated device 104 (which stores nicknames only for devices associated with the same user as the voice-activated device) and / or in the voice assistance server system 112 (which stores device nicknames associated with different users). For example, the voice assistance server system 112 stores multiple device nicknames and mappings across different devices and users, and the voice-activated device 104 associated with a particular user downloads the nicknames and mappings for the devices associated with the particular user for local storage.

[0048] In some embodiments, a user may group one or more of the voice-activated devices 104 and / or controllable devices 106 into user-created device groups. Similar to referencing individual devices by nicknames, groups may be given names and groups of devices may be referenced by group names. Device Nicknames Similarly, device groups and group names may be stored in the voice-activated device 104 and / or the voice assistance server system 112 .

[0049] The voice input from the user may explicitly specify a target controllable device 106, or a target group of devices, for the request in the voice input. For example, a user may issue the voice input "play classical music on the living room speakers." The target devices in the voice input are the "living room speakers"; the request in the voice input is a request to have the "living room speakers" play classical music. As another example, a user may issue the voice input "play classical music on the house speakers," where "house speakers" is the name of a group of devices. The target device group in the voice input is the "house speakers"; the request in the voice input is a request to have the devices in the "house speakers" group play classical music.

[0050] The voice input from the user may not have an explicit designation of a target device or device group; there is no reference to a target device or device group by name in the voice input. For example, following the exemplary voice input above, "Play classical music on living room speakers," the user may utter a subsequent voice input, "pause." The voice input does not include a designation of a target device for a request for a pause operation. In some embodiments, the designation of a target device in the voice input may be ambiguous. For example, the user may have uttered a device name incompletely. In some embodiments, if there is no explicit target device designation or the designation of the target device is ambiguous, a target device or device group may be assigned to the voice input, as described below.

[0051] In some embodiments, when the voice-activated device 104 receives a voice input with an explicit designation of a target device or device group, the voice-activated device 104 establishes a focus session with respect to the specified target device or device group. In some embodiments, the voice-activated device 104 stores for the focus session a session start time (e.g., a timestamp of the voice input based on which the focus session was started) and the specified target device or device group as the focused device for the focus session. In some embodiments, the voice-activated device 104 also logs subsequent voice inputs in the focus session. The voice-activated device 104 logs at least the most recent voice input in the focus session, and optionally logs and retains previous voice inputs in the focus session as well. In some embodiments, the voice assistance server system 112 establishes the focus session. In some embodiments, the focus session may be terminated by a voice input that explicitly designates a different target device or device group.

[0052] While a focus session for a device is active and the voice-activated device receives voice input, the voice-activated device 104 makes one or more determinations regarding the voice input. In some embodiments, the determinations include: whether the voice input includes an explicit target device designation, whether the requirements in the voice input are one that can be fulfilled by the focused device, and the time of the voice input compared to the time of the last voice input in the focus session and / or the session start time. If the voice input does not include an explicit target device designation, can be fulfilled by the focused device, and meets predefined time criteria with respect to the time of the last voice input in the focus session and / or the session start time, then the focused device is assigned as the target device for the voice input. Further details regarding focus sessions are described below.

[0053] Devices in the operating environment FIG. 2 is a block diagram illustrating an exemplary voice-activated device 104 adapted as a voice interface for collecting user voice commands in an operating environment (e.g., operating environment 100) according to some embodiments. The voice-activated device 104 typically includes one or more processing units (CPUs) 202, one or more network interfaces 204, memory 206, and one or more communication buses 208 for interconnecting these components (sometimes referred to as a chipset). The voice-activated device 104 includes one or more input devices 210 for facilitating user input, such as buttons 212, a touch-sensitive array 214, and one or more microphones 216. The voice-activated device 104 also includes one or more output devices 218, including one or more speakers 220, optionally an array of LEDs 222, and optionally a display 224. In some embodiments, the array of LEDs 222 is an array of full-color LEDs. In some embodiments, the voice-activated device 104 includes either an array of LEDs 222, a display 224, or both, depending on the type of device. In some embodiments, the voice-activated device 104 also includes a location detection device 226 (eg, a GPS module) and one or more sensors 228 (eg, an accelerometer, a gyroscope, a light sensor, etc.).

[0054] Memory 206 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. Memory 206 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 206, or the non-volatile memory within memory 206, includes a non-transitory computer-readable storage medium. In some embodiments, memory 206, or the non-transitory computer-readable storage medium of memory 206, stores the following programs, modules, and data structures, or a subset or superset thereof: · an operating system 232 that contains procedures for handling various basic system services and for performing hardware-dependent tasks; a network communications module 234 for connecting the voice-activated device 104 to other devices (e.g., a voice assistance service 140, one or more controllable devices 106, one or more client devices 103, and other voice-activated device(s) 104) via one or more network interfaces 204 (wired or wireless) and one or more networks 110, such as the Internet, other wide area networks, local area networks (e.g., local network 108), metropolitan area networks, etc.; An input / output control module 236 for receiving input via one or more input devices and enabling presentation of information on the voice-activated device 104 via one or more output devices 218, including: a voice processing module 238 for processing voice input or voice messages collected in the environment surrounding the voice-activated device 104 or for preparing the collected voice input or voice messages for processing in the voice assistance server system 112; an LED control module 240 for generating visual patterns on the LEDs 222 according to the device state of the voice-activated device 104; and ○ A touch sense module 242 for detecting touch events on the top surface of the voice-activated device 104 (e.g., on the touch sensor array 214); A voice-activated device data 244 for storing at least data related to the voice-activated device 104, including: ○ Common device settings (service layer, device model, memory capacity, processing capacity, communication capacity, etc.), information on one or more user accounts in the user domain, device nicknames and device groups, settings related to restrictions when dealing with unregistered users, and settings related to restrictions when dealing with unregistered users by LED222. a voice device configuration 246 for storing information relating to the voice-activated device 104 itself, including display specifications relating to one or more visual patterns to be displayed upon the voice-activated device 104; and ○ Voice control data 248 for storing voice signals, voice messages, response messages, and other data related to the voice interface functions of the voice-activated device 104; a response module 250 for executing instructions included in the voice request responses generated by the voice assistance server system 112 and, in some embodiments, generating responses to certain voice inputs; and A focus session module 252 for establishing, managing, and terminating a focus session with respect to the device.

[0055] In some embodiments, the audio processing module 238 includes the following modules (not shown): · a user identification module for identifying and disambiguating users providing voice input to the voice input device 104; a hotword recognition module for determining whether the voice input includes a hotword to activate the voice-activated device 104 and for recognizing such in the voice input; and A request recognition module for determining a user request contained in the speech input.

[0056] In some embodiments, the memory 206 also stores focus session data 254 for outstanding focus sessions, including: · Identifiers of the device or device group focused in an outstanding focus session (e.g., device nickname, device group name, device MAC address(es) for storing the device(s) on which the session is focused 256; Session start time 258 for storing a timestamp for the start of an outstanding focus session; and · Session command history 260 for storing a log of previous requests or commands in a focus session, including at least the most recent request / command. The log includes at least the timestamp(s) of the previous request(s) / command(s) recorded in the log.

[0057] Each of the above identified elements may be stored in one or more of the aforementioned memory devices and corresponds to a set of instructions for performing the functions described above. The above identified modules or programs (i.e., sets of instructions) need not be implemented as separate software programs, procedures, modules, or data structures, and thus various subsets of these modules may be combined or otherwise rearranged in various implementations. In some embodiments, memory 206 optionally stores a subset of the above identified modules and data structures. Additionally, memory 206 optionally stores additional modules and data structures not described above. In some embodiments, a subset of the programs, modules, and / or data stored in memory 206 may be stored on and / or executed by voice assistance server system 112.

[0058] In some embodiments, one or more of the modules in memory 206 described above are part of an audio processing library of modules that may be implemented and embedded in a wide variety of devices.

[0059] 3A-3B are block diagrams illustrating an example voice assistance server system 112 of a voice assistant service 140 of an operating environment (e.g., operating environment 100) according to some embodiments. The server system 112 typically includes one or more processing units (CPU(s)) 302, one or more network interfaces 304, memory 306, and one or more communication buses 308 for interconnecting these components (sometimes referred to as a chipset). The server system 112 may include one or more input devices 310 to facilitate user input, such as a keyboard, a mouse, a voice command input unit or microphone, a touch screen display, a touch sensitive input pad, a gesture capture camera, or other input buttons or controls. Additionally, the server system 112 may use a microphone and voice recognition, or a camera and gesture recognition, to supplement or replace a keyboard. In some embodiments, the server system 112 includes one or more cameras, scanners, or optical sensor units, for example for capturing images of a graphic series code printed on an electronic device. The server system 112 may also include one or more output devices 312 to enable presentation of a user interface and display content, including one or more speakers and / or one or more visual displays.

[0060] Memory 306 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and, optionally, non-volatile memory, such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. Memory 306 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 306, or the non-volatile memory within memory 306, includes a non-transitory computer-readable storage medium. In some embodiments, memory 306, or the non-transitory computer-readable storage medium of memory 306, stores the following programs, modules, and data structures, or a subset or superset thereof: · an operating system 316 that contains procedures for handling various basic system services and for performing hardware-dependent tasks; · a network communications module 318 for connecting the server system 112 to other devices (e.g., client devices 103, controllable devices 106, voice-activated devices 104) via one or more network interfaces 304 (wired or wireless) and one or more networks 110, such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc.; · a user interface module 320 for enabling presentation of information on the client device (e.g., application(s) 322-328, widgets, websites and their web pages, and / or graphical user interfaces for presenting games, audio and / or video content, text, etc.); An instruction execution module 321 for server-side execution (e.g., games, social network applications, smart home applications, and / or other web or non-web based applications for controlling client devices 103, controllable devices 106, voice-activated devices 104, and smart home devices and reviewing data captured by such devices), including one or more of the following: o A cast device application 322 that runs to provide server-side functionality for device provisioning, device control, and user account management associated with the cast device(s); o One or more media player applications 324 executed to provide server-side functionality for media presentation and user account management associated with corresponding media sources; one or more smart home device applications 326 that execute to provide server-side functionality for device provisioning, device control, data processing, and data review for a corresponding smart home device; and o Processing a voice message received from a voice-activated device 104 for organizing voice processing of the voice message or for extracting a user's voice command and one or more parameters for the user's voice command (e.g., designation of a cast device or another voice-activated device 104). a voice assistance application 328 that directly processes the Server system data 330 that stores at least data related to the automated control of media display (e.g., in automatic media output mode and follow-up mode) and other data, including one or more of the following: o Client device configuration 332 for storing information associated with one or more client devices, including common device configuration (e.g., service tier, device model, storage capacity, processing capabilities, communication capabilities, etc.) and information for automated media display control; o Cast device settings 334 for storing information associated with a user account of the cast device application 322, including one or more of account access information, information for device configuration (e.g., service tier, device model, storage capacity, processing capabilities, communication capabilities, etc.), and information for automatic media display control; o media player application settings 336 for storing information associated with a user account of one or more media player applications 324, including one or more of account access information, user preferences for media content types, review history data, and information for automated media display control; o smart home device settings 338 for storing information related to a user account of the smart home application 326, including one or more of the following: account access information, information for one or more smart home devices (e.g., service tier, device model, storage capacity, processing capabilities, communication capabilities, etc.); o Voice assistance data 340 for storing information related to a user account of the voice assistance application 328, including one or more account access information, information for one or more voice-activated devices 104 (e.g., service tier, device model, storage capacity, processing capabilities, communication capabilities, etc.); ○ User Data 342 for storing information about users in the user domain, including user subscriptions (e.g., music streaming service subscriptions, video streaming service subscriptions, newsletter subscriptions), user devices (e.g., devices registered in the device registry 118 associated with each user, device nicknames, device groups), user accounts (e.g., users' email accounts, calendar accounts, financial accounts, etc.), and other user data; ○ A user voice profile 344 for storing a voice profile of a user in a user domain, including, for example, a voice model or voice fingerprint of the user and the user's comfortable volume level threshold; and o Focus Session Data 346 for storing focus session data for multiple devices.

[0061] · a device registration module 348 for managing the device registry 118; a voice processing module 350 for processing voice input or voice messages collected in the environment surrounding the electronic device 104; and A focus session module 352 for establishing, managing, and terminating a focus session with respect to the device.

[0062] With reference to FIG. 3B , in some embodiments, the memory 306 also stores focus session data 346 for one or more outstanding focus sessions 3462-1 through 3462-M, including: · a session source device 3464 for storing an identifier of a device with which a focus session is established; · session focused device(s) 3466 for storing an identifier of the focused device or device group in an outstanding focus session (e.g., device nickname, device group name, device MAC address(es)); Session timestamp for remembering the start of an outstanding focus session tion start time 3468; and · Session command history 3470 for storing a log of previous requests or commands in a focus session, including at least the most recent request / command.

[0063] In some embodiments, the voice assistance server system 112 is primarily responsible for processing voice input, and thus one or more of the programs, modules, and data structures in memory 206 described above with reference to FIG. 2 are included in respective modules in memory 306 (e.g., the programs, modules, and data structures included in voice processing module 238 are included in voice processing module 350). The voice-activated device 104 communicates the captured voice input to the voice assistance server system 112 for processing, or first pre-processes the voice input and communicates the pre-processed voice input to the voice assistance server system 112 for processing. In some embodiments, the voice assistance server system 112 and the voice-activated device 104 have some shared responsibilities and some divided responsibilities with respect to processing the voice input, and the programs, modules, and data structures shown in FIG. 2 may be included in both the voice assistance server system 112 and the voice-activated device 104, or divided among the voice assistance server system 112 and the voice-activated device 104. Other programs, modules, and data structures shown in FIG. 2, or the like, may also be included in the voice assistance server system 112.

[0064] Each of the above elements may be stored in one or more of the memory devices mentioned above and corresponds to an instruction set for performing the functions described above. The above modules or programs (i.e., instruction sets) need not be implemented as separate software programs, procedures, modules, or data structures, and therefore various subsets of these modules may be combined or rearranged in various embodiments. In some embodiments, memory 306 stores a subset of the above modules and data structures, if desired. Additionally, memory 306 stores additional modules and data structures not described above, if desired.

[0065] Focus Session Example 4A-4D illustrate an example of a focus session according to some embodiments. In an operating environment with a voice-activated device 104 (e.g., operating environment 100) and multiple controllable devices 106, when a user in the environment provides voice input designating one of the controllable devices 106 as a target device, a focus session may be established with the target device as the focused device.

[0066] FIG. 4A shows a voice-activated device 404 (e.g., voice-activated device 104) and three controllable devices 406, 408, and 410 (e.g., controllable device 106) in an operating environment (e.g., operating environment 100). The devices may be in the same space (e.g., in the same room) as user 402 or may be spread throughout the structure in which the user is located. Device 406 is a speaker system nicknamed "Master Bedroom Speakers." Device 408 is a media device nicknamed "Living Room TV." Device 410 is a media device nicknamed "Game Room TV." There is currently no focus session; focus session 418 is empty.

[0067] A user 402 issues a voice input 403 of "play a cat video on the game room TV," and a voice-activated device 404 receives the voice input. The voice-activated device 404 determines that the request in the voice input 403 is a request to play a cat video, and the target device is the "game room TV" device 410 explicitly specified in the voice input 403. A session 418, in which the focused device is the "Game Room TV" device 410, is established with the voice-activated device 404, as shown in Figure 4B. A command to play the cat video is sent (by device 404 or by the voice assistance server system 112) to the "Game Room TV" device 410, which performs operation 416.

[0068] 4C , while the session 418 is active with the “Game Room TV” 410 in focus and the operation 416 is being performed by the device 410, the user 402 issues another voice input “pause” 420. The voice-activated device 404 determines whether the voice input 420 includes a designation of a target device and whether the request in the voice input 420 can be performed by the focused device 410. For the particular voice input 420 “pause”, the voice-activated device 404 determines that the voice input 420 does not include a designation of a target device and that the request in the voice input (to “pause whatever is playing”) can be performed by the focused device. In some embodiments, determining whether the voice input 420 includes a designation of a target device includes looking for a match to a device nickname in the voice input (e.g., performing speech-to-text recognition on the voice input and parsing the text to look for the device nickname). In some embodiments, determining whether the request in the voice input can be executed by the focused device includes determining what the request in the voice input is and comparing the request to the command history (e.g., history 260) of the current focus session 418 for consistency with the last command in the session (e.g., a “pause music” request is inconsistent with the most recent command being “pause music”), and comparing the request to the capabilities of the focused device for consistency (e.g., a “pause music” request is inconsistent with the capabilities of a smart light).

[0069] In some embodiments, the voice-activated device 404 also determines whether the audio input 420 meets one or more focus session maintenance criteria. In some embodiments, the focus session maintenance criteria is that the timestamp of the audio input 420 is within a certain time from the timestamp of the last audio input 403 in the active session (e.g., the second audio input is received within a certain time of the previous first audio input). In some embodiments, there are multiple time thresholds for this criterion. For example, there may be a first shorter time threshold (e.g., 20 minutes) and a second longer time threshold (e.g., 4 hours). If the audio input 420 is received within the first shorter threshold of the last audio input 403 and the other criteria above are met, the focused device is set as the target device for the audio input 420 (and in some embodiments, this target device setting is also conveyed when conveying the audio input 420 to the voice assistance server system 112 for processing). For example, it is determined that audio input 420 does not include a designation of a target device and that the request "pause" is consistent with the last command "play the cat video." If audio input 420 is received within the shorter time threshold of audio input 403, then the focused device, "Game Room TV" device 410, is set as the target device for audio input 420, and operation 416 being performed on "Game Room TV" device 410 is pausing the cat video in accordance with audio input 420, as shown in FIG. 4D.

[0070] If the voice input 420 is received after the first short threshold of the last voice input 403 and within the second long threshold, and the other criteria above are met, the voice-activated device 404 outputs a voice prompt to request confirmation from the user that the focused device is the desired target device for the voice input 420. If the voice-activated device 404 receives confirmation that the focused device is the desired target device, it maintains the session 418 and sets the focused device as the target device for the voice input 420 (and, in some embodiments, selects the focused device as the voice assistant for processing). (The voice-activated device 404 also communicates this target device setting when communicating the voice input 420 to the service server system 112. If the user does not confirm the target device, the voice-activated device 404 may request that the user specify a target device, that the user repeat the voice input but include the target device specification, and / or terminate the session 418. In some embodiments, the session 418 is terminated if the voice input 420 is received after the second longer threshold since the last voice input 403, or if the other criteria above are not met. In some embodiments, the values ​​of these time thresholds are stored in memory 206 and / or memory 306. The elapsed time between voice inputs is compared to these thresholds.

[0071] In some embodiments, the absence of an explicitly specified target device in the voice input and the consistency of the voice input request with the last voice input and the capabilities of the focused device are also considered focus session maintenance criteria.

[0072] Process Example 5 is a flow diagram illustrating a method 500 for responding to a user's voice input, according to some embodiments. In some embodiments, the method 500 is implemented on a first electronic device (e.g., a voice-activated device 104) that includes one or more microphones, a speaker, one or more processors, and a memory that stores one or more programs for execution by the one or more processors. The first electronic device is a member of a local group of connected electronic devices (e.g., voice-activated devices 104 and controllable devices 106 associated with a user account; controllable devices 106 associated with a particular voice-activated device 104, etc.) that are communicatively coupled (via network 110) to a common network service (e.g., voice assistance service 140).

[0073] A first electronic device receives a first voice command including a request for a first operation (502). For example, a voice-activated device 404 receives a first voice input 403.

[0074] The first electronic device determines (504) a first target device for a first operation from among the local group of connected electronic devices. The voice-activated device 404 determines (e.g., based on processing by the voice processing module 238) a target device (or group of devices) for the audio input 403 from among devices 406, 408, and 410. The voice-activated device 404 recognizes the target device designation “Game Room TV” in the audio input 403 as the “Game Room TV” device 410.

[0075] The first electronic device establishes 506 a focus session with the first target device (or device group). The voice-activated device 404 (e.g., focus session module 252) establishes a focus session 418 with the “Game Room TV” device 410 as the focused device.

[0076] The first electronic device causes the first operation to be performed by the first target device (or group of devices) via operation of the common network service (508). The voice-activated device 404 or the voice assistance server system 112 communicates an instruction to the device 410 via the voice assistance service 140 to perform the operation requested in the voice input 403.

[0077] The first electronic device receives 510 a second voice command including a request for a second operation. The voice-activated device 404 receives 420 a second voice input.

[0078] The first electronic device receives a second voice command to identify a second target device (or group of devices). The voice-activated device 404 determines 512 a target device for the voice input 420 (e.g., based on processing by the voice processing module 238) and recognizes that the voice input 420 does not include a designation of a target device.

[0079] The first electronic device determines 514 that the second operation can be performed by the first target device (or device group). The voice-activated device 404 determines that the operation requested in the voice input 420 can be performed by the focused device 410 and is consistent with the last operation requested in the voice input 403 and performed by the focused device 410.

[0080] The first electronic device determines whether the second voice command satisfies one or more focus session maintenance criteria (516). The voice-activated device 404 determines whether the voice input 420 is received within a certain time period of the voice input 403.

[0081] In accordance with the determination that the second voice command satisfies the focus session maintenance criteria, the first electronic device causes the second operation to be executed by the first target device (or group of devices) via operation of the common network service (518). The voice-activated device 404 determines that the voice input 420 was received within the first shorter time threshold of the voice input 403 and, in accordance with the determination, sets the target device for the voice input 420 to the focused device 410. The voice-activated device 404 or the voice assistance server system 112 communicates an instruction to the device 410 via the voice assistance service 140 to perform the operation requested in the voice input 420.

[0082] In some embodiments, determining a first target device for the first operation from among the local group of connected electronic devices includes obtaining an explicit designation of the first target device from the first voice command. The voice-activated device 404 may pre-process the voice input 403 to determine whether the voice input 403 includes an explicit designation of the target device. Alternatively, the voice-activated device 404 may receive the explicit designation of the target device from the voice assistance server system 112 that processed the voice input 403.

[0083] In some embodiments, determining a first target device for a first operation among the local group of connected electronic devices includes determining that the first voice command does not include an explicit designation of the first target device, determining that the first operation can be performed by a second electronic device among the local group of connected electronic devices, and selecting the second electronic device as the first target device. If the first voice input does not include an explicit designation of a target, but the request included in the first voice input is one that can be performed by a single device in the group (e.g., a video-related command, and there is only one video-enabled device in the group), then that single device is set as the target device for the first voice input. Furthermore, in some embodiments, if there is only one controllable device in addition to the voice-activated device, then that controllable device is the default target device for the voice input, and the voice input does not explicitly designate a target device, and the requested operation of the voice input can be performed by the controllable device.

[0084] In some embodiments, a user's voice input history (e.g., collected by the voice assistance server system 112 and stored in memory 306, and collected by the voice-activated device 104 and stored in memory 206) may be analyzed (e.g., by the voice assistance server system 112 or the voice-activated device 104) to determine whether the history indicates that a particular voice-activated device 104 is frequently used to control a particular controllable device 106. If the history indicates such a relationship, then the particular controllable device 106 may be selected as the voice-activated device 104. may be set as the default target device for voice input to the voice-activated device.

[0085] In some embodiments, a designation (eg, an identifier) ​​of the default target device is stored in the voice-activated device 104 and / or the voice assistance server system 112 .

[0086] In some embodiments, the focus session is extended for the first target device in accordance with a determination that the second voice command satisfies the focus session maintenance criteria. In some embodiments, the focus session times out (i.e., ends) after a certain amount of time. If the second voice input 420 satisfies the focus session maintenance criteria, the focus session 418 may be extended in time (e.g., resetting a timeout timer).

[0087] In some embodiments, establishing a focus session with respect to the first target device includes storing a timestamp of the first voice command and storing an identifier of the first target device. When a focus session is established after receiving a voice input 403, the voice-activated device 404 stores the time of the voice input 403 (e.g., in the session command history 260) and the identifier of the focused device 410 (e.g., in the session focused device 256).

[0088] In some embodiments, the focus session maintenance criteria include a criterion that the second voice command is received by the first electronic device within a first predetermined time interval relative to receipt of the first voice command or a second predetermined time interval relative to receipt of the first voice command, the second predetermined time interval following the first predetermined time interval; and determining whether the second voice command satisfies the one or more focus session maintenance criteria includes determining whether the second voice command is received either within the first predetermined time interval or within the second predetermined time interval. The voice-activated device 404 determines whether the voice input 420 satisfies one or more focus session maintenance criteria, including whether the voice input 420 is received within a first time threshold or a second time threshold of the voice input 403.

[0089] In some embodiments, in accordance with determining that the second voice command is received within the first predetermined time interval, the first electronic device selects the first target device as the target device for the second voice command. If the voice input 420 is determined to be received within a first shorter time threshold from the voice input 403, the focused device 410 is set as the target device for the voice input 420.

[0090] In some embodiments, following a determination that the second voice command is received within the second predetermined time interval, the first electronic device outputs a request to confirm the first target device as a target device of the second voice command; and following a positive confirmation of the first target device in response to the request to confirm, selects the first target device as a target device for the second voice command. If it is determined from the voice input 403 that the voice input 420 is received outside the first shorter time threshold but within the second longer time threshold, the voice-activated device prompts the user to confirm the target device (e.g., asks the user whether the focused device 410 is the intended target device). If the user confirms that the focused device 410 is the intended target device, the focused device 410 is set as the target device for the voice input 420.

[0091] In some embodiments, the first electronic device receives a third voice command including a request for a third operation and an explicit designation of a third target device in the local group of connected electronic devices. The voice-activated device 404 may receive a command to terminate the focus session with respect to the first target device, establish a focus session with respect to the third target device, and cause a third operation to be performed by the third target device via operation of the common network service. The voice-activated device 404 may receive a new voice input after the voice input 420 that includes an explicit designation of a target device other than the device 410 (e.g., device 406 or 408). Pursuant to receiving the voice input, the focus session 418 with the focused device 410 is terminated and a new session is established with the new focused target device. The voice-activated device 404 or the voice assistance server system 112 communicates a command via the voice assistance service 140 to the new target device to perform the operation requested in the new voice input.

[0092] In some embodiments, the first target device is the first electronic device. The first electronic device receives a fourth voice command including a request for a fourth operation and an explicit designation of a fourth target device in a local group of connected electronic devices, where the fourth target device is a third electronic device member of the local group of connected electronic devices, the third electronic device being different from the first electronic device; the first electronic device maintains a focus session with respect to the first target device; and causes the fourth operation to be performed by the fourth target device through operation of a common network service. If the focused device for the active focus session 418 on the voice-activated device 404 is the voice-activated device 404 itself, and a new voice input is received after the voice input 420 that designates another device as the target, the voice-activated device 404 or the voice assistance server system 112 communicates an instruction to the other target device via the voice assistance service 140 to perform the operation requested in the new voice input, but the focus session is maintained with the voice-activated device 404 in focus.

[0093] In some embodiments, the second voice command is received after a fourth operation is caused to be performed by the fourth target device, the first operation being a media play operation and the second operation being a media stop operation. The first electronic device receives a fifth voice command including a request for a fifth operation and an explicit designation of a fifth target device from among a local group of connected electronic devices, in which the fifth target device is the third electronic device; the first electronic device terminates a focus session with respect to the first target device; establishes a focus session with respect to the fifth target device; and causes the fifth target device to perform the fifth operation via operation of the common network service. If the focused device for the active focus session 418 on the voice-activated device 404 is the voice-activated device 404 itself, the voice input 403 includes a request to start media playback, the voice input 403 includes a request to pause media playback as a result of the voice input 403, and new voice input is received after the voice input 420 that targets a different device, the voice-activated device 404 or the voice assistance server system 112 communicates a command to the different target device via the voice assistance service 140 to perform the operation requested in the new voice input, and the focus session with the focused voice-activated device is terminated and a new focus session with the new focused target device is established.

[0094] In some embodiments, the first electronic device receives a fifth voice command including a request to end the predetermined operation, and pursuant to receiving the fifth voice command, causes the first operation to be stopped from being performed by the first target device and ends the focus session with respect to the first target device. If the voice-activated device 404 receives the predetermined end command (e.g., "Stop"), the voice-activated device 404 or the voice assistance server system 112 communicates a command to the device 410 via the voice assistance service 140 to cease performing the operation 416 and the focus session 418 is ended.

[0095] In some embodiments, the first operation is a media play operation and the second operation is one of a media stop operation, a media rewind operation, a media fast forward operation, a volume up operation, and a volume down operation. The request at the audio input 403 may be a request to start playing media content (e.g., video, music) and the request at the audio input 420 may be a request to control the playback (e.g., pause, rewind, fast forward, volume up / down, next item / track, previous item / track).

[0096] In some embodiments, the first operation is a device state change operation to a first state of the plurality of device states and the second operation is a device state change operation to a second state of the plurality of device states. The request in audio input 403 may be a request to transition to the first state (e.g., power on a light or device, enter sleep mode) and the request in audio input 420 may be a request to transition to the second state (e.g., power off a light or device, return from sleep mode).

[0097] In some embodiments, the first operation is a magnitude change operation in a first direction on the magnitude scale and the second operation is a magnitude change operation in a second direction on the magnitude scale that is opposite to the first direction.

[0098] The request at audio input 403 may be a request to change the magnitude in one direction (e.g., brighten, increase the volume) and the request at audio input 420 may be a request to change the magnitude in the opposite direction (e.g., dim, decrease the volume).

[0099] In some embodiments, the first electronic device includes an array of one or more LEDs. The first electronic device indicates the status of the focus session by illuminating one or more LEDs in the array of LEDs. The voice-activated device 404 may indicate that there is an active focus session, or other status and information related to the focus session (e.g., an indication of how long the focus session has been active or how much time has passed since the last voice input), by displaying a pattern on the LED array.

[0100] In some embodiments, a focus session may be established for each identified user. For example, if a user provides voice input specifying a target device, the user is identified and a focus session is established for the identified user with the target device specified by the voice input in focus. If a different user provides voice input specifying a different target device, a different user is identified and another focus session is established for the different identified users with the different target devices in focus. Voice inputs spoken by different users and that do not specify a target device may be assigned different target devices based on the active focus sessions corresponding to each identified user.

[0101] Adding a focus session example The following embodiments describe embodiments in the context of a voice-activated device 104 in the same room as one or more controllable devices 106 that are media devices. It should be understood that the embodiments described below can be adapted to other types of controllable devices 106 (e.g., smart home devices) and other device configurations.

[0102] In some embodiments, if there is no media already playing on the voice-activated device, a focus session may be initiated using a focus device that is a controllable device other than the voice-activated device. When media is paused, a focus session can be started with a controllable device other than a voice-activated device as the focus device.

[0103] In some embodiments, a focus session is initiated when a user issues any valid request that has an explicit target device directed to a device or group of devices associated with the voice-activated device (and, optionally, on the same Wi-Fi network as the voice-activated device). Examples of such valid requests include "play music on living room speakers," "turn up volume on bedroom TV," "next song on homegroup," "pause living room speakers," etc. The explicit target device becomes the focused device for the focus session.

[0104] In some embodiments, if the request is explicitly a video-related request and there is a single video-capable device among the associated controllable devices, a focus session may be established with the video-capable device as the focused device.

[0105] In some embodiments, while a voice-activated device is actively playing media, if a request is received with another device as the target device, the focus remains on the voice-activated device, but once the voice-activated device stops or pauses the session, any new requests to play or control media on another device will move the focus to the other device.

[0106] For example, the user requests "play Lady Gaga," and the voice-activated device begins playing Lady Gaga music, beginning a focus session with the voice-activated device in focus. The user then requests "pause," and the voice-activated device pauses the Lady Gaga music (and maintains the focus session for, say, two hours). After one hour has passed, the user requests "play cat videos on my TV." Focus moves to the TV, and the TV begins playing cat videos.

[0107] As another example, a user requests "play Lady Gaga," and the voice-activated device begins playing Lady Gaga music and initiates a focus session with the voice-activated device in focus. The user then requests "show cat videos on my TV," and cat videos begin to appear on the TV, but focus remains on the voice-activated device. The user then requests "next," and the voice-activated device complies with the request and advances to the next track in the Lady Gaga music. The user then requests "pause," and the music on the voice-activated device is paused. The user then requests "next slide on my TV," and the next slide begins on the TV and focus is moved to the TV.

[0108] In some embodiments, valid requests include starting music, starting a video, starting a news read (such as reading a news article), starting a podcast, starting a photo (such as displaying or showing a photo slideshow), and any media control command (other than a predefined STOP command that ends any current focus session).

[0109] In some embodiments, the focus session ends when any of the following occurs: · the focus session is transferred to a different device (e.g. via voice input, explicitly specifying a different device), in which case the focus session is started with the different device; via voice input or casting from another device (e.g. via voice: "Play Lady Gaga on <nickname of voice interface device>", "Play Lady Gaga locally", etc.; via casting: user casts content to the voice-activated device via an application on a client device), on the voice-activated device A focus session is started or resumed (from a paused state); However, if the voice-activated device is a member (follower or leader) of the group trying to play media, it will not stop focus (even if it is playing), so focus will remain with the leader of the group (which could be another voice-activated device); · when the request is a "stop" command of a given (including all associated grammar) to the focused controllable device; Timeout related commands: ○ The timeout may be measured from the last request or command other than a given "stop" command given to the controllable device, whether the controllable device is explicitly specified or set based on the focused device of a focus session; ○ The timeout is 240 minutes across the various possible commands; and When a user presses a pause / play button on a voice-activated device (and any paused content is resumed locally on the voice-activated device).

[0110] In some embodiments, the voice-activated device requests confirmation from the user of the target device. The user is prompted if they want to play media on the controllable device, as follows: Confirmation requests are triggered for media starts (e.g. starting music when nothing is playing) (for media controls like fast forward or next track); When a focus session becomes active, a confirmation request is triggered; and The confirmation request is triggered a certain amount of time (e.g. 20 minutes) after the last voice command, other than a given "stop" command, given from the current voice-activated device to a controllable device, whether the controllable device is specified explicitly or set based on the focused device of a focus session.

[0111] A verification request might look like this: Voice-activated devices should output "Do you want me to play on <controllable device name>?"

[0112] ○ The user responds "Yes," and the requested media is played on the focused controllable device and focus is maintained on that device.

[0113] ○ The user responds "No," and the requested media is played on the voice-activated device and the focus session is ended.

[0114] Other: For example, if the user's response is unclear, the voice-activated device may output "Sorry, I didn't understand your response."

[0115] In some embodiments, when a focus session is initiated, media initiation and voice-based control commands are applied to the focused controllable device. Non-media requests (e.g., searches, questions) are answered by the voice-activated device, and non-media requests do not end the focus session.

[0116] In some embodiments, even when a focus session is initiated, physical interactions still control the voice-activated device, so that physical interactions with the voice-activated device (e.g., pressing a button, touching a touch-sensitive area) to change volume and pause / play affect the voice-activated device and not necessarily the controllable device.

[0117] In some embodiments, a timer / alarm / text being played on a voice-activated device A request or command issued to speak text has a higher priority than a similar request or command to a focused controllable device. For example, if a voice-activated device is ringing a timer or alarm and the user says "stop," the voice-activated device will stop the timer or alarm. If the user then says "turn volume up / down," the timer or alarm will still be stopped and the volume on the controllable device will be changed, turned up or down.

[0118] As another example, if a voice-activated device is playing text-to-speech (e.g. reading the user's email) and the user says "stop," the voice-activated device will stop reading the text. If the user then says "volume <up / down>," the volume on the voice-activated device will be changed, either up or down.

[0119] As yet another example, if a voice-activated device is paused, paused, or an application is loaded and the user says "stop," media playback on the controllable device will be stopped and the focus session will be ended. If the user then says "volume <up / down>," the volume on the controllable device will be changed, either up or down.

[0120] Physical characteristics of voice-activated electronic devices 6A and 6B are front and rear views 600 and 620 of a voice-activated electronic device 104 (FIG. 1) according to some embodiments. The electronic device 104 includes one or more microphones 602 and an array of full-color LEDs 604. The full-color LEDs 604 may be hidden under a top surface of the electronic device 104 and may be invisible to a user when they are not lit. In some embodiments, the array of full-color LEDs 604 is physically arranged in a ring. Additionally, the rear of the electronic device 104 optionally includes a power connector 608 configured to couple to a power source.

[0121] In some embodiments, the electronic device 104 presents a clean appearance with no visible buttons, and interaction with the electronic device 104 is based on voice and touch gestures. Alternatively, in some embodiments, the electronic device 104 includes a limited number of physical buttons (e.g., button 606 on its back surface), and interaction with the electronic device 104 is based on button presses in addition to voice and touch gestures.

[0122] One or more speakers are provided in the electronic device 104. Figure 6C is a perspective view 660 of a voice-activated electronic device 104 showing a speaker 622 housed in a base 610 of the electronic device 104 in an open configuration, according to some embodiments. The electronic device 104 includes an array of full-color LEDs 604, one or more microphones 602, a speaker 622, a dual-band WiFi 802.11ac radio, a Bluetooth LE radio, an ambient light sensor, a USB port, a processor, and a memory that stores at least one program for execution by the processor.

[0123] 6D , electronic device 104 further includes a touch sense array 624 configured to detect touch events on a top surface of electronic device 104. Touch sense array 624 may be disposed and hidden beneath the top surface of electronic device 104. In some embodiments, the touch sense array is arranged on a top surface of a circuit board that includes an array of via holes, and full color LEDs 604 are disposed within the via holes of the circuit board. When the circuit board is disposed beneath the top surface of electronic device 104, both full color LEDs 604 and touch sense array 624 are disposed beneath the top surface of electronic device 104 as well.

[0124] 6E(1)-6E(4) show four touch events detected on the touch sense array 624 of the voice-activated electronic device 104, according to some embodiments. 6E(2), the touch sense array 624 detects a rotational swipe on the top surface of the voice-activated electronic device 104. In response to detecting a clockwise swipe, the voice-activated electronic device 104 increases the volume of its audio output, and in response to detecting a counterclockwise swipe, the voice-activated electronic device 104 decreases the volume of its audio output. With reference to FIG. 6E(3), the touch sense array 624 detects a single tap touch on the top surface of the voice-activated electronic device 104. In response to detecting a first tap touch, the voice-activated electronic device 104 performs a first media control operation (e.g., playing a particular media content), and in response to detecting a second tap touch, the voice-activated electronic device 104 performs a second media control operation (e.g., pausing a particular media content that is currently being played). With reference to FIG. 6E(4), the touch sense array 624 detects a double tap touch (e.g., two consecutive touches) on the top surface of the voice-activated electronic device 104. The two consecutive touches are separated by less than a predetermined amount of time. However, when they are separated by more than a predetermined amount of time, the two consecutive touches are considered as two single tap touches. In some embodiments, in response to detecting the double tap touch, the voice-activated electronic device 104 initiates a hotword detection state in which the electronic device 104 listens for and recognizes one or more hotwords (e.g., predetermined keywords). Until the electronic device 104 recognizes the hotword, the electronic device 104 does not send any voice input to the voice assistance server 112 or the crowdcast service server 118. In some embodiments, in response to detecting the one or more hotwords, a focus session is initiated.

[0125] In some embodiments, the array of full-color LEDs 604 is configured to display a set of visual patterns according to an LED design language to indicate detection of a clockwise swipe, a counterclockwise swipe, a single tap, or a double tap on the top surface of the voice-activated electronic device 104. For example, the array of full-color LEDs 604 may be illuminated sequentially to track a clockwise or counterclockwise swipe, as shown in Figures 6E(1) and 6E(2), respectively. Further details regarding the visual patterns associated with the voice processing states of the electronic device 104 are described below with reference to Figures 6F and 6G(1)-6G(8).

[0126] 6E(5) illustrates an exemplary user touch or press of a button 606 on the back of a voice-activated electronic device 104, according to some embodiments. In response to a first user touch or press of button 606, the microphone of electronic device 104 is muted, and in response to a second user touch or press of button 606, the microphone of electronic device 104 is activated.

[0127] An LED Design Language for Visual Comfort of Speech User Interfaces In some embodiments, the electronic device 104 includes an array of full-color light-emitting diodes (LEDs) rather than a full display screen. An LED design language is employed to configure the illumination of the array of full-color LEDs and to enable different visual patterns that indicate different audio processing states of the electronic device 104. The LED design language consists of a grammar of colors, patterns, and specific actions that are applied to a fixed set of full-color LEDs. Elements in the language are combined to visually indicate specific device states during use of the electronic device 104. In some embodiments, the illumination of the full-color LEDs is intended to clearly delineate, among other important states, the passive and active listening states of the electronic device 104. States that can be visually indicated by the LEDs (e.g., LED 604) using similar LED design language elements include the state of one or more focus sessions, the identity of one or more users associated with one or more particular focus sessions, and / or the duration of one or more active focus sessions. For example, in some embodiments, a focus session can be visually indicated as being active, extended due to detection of a second audio input, and / or recently terminated due to a lack of user audio interaction with the electronic device 104. A different light pattern, color combination, and / or a particular movement of the LEDs 604 can be used to indicate expiration. One or more identities of one or more users associated with a particular focus session can also be indicated with different light patterns, color combinations, and / or a particular movement of the LEDs 604 that visually identify the particular users. The arrangement of full-color LEDs conforms to the physical constraints of the electronic device 104, and an array of full-color LEDs can be used in speakers manufactured by a third party original equipment manufacturer (OEM) based on a particular technology (e.g., Google Assistant).

[0128] In a voice-activated electronic device 104, passive listening occurs when the electronic device 104 processes voice input collected from its surrounding environment but does not store the voice input or communicate the voice input to any remote server. In contrast, active listening occurs when the electronic device 104 stores the voice input collected from its surrounding environment and / or shares the voice input with a remote server. According to some embodiments of the present application, the electronic device 104 only passively listens for voice input in its surrounding environment without violating the privacy of the user of the electronic device 104.

[0129] FIG. 6G is a top view of the voice-activated electronic device 104, according to some embodiments, and FIG. 6H shows six exemplary visual patterns displayed by an array of full-color LEDs to indicate voice processing states, according to some embodiments. In some embodiments, the electronic device 104 does not include any display screen, and the full-color LEDs 604 provide a simple, low-cost visual user interface compared to a full display screen. The full-color LEDs may be hidden under the top surface of the electronic device and invisible to the user when not lit. With reference to FIG. 6G and FIG. 6H, in some embodiments, the array of full-color LEDs 604 is physically arranged in a ring. For example, as shown in FIG. 6H(6), the array of full-color LEDs 604 may be sequentially lit to track clockwise or counterclockwise swipes, as shown in FIG. 6F(1) and 6F(2), respectively.

[0130] A method for visually indicating an audio processing state is implemented in the electronic device 104. The electronic device 104 collects audio input from an environment proximate the electronic device via one or more microphones 602 and processes the audio input. The processing includes one or more of identifying audio input from a user in the environment and responding to the audio input. The electronic device 104 determines a state of the processing from among a plurality of predefined audio processing states. For each of the full-color LEDs 604, the electronic device 104 identifies a respective predefined LED lighting specification associated with the determined audio processing state. The lighting specification includes one or more of LED lighting duration, pulse repetition rate, duty cycle, color sequence, and brightness. In some embodiments, the electronic device 104 determines that the audio processing state (including, in some embodiments, a state of a focus session) is associated with one of the plurality of users and identifies the predefined LED lighting specification for the full-color LEDs 604 by customizing at least one of the predefined LED lighting specifications (e.g., color sequence) of the full-color LEDs 604 according to the identity of the one of the plurality of users.

[0131] Further, in some embodiments, the colors of the full-color LEDs include a set of predetermined colors according to the determined audio processing state. For example, referring to Figures 6G(2), 6G(4), and 6G(7)-(10), the set of predetermined colors includes Google brand colors including blue, green, yellow, and red, and the array of full-color LEDs is divided into four quadrants, each associated with one of the Google brand colors.

[0132] According to the identified LED lighting specifications of the full-color LEDs, the electronic device 104 synchronizes the lighting of the array of full-color LEDs to match the determined audio processing state (in some embodiments, In some embodiments, the visual pattern indicative of the audio processing state includes a plurality of discrete LED illumination pixels. In some embodiments, the visual pattern includes a start segment, a loop segment, and an end segment. The loop segment lasts for a period related to the LED illumination duration of the full color LEDs and is configured to match the length of the audio processing state (e.g., the duration of the active focus session).

[0133] In some embodiments, the electronic device 104 has more than 20 different device states (including a plurality of predefined speech processing states) represented by an LED design language. Optionally, the plurality of predefined speech processing states includes one or more of a hot word detection state, a listening state, a thinking state, and a response state. In some embodiments, the plurality of predefined speech processing states includes one or more focus session states, as described above.

[0134] Reference has been made above to the embodiments, examples of which are illustrated in the accompanying drawings. In the foregoing detailed description, numerous specific details have been set forth in order to provide a thorough understanding of the various embodiments being described. However, it will be apparent to those skilled in the art that the various embodiments described may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0135] In some instances, terms such as first, second, etc. may be used herein to describe various elements, but it will also be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first device can be referred to as a second device, and similarly, a second device can be referred to as a first device, without departing from the scope of various described embodiments. The first device and the second device are both types of devices, but are not the same device.

[0136] The terminology used in the description of the various embodiments described herein is for the purpose of describing particular implementations only and is not intended to be limiting. As used in the description of the various implementations described and in the appended claims, the terms "a," "an," and "the" are used interchangeably. Numerals are intended to include the plural unless the context clearly indicates otherwise. The term "and / or" as used herein is also understood to refer to and include any and all possible combinations of one or more of the associated listed items. It is further understood that the terms "comprises," "including," "comprising," and / or "comprising," as used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0137] The term "if" as used herein is interpreted, optionally depending on the context, to mean "when" or "then" or "in response to determining" or "in response to detecting" or "following a determination that." Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" is interpreted, optionally depending on the context, to mean "upon determining" or "in response to determining" or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]" or "following a determination that [a stated condition or event] is detected."

[0138] In situations where the above-described systems collect information about the user, the user is given the opportunity to opt-in or opt-out of programs or features that may collect personal information (e.g., information about the user's preferences or use of the smart device). In some embodiments, some data is anonymized in one or more ways before it is stored or used, such that personally identifiable information is removed. For example, a user's identity may be anonymized so that personally identifiable information cannot be determined or associated with a user, and user preferences or user interactions are generalized (e.g., generalized based on user statistics) rather than associated with a particular user.

[0139] Although some of the various figures depict multiple logical steps in a particular order, steps that are not order dependent may be reordered and other steps may be combined or separated. While certain reorderings or other groupings are specifically mentioned, others will be apparent to those of ordinary skill in the art, and thus the ordering and groupings presented herein are not an exhaustive list of alternatives. Furthermore, it should be recognized that steps could be implemented in hardware, firmware, software, or any combination thereof.

[0140] The foregoing description has been provided for illustrative purposes with reference to specific implementations. However, the illustrative discussion above is not intended to be exhaustive or to limit the scope of the claims to the precise form disclosed. Numerous modifications and variations are possible in light of the above teachings. The implementations have been selected to best explain the principles underlying the claims and their practical application, thereby enabling those skilled in the art to best employ the implementations with various modifications as appropriate for the particular use contemplated.

Claims

1. 1. A method performed by a first electronic device including one or more microphones, a speaker, one or more processors, and a memory storing one or more programs for execution by the one or more processors, the first electronic device being a member of a local group of connected electronic devices communicatively coupled to a common network service executed by a server, the first electronic device comprising: Receiving a first voice command including a request for a first operation; assigning a first target device from the local group of connected electronic devices as a focused device for performing the first operation; causing the first operation to be performed by the first target device through operation of a common network service performed by the server in accordance with the assignment of the first target device as the focused device for performing the first operation; receiving a second voice command including a request for a second operation; determining that the second voice command does not include an explicit designation of a second target device; determining that the second operation can be performed by the first target device; (i) in accordance with the determination that the second voice command does not include an explicit designation of a second target device, and (ii) in accordance with the determination that the second operation can be performed by the first target device; assigning the first target device as the focused device for performing the second operation; causing the second operation to be performed by the first target device via operation of a common network service performed by the server in accordance with the assignment of the first target device as the focused device for performing the second operation; In the case where the first target device is the first electronic device, the method further comprises: receiving a third voice command including a request for a third operation and an explicit designation of a third target device within the local group of connected electronic devices, the third target device being a second electronic device within the local group of connected electronic devices, the second electronic device being different from the first electronic device; maintaining an assignment of the first target device as the focused device; and causing the third operation to be performed by the third target device via operation of the common network service.

2. assigning the first target device as the focused device for performing the first operation includes: obtaining an explicit designation of the first target device from the first voice command; and determining the first target device for the first operation from among the local group of connected electronic devices based on the explicit designation.

3. assigning the first target device as the focused device for performing the first operation includes: determining that the first voice command does not include an explicit designation of the first target device; determining that the first operation can be performed by a third electronic device in the local group of connected electronic devices; and selecting the third electronic device as the first target device.

4. The method of any one of claims 1 to 3, further comprising maintaining an assignment of the first target device as the focused device according to a determination that the second voice command satisfies a session maintenance criterion.

5. Assigning the first target device as the focused device comprises: storing a timestamp of the first voice command; The method of any one of claims 1 to 4, comprising storing an identifier of the first target device.

6. In the case where the first target device is not the first electronic device, the method further comprises: receiving a fourth voice command including a request for a fourth operation and an explicit designation of a fourth target device within the local group of connected electronic devices; ceasing to assign the first target device as the focused device; and assigning the fourth target device as the focused device; The method of any one of claims 1 to 5, further comprising causing the fourth operation to be performed by the fourth target device via operation of the common network service.

7. the second voice command is received after causing the third operation to be performed by the third target device; the first operation is a media playback operation; the second operation is a media stop operation; The method comprises: receiving a fifth voice command including a request for a fifth operation and an explicit designation of a fifth target device within the local group of connected electronic devices, the fifth target device being the second electronic device, the method further comprising: ceasing to assign the first target device as the focused device; and assigning the fifth target device as the focused device; The method of any one of claims 1 to 6, further comprising causing the fifth operation to be performed by the fifth target device via operation of the common network service.

8. 1. An electronic device comprising: one or more microphones; Speaker, one or more processors; and a memory for storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions, the instructions comprising: Receiving a first voice command including a request for a first operation; assigning a first target device from a local group of connected electronic devices as a focused device for performing the first operation; causing the first operation to be performed by the first target device via operation of a common network service performed by a server in accordance with the assignment of the first target device as the focused device for performing the first operation; receiving a second voice command including a request for a second operation; determining that the second voice command does not include an explicit designation of a second target device; determining that the second operation can be performed by the first target device; (i) in accordance with the determination that the second voice command does not include an explicit designation of a second target device, and (ii) in accordance with the determination that the second operation can be performed by the first target device; assigning the first target device as the focused device for performing the second operation; causing the second operation to be performed by the first target device through operation of a common network service performed by the server in accordance with the assignment of the first target device as the focused device for performing the second operation; In the case where the first target device is the electronic device, the one or more programs include: receiving a third voice command including a request for a third operation and an explicit designation of a third target device within the local group of connected electronic devices, the third target device being a second electronic device within the local group of connected electronic devices, the second electronic device being different from the first electronic device; maintaining an assignment of the first target device as the focused device; and and causing the third operation to be performed by the third target device via operation of the common network service.

9. The one or more programs are receiving a fifth voice command including a request to end a predetermined operation; In response to receiving the fifth voice command, ceasing execution of the first operation by the first target device; and and ceasing to assign the first target device as the focused device.

10. the first operation is a media playback operation; 10. The electronic device of claim 8 or 9, wherein the second action is one of a media stop action, a media rewind action, a media fast forward action, a volume up action, and a volume down action.

11. the first operation is a device state change operation to a first state among a plurality of device states; The electronic device according to claim 8 or 9, wherein the second operation is a device state changing operation to a second state of a plurality of device states.

12. the first operation is a magnitude change operation in a first direction on a magnitude scale; 10. The electronic device according to claim 8 or 9, wherein the second operation is a magnitude change operation in a second direction on the magnitude scale opposite to the first direction.

13. further comprising an array of one or more LEDs; The one or more programs are 13. The electronic device of claim 8, further comprising instructions for indicating a focused state of the electronic device by illuminating one or more of the LEDs in the array of LEDs.

14. A computer readable program comprising instructions that, when executed by a first electronic device comprising one or more microphones, a speaker, and one or more processors, cause the first electronic device to perform operations of a method, the method comprising: Receiving a first voice command including a request for a first operation; assigning a first target device from a local group of connected electronic devices as a focused device for performing the first operation; causing the first operation to be performed by the first target device via operation of a common network service performed by a server in accordance with the assignment of the first target device as the focused device for performing the first operation; receiving a second voice command including a request for a second operation; determining that the second voice command does not include an explicit designation of a second target device; determining that the second operation can be performed by the first target device; (i) in accordance with the determination that the second voice command does not include an explicit designation of a second target device, and (ii) in accordance with the determination that the second operation can be performed by the first target device; assigning the first target device as the focused device for performing the second operation; causing the second operation to be performed by the first target device through operation of a common network service performed by the server in accordance with the assignment of the first target device as the focused device for performing the second operation; In the case where the first target device is the first electronic device, the method further comprises: receiving a third voice command including a request for a third operation and an explicit designation of a third target device within the local group of connected electronic devices, the third target device being a second electronic device within the local group of connected electronic devices, the second electronic device being different from the first electronic device; maintaining an assignment of the first target device as the focused device; and causing the third operation to be performed by the third target device via operation of the common network service.

15. Assigning the first target device as the focused device comprises:

15. The computer readable program of claim 14, comprising determining that the second voice command is received by the first electronic device within a first predetermined time interval relative to receipt of the first voice command or at a second predetermined time interval relative to receipt of the first voice command, the second predetermined time interval following the first predetermined time interval.

16. Assigning the first target device as the focused device comprises: determining that the second voice command is received within the first predetermined time interval; and refraining from outputting a request to identify the first target device as the target device for the second voice command.

17. Assigning the first target device as the focused device comprises: determining that the second voice command is received within the second predetermined time interval; outputting a request to identify the first target device as a target device for the second voice command; and receiving a positive confirmation of the first target device in response to the request to confirm.

18. the second voice command is received after causing the third operation to be performed by the third target device; the first operation is a media playback operation; the second operation is a media stop operation; The method comprises: receiving a fifth voice command including a request for a fifth operation and an explicit designation of a fifth target device within the local group of connected electronic devices, the fifth target device being the second electronic device, the method further comprising: ceasing to assign the first target device as the focused device; and assigning the fifth target device as the focused device; 18. The computer readable program of claim 14, further comprising causing the fifth operation to be performed by the fifth target device via operation of the common network service.

19. 8. The method of claim 1, further comprising: the first electronic device identifying users inputting the first and second voice commands; and establishing a different focus session for each identified user.

20. The one or more programs are 14. The electronic device of claim 8, further comprising instructions for identifying users who input the first voice command and the second voice command and establishing a different focus session for each identified user.

21. 19. The computer readable program of claim 14, wherein the method further comprises identifying users inputting the first voice command and the second voice command, and establishing a different focus session for each identified user.

Citation Information

Patent Citations

  • Remote control system, controller, program for imparting function of controller to computer, storage medium with the program stored thereon, and server

    JP2006033795A

  • Controller for electronic apparatus, and control method therefor

    JP2007243602A

  • Remote controller, remote control system and remote control method

    JP2009044609A

  • Voice operation system for plural devices, voice operation method and program

    JP2015201739A